Sofia, Bulgaria
Our client is the leading source of actionable intelligence for the global insurance, reinsurance and ILS markets. They are building a new AI-powered data product and need someone to help them turning their growing pool of data into the structured, ready-to-use formats that power search, personalisation and new features.
They need a Data Scientist with a software or data engineering background who has moved into data science over the past few years and wants to keep developing in that direction. Or a generalist who's happy moving between engineering and data science work rather than a deep specialist in either. They'll soon be bringing in a lot of new data and new data types from external sources, which will make this data structuring and manipulation work even more central to the role.
Your day-to-day impact:
✔ Turn raw, messy or unstructured data - article text, usage data, and new data arriving from external sources - into clean, ready-to-use formats for downstream products.
✔ Apply light ML/NLP techniques (e.g. entity extraction, classification, tagging) to structure unstructured text.
✔ Adapt existing pipelines to ingest and normalise new data types as they come online.
✔ Bring in new datasets - scraped or paid-for/third-party - and merge them with our existing data to create new, value-add content sets.
✔ Build personalisation logic that blends usage data with content suggestions, helping surface the right content to the right customer.
✔ Work with product and editorial to define the signals and rules behind personalised recommendations and content creation.
Data Science tasks:
✔ Use statistical and ML methods (e.g. classification, clustering, embeddings) to analyse usage and content data, surfacing patterns and insights that inform personalisation and product decisions.
✔ Use LLMs to extract structure - keywords, entities, topics - from unstructured article text, and check the accuracy and quality of that output.
✔ Prototype and test simple models or scoring logic (e.g. for recommendations or content matching), working with product to validate they add value before scaling.
✔ Try different approaches to a data or personalisation problem, evaluate what works, and explain the trade-offs in plain terms to non-technical stakeholders.
Data Engineering tasks:
✔ Maintain and extend our Databricks/Python pipelines - ingestion, transformation, scheduling and monitoring.
✔ Identify new data sources relevant to the product and help scope their integration.
AI-Assisted Development:
✔ Use AI coding tools (e.g. GitHub Copilot, Claude, Cursor) to build and iterate on data pipelines and structuring work quickly.
✔ Work across product, editorial and commercial teams to understand data needs and communicate findings clearly to non-technical audiences.
✔ Monitor data quality in production, flagging issues and improving tooling and documentation.
Your profile looks like that:
✔ 3–5 years' hands-on experience in data engineering, with strong Python and SQL.
✔ Experience building and maintaining data pipelines, ideally with Databricks, Spark or a comparable cloud data platform.
✔ Some exposure to - or a strong interest in developing - data science/ML techniques, e.g. classification, clustering, NER, embeddings.
✔ Experience turning raw or messy data into structured, ready-to-use formats for downstream use.
✔ Comfortable working independently and using AI coding tools to move quickly.
You are a great fit if you have:
✔ Experience with personalisation, recommendation systems, or blending usage/behavioural data with content.
✔ Exposure to using LLMs to extract structure (keywords, entities, topics) from unstructured text.
✔ A degree in Computer Science, Data Science, Engineering or a related field - or equivalent experience.
✔ Experience with orchestration tools such as Airflow, Dagster or Databricks Workflows.
They offer, not just a nice team and environment, but:
✔︎ Flexibility with true hybrid working - expected to be in the office 1-2 days a week.
✔︎ 25 holiday days per year, plus your birthday off.
✔︎ Opportunities for professional growth and development.
✔︎ Competitive compensation and benefits package.
Are you still here? Great! Let’s discuss further at katrin@cadabra.bg
(Recruitment License № 2709/ 17.01.2019)