Data Engineer
About the work
We build and maintain industry intelligence databases - organizations, people, products and their relationships across sectors such as HealthTech, QuantumTech, LegalTech and wellness, covering multiple regions. Every one of them starts as messy public data and has to end up as a clean, deduplicated, verifiable dataset that analysts and downstream products can rely on.
That means the hard problems here are not "write a scraper". They are: how do you know two records are the same organization? How do you prove where a field came from six months later? How do you re-run a collection over a source that changed its markup, without corrupting what you already have?
You will work directly with the Head of Data Science and a team of AI and data engineers.
What you will do
- Build and own collection pipelines over public web sources - structured, semi-structured and awkward
- Design the enrichment layer: normalization, deduplication, entity resolution, confidence scoring
- Keep provenance on every field, so any value can be traced back to its source and date
- Use LLMs where they genuinely help - extraction from unstructured text, classification, taxonomy assignment - and know where they don't
- Take datasets from one-off collection to scheduled, monitored, re-runnable pipelines
- Define and enforce quality checks: coverage, freshness, duplicate rate, field completeness
What we're looking for
Both levels
- Strong Python - you write services and pipelines, not notebooks
- Real production scraping experience: sessions, rate limits, retries, blocks, anti-bot, JS-rendered pages, APIs that are not documented
- SQL and PostgreSQL beyond CRUD - indexing, query plans, schema design
- Practical experience with LLM APIs for extraction and classification, including how you validate their output
- Data quality instinct: you check what you collected before you hand it over
Additionally for the Senior opening
- You have owned a data platform, not just tasks on one - orchestration, scheduling, monitoring, backfills
- Experience with entity resolution or record linkage across sources that disagree with each other
- Airflow, Dagster or Prefect in production
- You have set the standards other people then worked to: schemas, conventions, review
Nice to have
- Docker and CI/CD
- Knowledge graphs, ontologies, or taxonomy design
- Vector search and RAG pipelines
- OSINT methodology and source verification
- Experience with data licensing or compliance constraints on public-web collection
How we hire
- Intro call - 30 minutes, mutual fit and what you've actually built
- Paid proof task - a real, bounded piece of our work. Paid at market rate, roughly 4 hours. We are not asking anyone to work for free
- Technical review - we go through your solution together: your trade-offs, what you'd do with more time, what you'd do differently at ten times the volume
- Offer - band determined by the level you land at, not by the title you applied under
We evaluate the proof task on judgment, not polish. Reproducibility, honest handling of edge cases, and clear reasoning about what you chose not to do count for more than a clever one-liner.
What we care about
Two things, mostly.
You question your own data. The engineers who work out well here are the ones who notice that a source silently changed its schema, that a "97% match rate" is hiding a systematic bias, that a field is populated but wrong. Volume without verification is worth nothing to us.
You explain your reasoning. Much of this work involves judgment calls that nobody can check quickly - which record wins a merge, what confidence threshold to use, when to stop enriching. We need those decisions written down and defensible, not buried in a script.
To apply: send your CV plus one short paragraph on the hardest data collection or deduplication problem you have solved, and what you got wrong on the way.