logo[мetahunt]
> Djinni
senior

Data Engineer

Deep Knowledge Group
pythonairflowsqlpostgresqldockerci/cdllm
dagsterprefectknowledge graphsvector searchrag
test task:yes
domainAI
> full description

About Deep Knowledge Group

Deep Knowledge Group (https://www.dkv.global/) builds large-scale analytical databases, investment-intelligence systems, industry ecosystem platforms, competitiveness indices and interactive dashboards across frontier industries including AI, Longevity, HealthTech, DeepTech and QuantumTech.

The Data Science department builds and maintains the industry intelligence databases behind those products — organisations, people, products and the relationships between them, across multiple sectors and regions. Every one of them starts as messy public data and has to end up as a clean, deduplicated, verifiable dataset that analysts and downstream products can rely on.
 

About the Role

We currently build each database the way the last one was built. Collectors are source-specific scripts, scheduling is manual, quality is whatever the person who built it checked, and nothing links one database to another. That works at our current size and will not work at the next one.

We are looking for a Senior Data Engineer to own the shared platform that replaces this: one collection and enrichment framework that every database runs on, with orchestration, monitoring, provenance and entity resolution as properties of the system rather than as things individual engineers remember to do.

This is an ownership role. You will work directly with the Head of Data Science on architecture, and you will be the person other engineers bring design decisions to. We are not looking for an extra pair of hands; we are looking for a second point of technical decision-making in the department.
 

Key Responsibilities

  • Design and own the shared collection and enrichment framework: what is reusable, what stays source-specific, and where the line between them sits.
  • Own orchestration end to end — scheduling, dependencies, retries, backfills, incremental and idempotent re-runs.
  • Make entity resolution a permanent service rather than a per-project effort: blocking, matching, thresholds, and a defensible account of why a merge was made.
  • Design the provenance model: every value traceable to its source, its retrieval date and the logic that transformed it, six months after the fact.
  • Establish the staging-and-promotion rule and enforce it: nothing writes to production tables directly.
  • Build monitoring and observability for data, not only for jobs — freshness, coverage, null rates, distribution drift, source schema changes.
  • Define the quality standard: accuracy, completeness, duplicate rate, freshness, provenance — and the instrument that measures it.
  • Set the schemas, conventions and review standards that the rest of the department works to.
  • Design the deterministic backbone of our agent network, and decide where an agent is warranted and where plain code is cheaper, faster and testable.
  • Migrate existing databases onto the shared framework alongside the engineers who built them.
  • Mentor and unblock other engineers; take architectural decisions off the Head of Department's desk.
     

Required Skills

  • Strong Python — you write services and pipelines, not notebooks.
  • You have owned a data platform, not just tasks on one — orchestration, scheduling, monitoring, backfills.
  • Airflow, Dagster or Prefect in production, with a considered view on why that one.
  • SQL and PostgreSQL in depth — indexing, query plans, schema design, migrations.
  • Entity resolution or record linkage across sources that disagree with each other, including how you measured whether it worked.
  • You have set the standards other people then worked to: schemas, conventions, review.
  • Docker, CI/CD, and comfort running production systems on plain infrastructure rather than a managed platform.
  • Practical experience with LLM APIs and with validating non-deterministic output before it reaches a production table.
  • Data quality discipline: you can state a dataset's duplicate rate or field staleness with evidence, not by inspection.
  • The ability to write a design down clearly enough that someone can argue with it.
     

Strong Advantages

  • Production collection from public web sources: anti-bot, JS-rendered pages, undocumented APIs, source drift.
  • Knowledge graphs, ontologies and canonical entity models across multiple domains.
  • Data lineage and governance tooling in production, used by people other than its author.
  • Agent orchestration frameworks, and a clear view of their limits.
  • Vector search and RAG pipelines.
  • Experience turning a team's ad-hoc scripts into a framework that team then adopted.
  • OSINT methodology and source verification.
     

What We Will Evaluate

Two things, mostly.

You question your own data. The engineers who work out well here are the ones who notice that a source silently changed its schema, that a 97% match rate is hiding a systematic bias, that a field is populated but wrong. Volume without verification is worth nothing to us.

You explain your reasoning. Much of this work involves judgment calls that nobody can check quickly — which record wins a merge, what confidence threshold to use, when to stop enriching. We need those decisions written down and defensible, not buried in a script.

The hiring process is an intro call, a short practical task on real data, a technical review of your solution, and an offer. We evaluate the task on judgment, not polish: reproducibility, honest handling of edge cases, and clear reasoning about what you chose not to do count for more than a clever one-liner.