Data Engineer
About Zoral.
Zoral is an IT product and professional services company serving banks, insurers, wealth
managers and fintechs. Alongside its flagship product, Zoral fOS (Financial Operating System), Zoral
delivers data, analytics and AI programmes for regulated financial-services firms in Europe, the UK, the US and Asia.
About the role.
We are looking for a Senior Data Engineer to assess and build enterprise data platforms
on Azure Databricks for regulated financial-services clients. Typical platforms integrate tens of source
feeds (databases, files and APIs) and millions of customer and policy records, and feed several hundred reports. You review existing lakes and pipelines, profile and measure the quality of the data, prove cleansing and customer-matching approaches on real data, and then build and lead the build of production pipelines.
Responsibilities
Review an existing data lake, its pipelines, jobs, notebooks and code as deployed; assess code
quality, operability, testing and deployment practice.
Build the inventory of sources, feeds, entities, keys, volumes and refresh patterns.
Profile data at scale and establish a data-quality baseline (completeness, validity, consistency,
uniqueness, timeliness); propose a candidate set of data-quality rules.
Design and build ingestion (batch, change-data capture, files, APIs) and transformation pipelines
across layered architectures, including Data Vault 2.0 integration layers and Kimball dimensional
models.
Implement data cleansing and standardisation, de-duplication and customer matching (entity
resolution), and build a proof of concept on a data sample with measured precision and recall.
Implement data-quality rules as code: quality gates, quarantine of failed records, monitoring, and
reconciliation of outputs to source totals; capture audit and lineage information for every run.
Build idempotent, re-runnable and testable pipelines with automated tests and deployment through
CI/CD.
Review the work of other engineers, pair with client engineers and document runbooks and hand
over material.
Requirements
6+ years of data engineering, including at least three on Databricks in production.
Expert SQL and strong Python with PySpark.
Hands-on experience with Delta Lake, Unity Catalog, Lakeflow Declarative Pipelines (formerly Delta
Live Tables) and Lakeflow Jobs (formerly Workflows), Auto Loader and Azure Data Factory.
Data modelling: Data Vault 2.0 and dimensional modelling, slowly changing dimensions, history and
as-at data.
Data profiling, data-quality rule design and quality frameworks (pipeline expectations, Great
Expectations, Soda or similar).
At least one delivered entity-resolution or de-duplication project on customer data (Splink, Zingg,
Dedupe or a commercial master-data tool), including evaluation of match quality.
Git, code review, CI/CD and automated testing of data pipelines.
Experience of migrating pipelines and data from legacy platforms, including reconciliation.
Good written and spoken English; willingness to work on site at client premises (including in the
UK) for periods of several weeks.
Nice to have
Databricks Asset Bundles, Terraform.
Metadata-driven pipeline and model generation.
Power BI semantic models and Databricks SQL.
Address and name standardisation; working with personal data under masking and
pseudonymisation.
Experience in banking, insurance or wealth management.
Databricks Data Engineer Professional certification.
Tech stack:
Azure Databricks (Unity Catalog, Delta Lake, Lakeflow Declarative Pipelines, Lakeflow Jobs,
Lakeflow Connect, Auto Loader, Databricks SQL), Azure Data Lake Storage, Azure Data Factory, SQL,
Python/PySpark, Splink or similar, Great Expectations or Soda, Git, Azure DevOps or GitHub Actions,
Terraform, Power BI.
Conditions.
Client work in regulated financial services requires pre-assignment background screening
(identity, right to work, employment history, criminal-record, credit and sanctions checks, to the standard the client specifies) and adherence to Zoral’s information-security and data-protection policies.