Data Engineer
About the role
We're hiring two Senior Data Ingestion Engineers for a major client in the insurance sector. You'll design and implement scalable ingestion pipelines for high-volume structured and unstructured insurance documents: PDFs, scans, emails, Word, Excel and PowerPoint files.
The pipelines feed downstream AI solutions. The work covers OCR and document extraction, text cleaning and normalisation, semantic chunking, metadata tagging, vector storage and RAG preparation, validation, and monitoring.
What you'll do
- Design and build scalable document-ingestion pipelines on AWS (S3, Step Functions, CloudWatch).
- Process unstructured documents: PDFs, scanned documents, Word, Excel, PowerPoint and emails.
- Implement OCR and document extraction with AWS Textract or equivalent tools, and validate the extraction results.
- Build connectors and integrations with sources such as SharePoint and email.
- Clean and normalise text, and extract and tag metadata.
- Prepare data for RAG: semantic chunking, vector-storage schema design and retrieval.
- Set up automated error monitoring, alerting and quality checks.
- Integrate ingestion outputs with downstream AI solutions, following enterprise security and SLA requirements.
Must-have
- 7+ years of commercial experience as a Data Engineer with Python and SQL.
- 5+ years of hands-on AWS experience, including S3, Step Functions and CloudWatch.
- Proven experience building data and document ingestion pipelines in a public cloud.
- Processing of unstructured documents: PDF, scanned documents, Word, Excel, PowerPoint and emails.
- OCR and document extraction technologies (AWS Textract or equivalent), including validation of OCR and extraction results.
- Connectors and integrations with sources such as SharePoint and email.
- Text cleaning and normalisation, and metadata extraction and tagging.
- Automated error monitoring for data pipelines.
- Git, CI/CD and testing practices.
- English at C1 level.
For every skill above, you should be able to give concrete examples from real projects.
Nice to have
- Vector databases, RAG, semantic chunking, vector-storage schema design and retrieval mechanisms
- AWS Textract specifically
- Azure and/or Databricks
- Experience in insurance or financial services
- Experience with enterprise security and low-latency SLA requirements
Location
Fully remote, but you must live in an EU member country and hold a valid residence permit there.
What we offer
- A 1-year full-time engagement with a major enterprise client
- The option of a direct contract with the client (the details of the cooperation model are discussed individually)
- Hands-on work at the intersection of document processing, data engineering and GenAI
- Fully remote work within the EU
How to apply
Please send a CV tailored to this role. The client screens CVs strictly and won't review a generic CV or a list of keywords. For each relevant project, please include:
- start and end dates, the context, and your personal role and responsibilities
- the AWS services and tools you used (S3, Step Functions, CloudWatch, Textract and so on) and how you used them
- the document types you processed, how you did the OCR, extraction, cleaning, metadata tagging and validation, and which sources you integrated (SharePoint, email)
- any experience with RAG, vector databases, chunking or retrieval, plus your Azure, Databricks and insurance or financial-services experience
Please also include:
- your expected hourly rate (USD)
- your English level
- years of experience
- your country of residence and residence status
- your citizenship
- your availability date