logo[мetahunt]
> Djinni
senior

Data Engineer

pythonsqlawsaws s3aws step functionscloudwatchdata pipelinesocrsharepointci/cd
vector databasesragazuredatabricks
englishC1–C2
experience7+ years
> full description

The Data/Ingestion Engineer will design and implement scalable ingestion pipelines for high-volume structured and unstructured insurance documents, including PDFs, scans, emails, Word, Excel and PowerPoint files. 

 

The role includes OCR/document extraction, text cleaning and normalization, semantic chunking, metadata tagging, vector storage/RAG preparation, validation, monitoring and integration with downstream AI solutions.

  • At least 7 years of experience as Data Engineer with Python and SQL
  • At least 5 with AWS, the rest should be more or less but should be real experience that the candidate is able to talk about and give a concrete example/
  • Must to have : 
    • Python
    • SQL
    • AWS
    • S3
    • Step Functions
    • CloudWatch
    • Data/document ingestion pipelines
    • Processing of unstructured documents:
    • PDF
    • scanned documents
    • Word
    • Excel
    • PowerPoint
    • emails
    • OCR / document extraction technologies
    • AWS Textract or equivalent
    • Connectors/integrations with sources such as:
    • SharePoint
    • email
    • Text cleaning and normalization
    • Metadata extraction/tagging
    • Automated error monitoring
    • Validation of OCR/extraction results
    • Git
    • CI/CD
    • Testing
    • Public-cloud data processing/document ingestion experience
    • Good to have :
    • Vector databases
    • RAG
    • Semantic chunking
    • Vector-storage schema design
    • Retrieval mechanisms
    • Azure
    • Databricks
    • AWS Textract specifically
    • Experience within insurance / financial services
    • Experience handling enterprise security and low-latency SLA requirements
  • English Level C1
  • Duration of the mission :  1 year