logo[мetahunt]
> Djinni
senior

Data Engineer

ZentixSoft6.4k USD
format:Remotetype:Contract
pythonsqlawsaws s3aws step functionscloudwatchocrsharepointci/cddata pipelines
vector databasesragazuredatabricks
englishC1–C2
experience7+ years
> full description

 

ZentixSoft is looking for two Senior Data Ingestion Engineers! 🚀

Format: Direct Contract Engagement (1-Year Contract, 100% Full-Time Remote). 
 

💡 Why ZentixSoft Partner Network?
 💸 Transparent Compensation: Direct payroll arrangement with $40 USD/hour base compensation. 
⚖️ Work-Life Balance: Sustainable engineering workflows with predictable long-term deliverables. 
🦾 Trust & Transparency: Zero micromanagement and full autonomy over your technical pipeline architecture. 
🎁 Culture & Growth: Long-term project stability within large-scale financial and insurance data domains.
 

🧩 Responsibilities:

  • Pipeline Design & Engineering: Design and deploy scalable data ingestion pipelines for high-volume structured/unstructured documents (PDFs, scans, emails, Word, Excel, PowerPoint).
  • Document Extraction & OCR: Implement OCR and document processing workflows using AWS Textract (or equivalent) for text extraction, cleaning, normalization, and metadata tagging.
  • RAG & Vector Storage Prep: Execute semantic chunking, metadata extraction, vector storage schema design, and retrieval mechanism preparation for downstream AI models.
  • Integrations & Connectors: Build robust connectors with enterprise sources, including SharePoint, email servers, and public cloud repositories.
  • Validation & Observability: Implement automated error monitoring, OCR extraction validation, and CI/CD automated testing using AWS Step Functions and CloudWatch.
     

🎓 Our Perfect Match (Requirements):

  • Senior Data Engineering: 7+ years of commercial Data Engineering experience, with strong proficiency in Python and SQL.
  • AWS Expertise: 5+ years of hands-on experience in AWS environments, specifically AWS S3, Step Functions, CloudWatch, and public cloud data processing.
  • Unstructured Data & OCR: Proven background in building document extraction pipelines (handling PDFs, scanned images, emails, Office documents) and utilizing OCR technology (AWS Textract or similar).
  • Source Connectors & Pipeline Operations: Practical experience integrating sources like SharePoint and email, performing text normalization, metadata tagging, and maintaining CI/CD/Git testing standards.
  • EU Residency & Location: Candidates MUST reside in an EU member country (with active legal residency; citizenship can be non-EU/global).
  • Language & Communication: C1 Advanced English (verbal and written) for direct technical collaboration.
     

➕ Good to Have:

  • Experience with Vector Databases, RAG architectures, semantic chunking, and vector retrieval mechanisms.
  • Background in Insurance or Financial Services industries handling enterprise security standards.
  • Exposure to Azure or Databricks platform components.
     

📍 Project Details:

⏳ Duration: 1 Year (Full-time contract).

🚀 Target Start: End of October.

📍 Location: 100% Remote (Must be physically located in an EU member state with valid residency).