logo[мetahunt]
> Djinni

Data Engineer

pythonsqlapache sparkkafkaairflowsap s/4hanadatabrickspostgresql
sap btpweaviatemilvusdbtdagsterprefectdockerkubernetes
experience7+ years
domainDefTech
> full description

Responsibilities:
- Architect and implement data lakehouse solutions to centralize and harmonize supply chain
and procurement data from multiple enterprise systems
- Design and deploy data ingestion pipelines for structured and unstructured data, including
ERP sources (SAP S/4HANA, Ariba), external market feeds, and technical documents
- Develop and maintain unified data models and taxonomies to support analytics and AI-driven
forecasting
- Build and optimize pipelines for processing unstructured data (PDFs, CAD files, regulatory
documents) into formats suitable for AI and RAG applications
- Manage and optimize vector databases to enable high-speed retrieval of engineering and
procurement data for generative AI tools
- Establish and enforce data lineage, traceability, and governance protocols to ensure data
integrity and compliance
- Implement and monitor data quality controls to validate completeness and accuracy of critical
datasets

- Collaborate with cross-functional teams to map enterprise data sources and define
requirements for AI and analytics use cases
- Optimize data workflows for secure, on-premise, and air-gapped environments, ensuring
efficient use of infrastructure
- Support the technical execution of foundational data platform initiatives within structured sprint
cycles
 

Skills:
- Expert proficiency in Python, SQL, and modern data engineering frameworks (Apache Spark,
Kafka, Airflow).
- Enterprise ERP: Strong experience extracting data from complex ERP environments,
specifically SAP S/4HANA and SAP Ariba. Familiarity with SAP BTP is a plus.
- Database Technologies: Deep understanding of Data Lakehouse architectures
(Databricks/Delta Lake), Relational Databases (PostgreSQL), and Vector Databases
(Weaviate/Milvus).
- Data Pipeline Development: Experience building pipelines for RAG solutions, Conversational
agents and classical ML models with tools like dbt, dagster, or prefect
- DevOps/DataOps: Proficiency with containerization (Docker, Kubernetes) and CI/CD pipelines
for deploying data workflows in secure environments.
- Experience: 7+ years of experience in Data Engineering, with at least 2 years focused on
building pipelines for Machine Learning or Generative AI applications in an enterprise setting.
- Domain Knowledge: Experience in Supply Chain, Manufacturing, or Defense sectors is highly
desirable. Ability to understand "Bill of Materials" (BOM) structures and procurement lifecycles.
- Problem Solving: Ability to navigate the "Governance Collision" between agile data work and
rigid systems engineering requirements, ensuring data deliverables meet formal Stage Gate
reviews.
- Collaboration: Proven ability to work alongside Data Scientists and Backend Engineers to
define data schemas that support predictive modeling and AI agents.