> DOUABCloudz
Data Engineer
format:Remotecompany:Outsource
pythonairflowsqlaws s3athenaaws iambigqueryprestotrinoany one ofbigqueryci/cdmlopsetldata modelingdockerlinux
kuberneteskafkakinesis
> full description
We’re looking for a skilled AI and Data Engineer to join our team. You’ll be at the forefront of designing, building, and maintaining the infrastructure that powers our data-driven products and machine learning models. This role requires a strong blend of data engineering expertise, a deep understanding of MLOps principles, and the ability to work with large-scale datasets. You will be responsible for creating robust data pipelines, managing our data platforms, and collaborating with data scientists and software engineers to bring AI solutions to life.
Requirements
- Higher education in Computer Science, Statistics, Machine Learning, Data Science, or a related quantitative field
- 2+ years of experience developing technical solutions for automation or AI
- Strong programming skills in Python and experience with Big Data applications
- Strong hands-on experience with Apache Airflow, including writing DAGs from scratch, managing dependencies and scheduling, and debugging failed production runs
- Strong SQL skills, including complex transformations, query optimization, and working with large analytical datasets
- Several years of hands-on experience designing, building, and supporting reliable production data pipelines
- Hands-on experience with AWS data services, particularly Amazon S3, Athena, and IAM, or equivalent experience querying file-based datasets using Presto, Trino, or BigQuery
- Strong understanding of data lake fundamentals, including partitioned Parquet datasets, partitioning strategies, incremental loads, and schema evolution
- Solid understanding of data lifecycle management, data architecture, and governance
- Experience with CI/CD pipelines and MLOps tools
- Experience with data modelling, ETL, and analytics-ready datasets
- Working knowledge of Docker and containerized development/deployment workflows
- Comfortable working with and troubleshooting applications and data workloads in Linux environments
- Excellent communication skills with the ability to explain complex technical concepts to non-technical stakeholders
Will be a Plus
- Experience with Docker and Kubernetes for containerization and orchestration
- Knowledge of real-time data processing technologies like Apache Kafka or Amazon Kinesis
- Familiarity with data observability, testing, and automation frameworks
Responsibilities
- Design, build, and maintain scalable data pipelines using tools like Apache Airflow, and AWS Glue
- Manage and optimize cloud-based data platforms, including data lakes (S3) and data warehouses (Redshift, BigQuery)
- Solve complex data integration challenges using optimal ETL patterns, sourcing from structured and unstructured data
- Develop and implement MLOps frameworks to automate deployment, monitoring, and lifecycle management of ML models
- Transition experimental models into production-ready code and integrate them into applications with software engineering teams
- Apply best practices in CI/CD, using tools like Jenkins, GitLab CI, and MLflow or Kubeflow
- Build interactive dashboards and data products to support analytics and decision-making
- Collaborate with product, QA, architecture, and engineering teams to deliver cleanly integrated data services
- Support teams with robust, well-documented data models and infrastructure
- Continuously monitor and improve the performance and cost-efficiency of data and AI infrastructure
- Drive innovation by evaluating open-source tools and emerging technologies, and recommending solutions
- Influence product and cross-functional teams to identify data opportunities that drive measurable impact