logo[мetahunt]
> DOU

Data Engineer

ABCloudz
format:Remotecompany:Outsource
pythonairflowsqlaws s3athenaaws iambigqueryprestotrinoany one ofbigqueryci/cdmlopsetldata modelingdockerlinux
kuberneteskafkakinesis
experience2+ years
> full description

We’re looking for a skilled AI and Data Engineer to join our team. You’ll be at the forefront of designing, building, and maintaining the infrastructure that powers our data-driven products and machine learning models. This role requires a strong blend of data engineering expertise, a deep understanding of MLOps principles, and the ability to work with large-scale datasets. You will be responsible for creating robust data pipelines, managing our data platforms, and collaborating with data scientists and software engineers to bring AI solutions to life.

Requirements

  • Higher education in Computer Science, Statistics, Machine Learning, Data Science, or a related quantitative field
  • 2+ years of experience developing technical solutions for automation or AI
  • Strong programming skills in Python and experience with Big Data applications
  • Strong hands-on experience with Apache Airflow, including writing DAGs from scratch, managing dependencies and scheduling, and debugging failed production runs
  • Strong SQL skills, including complex transformations, query optimization, and working with large analytical datasets
  • Several years of hands-on experience designing, building, and supporting reliable production data pipelines
  • Hands-on experience with AWS data services, particularly Amazon S3, Athena, and IAM, or equivalent experience querying file-based datasets using Presto, Trino, or BigQuery
  • Strong understanding of data lake fundamentals, including partitioned Parquet datasets, partitioning strategies, incremental loads, and schema evolution
  • Solid understanding of data lifecycle management, data architecture, and governance
  • Experience with CI/CD pipelines and MLOps tools
  • Experience with data modelling, ETL, and analytics-ready datasets
  • Working knowledge of Docker and containerized development/deployment workflows
  • Comfortable working with and troubleshooting applications and data workloads in Linux environments
  • Excellent communication skills with the ability to explain complex technical concepts to non-technical stakeholders

Will be a Plus

  • Experience with Docker and Kubernetes for containerization and orchestration
  • Knowledge of real-time data processing technologies like Apache Kafka or Amazon Kinesis
  • Familiarity with data observability, testing, and automation frameworks

Responsibilities

  • Design, build, and maintain scalable data pipelines using tools like Apache Airflow, and AWS Glue
  • Manage and optimize cloud-based data platforms, including data lakes (S3) and data warehouses (Redshift, BigQuery)
  • Solve complex data integration challenges using optimal ETL patterns, sourcing from structured and unstructured data
  • Develop and implement MLOps frameworks to automate deployment, monitoring, and lifecycle management of ML models
  • Transition experimental models into production-ready code and integrate them into applications with software engineering teams
  • Apply best practices in CI/CD, using tools like Jenkins, GitLab CI, and MLflow or Kubeflow
  • Build interactive dashboards and data products to support analytics and decision-making
  • Collaborate with product, QA, architecture, and engineering teams to deliver cleanly integrated data services
  • Support teams with robust, well-documented data models and infrastructure
  • Continuously monitor and improve the performance and cost-efficiency of data and AI infrastructure
  • Drive innovation by evaluating open-source tools and emerging technologies, and recommending solutions
  • Influence product and cross-functional teams to identify data opportunities that drive measurable impact
Відгукнутись на вакансію