We are looking for a DevOps / MLOps Engineer to join our client’s distributed engineering team and help build and maintain the cloud infrastructure that powers AI-driven applications.
In this role, you will focus on infrastructure, deployment pipelines, automation, and observability, ensuring reliability, scalability, and efficient operations across production environments that support AI workloads.
Please note that this role is focused on infrastructure and platform engineering. You will not work on proprietary ML models or training assets.
Our client is a fast-growing US-based AI-first startup building next-generation technology for the fashion and e-commerce industry.
Their platform leverages artificial intelligence to deliver photorealistic virtual try-on, AI-powered sizing prediction, and personalized outfit recommendations, helping leading fashion brands create more engaging shopping experiences.
The working hours: Maintain ≥4 hours/day overlap with US Pacific working hours to ensure overlap with the US team
Experience / Skills required:
Must have:
2+ years of professional experience as a DevOps or MLOps EngineerStrong production experience with AWS, Kubernetes (EKS), Docker, and TerraformExperience with Slurm for workload orchestration in production environmentsExperience building and maintaining CI/CD pipelines using GitHub ActionsExperience with monitoring and observability tools, including Grafana, Prometheus, and AWS CloudWatchPython scripting experience for automation and infrastructure toolingExperience supporting production cloud infrastructureStrong communication skills and ability to collaborate with distributed teamsUpper-Intermediate English or higher
Good to have:
Experience supporting GPU-based infrastructure or AI/ML workloadsExperience with MLflowExperience with FinOps or cloud cost optimization
Responsibilities:
Build, manage, and optimize AWS EKS/Kubernetes infrastructure supporting production servicesContainerize and deploy applications using DockerProvision and maintain infrastructure using TerraformSupport GPU-enabled infrastructure and workload orchestration with SlurmBuild and maintain CI/CD pipelines using GitHub ActionsImplement and maintain monitoring and observability using Grafana, Prometheus, and AWS CloudWatchDevelop Python scripts and automation tools to improve operational efficiencyMonitor cloud infrastructure utilization and contribute to cloud cost optimization initiativesCollaborate closely with software engineers to ensure reliable and scalable platform operationsParticipate in code reviews and follow engineering, security, and infrastructure best practices
We offer:
Competitive salaryVacation (up to 20 working days)Paid sick leaves (10 working days)National Holidays as paid time offFlexible working schedule, remote formatDirect cooperation with the customerDynamic environment with low level of bureaucracy and great team spiritChallenging projects in diverse business domains and a variety of tech stacksCommunication with Top/Senior level specialists to strengthen your hard skillsOnline teambuildings
Відгукнутись на вакансію