logo[мetahunt]
> DOU
senior

DevOps Engineer

EPAM
format:Remotecompany:Outsource
pythonopentelemetrynew relickubernetesazurelangchainmlopsterraformci/cd
databricksazure ai foundry
englishB2
experience4+ years
domainAI
locationUkraine (Kyiv, Kharkiv, Lviv, +8)
> full description

EPAM’s Operational Intelligence practice is building a new capability: AI Reliability Engineering (AIRE) — applying SRE principles and cloud-native practices to the lifecycle of production AI/ML and LLM systems. As clients move GenAI applications and agentic systems from pilot into production, they discover that classic APM tells them nothing about token latency, cost per request, semantic drift or hallucinations. That gap is what this role closes.

You will instrument, monitor and harden production AI systems, define AI-native service level objectives and build the accelerators and reference implementations our practice reuses across accounts. This is an engineering role, not an L1/L2 support role — no 24/7 on-call rotation.

What You’ll Get

  • A genuinely new discipline inside EPAM — you define how it is done, not follow an existing runbook
  • Engineering focus without 24/7 on-call rotation
  • Funded certification and enablement tracks (Anthropic/Claude, Databricks, AI & Data Observability learning paths)
  • Cross-client exposure and a direct path into presales and solution engineering

Kindly note that this role supports remote work, but only from within Ukraine.

Responsibilities

  • Instrument production LLM, RAG and agentic applications with AI telemetry based on OpenTelemetry and APM-native AI monitoring
  • Implement distributed tracing across multi-model chains, agent workflows and retrieval-augmented generation pipelines to profile systemic latency and failure points
  • Define and measure AI-native SLIs and SLOs: time to first token (TTFT), throughput, error and refusal rates, cost per request, semantic drift, hallucination boundaries, contextual accuracy
  • Set up structured semantic logging and prompt/response monitoring for quality analysis
  • Build and run evaluation loops for output quality and safety — golden sets, LLM-as-a-judge, Ragas/DeepEval-style frameworks — and wire them into CI/CD and runtime
  • Track token-based cloud spend, model API rate limits and quota consumption; drive AI cost optimization
  • Configure AI gateways for API load balancing, failover and fallback models across multiple LLM providers
  • Implement guardrails: prompt injection and jailbreak filtering, output compliance, bias and safety constraints
  • Design detection, triage, restore and problem management workflows for AI incidents; integrate autonomous AI agents into RCA to parse logs, form hypotheses and correlate state changes
  • Support rollback, canary and fail-safe patterns for model, prompt and configuration releases; maintain reproducibility through versioning of data, code, prompts and models
  • Build practice accelerators, reference architectures and internal enablement materials; support presales and client assessments

Requirements

  • 4+ years in SRE, DevOps, platform or observability engineering, including hands-on work with production AI/ML or LLM workloads
  • Solid SRE fundamentals: Golden Signals, SLI/SLO definition, error budgets and burn rate, incident lifecycle, ITIL basics
  • Strong Python for instrumentation, automation and evaluation tooling
  • Practical experience with OpenTelemetry and at least one APM/observability platform: New Relic, Datadog, Grafana LGTM stack, Splunk or Elastic
  • Hands-on production experience with at least one cloud platform (Azure preferred, AWS or GCP) and Kubernetes
  • Working understanding of LLM application architecture: prompts, embeddings and vector stores, RAG, agent orchestration (LangChain / LangGraph or equivalent)
  • MLOps awareness: model lifecycle (training vs. inference), model endpoints, containerization, deployment and rollback patterns
  • Infrastructure as Code with Terraform; CI/CD experience (Azure DevOps, GitLab CI, GitHub Actions)
  • B2+ English — the role is client-facing and requires clear written and spoken technical communication

Nice to have

  • Experience with AI-specific observability and evaluation tooling: Traceloop/OpenLLMetry, Langfuse, Arize Phoenix, Ragas, DeepEval, MLflow
  • Distributed inference serving at scale: vLLM, KServe, Ray Serve, Kubernetes-native LLM orchestration, GPU capacity planning
  • AI security: OWASP LLM Top 10, prompt injection defense, guardrail frameworks (NeMo Guardrails, Llama Guard)
  • Databricks (incl. Mosaic AI / MLflow) or Azure AI Foundry experience
  • Certifications: Anthropic Claude, Azure AI Engineer, AWS ML Specialty, Databricks GenAI
  • FinOps for AI workloads — token and GPU cost modeling
  • Data Reliability Engineering background (data quality, pipeline SLOs) — AI reliability starts with data reliability
  • Mentoring or team lead experience

We offer/Benefits

With us you can:

  • Work on a flexible schedule remotely or from any of our comfortable offices or coworking spaces in Ukraine
  • Receive the necessary equipment to perform your work tasks
  • Change projects and technology stacks within EPAM
  • Gain experience in various business domains (Insurance, E-commerce, Healthcare, Finance, Travelling, Media, Artificial Intelligence, and more)
  • Relocation opportunities may be available for eligible candidates, depending on the role and openings at other EPAM locations
  • Participate in volunteer, charity programs and communities (both technical and interest-based)

We focus on your professional growth:

  • You can plan your individual career path together with your manager
  • Receive regular feedback from colleagues
  • Improve your English for free with certified teachers (Speaking Clubs, client interview preparation courses, etc.)
  • Get the opportunity to undergo free training and certification in AWS, GCP, or Azure Clouds
  • Use the internal E-learn training program (18,200+ specialized training and mentoring programs)
  • Access corporate accounts on LinkedIn Learning, Get Abstract and other partner resources
  • Study at EPAM Solution Architecture School with the instructors who are practicing architects
  • Develop as a leader, join Delivery Management, Resource Management, Leadership Essentials school and more
  • Participate in internal communities (500+ meetups, technical discussions, brainstorming sessions, online events and conferences annually)

What we offer:

  • Vacation and sick leave (including a sick leave without a medical certificate)
  • A wide range of Voluntary Medical Insurance programs providing both medical treatment and various preventive options (including sports activities)
  • Medical insurance for family members at corporate rates
  • Company support during significant life events (childbirth or adoption, marriage, etc.)
  • Support for psychological comfort: discounts on services from mental health specialists or coaches, thematic training
  • E-kids program — a free programming language training program for EPAMers’ children

Kindly be advised that the set of benefits, including learning, certification, and other opportunities, may vary depending on the role you apply for. Our recruiter will be able to share more details about the specific opportunity during your general interview.

ABOUT EPAM

EPAM strives to provide its global team of over 62,350 professionals in more than 55 countries with opportunities for professional growth from day one of collaboration. Our colleagues are the source of EPAM’s success, so we value cooperation, strive to always understand our clients’ business and aim for the highest quality standards. No matter where you are, you will join a dedicated, diverse community that will help you realize your potential to the fullest.

Відгукнутись на вакансію