DevOps Engineer
Role in one line
Build and operate the self-hosted AI inference platform — Kubernetes, NVIDIA H100 infrastructure, vLLM, observability, GitOps, and production reliability.
Context
We are building a self-hosted LLM inference platform running on NVIDIA H100 GPUs and bare-metal Kubernetes. The platform provides vLLM inference workers and an OpenAI-compatible API for internal services.
There is no managed cloud Kubernetes layer — the engineer will work directly with the infrastructure and own the platform from cluster bootstrap through production operations.
What you will work on
- Bootstrap and maintain self-hosted Kubernetes clusters with HA control plane, etcd backup/restore and upgrades.
- Set up and maintain NVIDIA GPU infrastructure, including GPU Operator, DCGM and the CUDA/container runtime stack.
- Configure GPU sharing and scheduling using MPS or time-slicing.
- Deploy and operate vLLM inference workers with zero-downtime rolling updates and graceful connection draining.
- Build Prometheus/Grafana monitoring and alerting for inference SLIs such as TTFT, queue depth, KV-cache utilisation and rejection rate.
- Implement IaC for bare-metal provisioning and cluster bootstrap using Terraform or Pulumi.
- Set up GitOps workflows with ArgoCD or Flux for controlled rollouts and rollbacks.
- Configure API Gateway, token-aware rate limiting, NetworkPolicy and mTLS.
- Build automation and runbooks for GPU-specific production incidents.
Must-have
- 5+ years in DevOps, Platform Engineering or SRE with production Kubernetes.
- Hands-on experience with self-hosted Kubernetes (kubeadm, RKE2 or k3s), including bootstrap, upgrades, etcd and HA.
- Practical experience with NVIDIA GPUs in Kubernetes, GPU Operator/Device Plugin and the CUDA driver stack.
- Strong understanding of GPU sharing and scheduling: MPS, time-slicing and multi-process workloads.
- Strong Linux/system administration skills, including NUMA, huge pages, CPU management and bare-metal troubleshooting.
- Experience with zero-downtime deployments of long-running or stateful workloads.
- Strong Prometheus and Grafana skills, including recording and alerting rules.
- Experience with Terraform or Pulumi.
- Bash/Python scripting for automation, health checks and monitoring.
- Basic understanding of LLM inference metrics and concepts: TTFT, token throughput, KV-cache and batching.
Nice-to-have
- Experience with vLLM, Triton Inference Server, KServe or similar.
- Production GitOps experience with ArgoCD or FluxCD.
- Experience with Kong, Envoy, APISIX or similar API gateways.
- Experience with Vault or External Secrets Operator.
- CKA, CKS or NVIDIA certifications.
- Experience benchmarking GPU workloads under mixed-model load.
Tech stack you will touch
Kubernetes, NVIDIA H100, NVIDIA GPU Operator, DCGM, CUDA, containerd, vLLM · Prometheus, Grafana · Terraform / Pulumi · ArgoCD / Flux · Linux, Bash, Python · API Gateway, NetworkPolicy, mTLS · Docker, Git, CI/CD.
Ways of working
- Remote, distributed engineering team.
- Production-focused environment with ownership of the infrastructure and inference stack.
- The role involves building the platform from the ground up and keeping it reliable in production.
Important: As this is a Germany-based project, we are primarily seeking candidates based in Western Ukraine, with Vinnytsia and Lviv being our preferred locations. Other secure western-region locations can be considered case by case.
Відгукнутись на вакансію