logo[мetahunt]
> Djinni
senior

DevOps Engineer

format:Remotecompany:Outsource
linuxdockercudakubernetesobservability
pytorch
englishB2
domainAI
> full description

On behalf of our client, we are looking for a Senior Infrastructure Engineer (GPU Platform)

 

Requirements:

 

- Production experience with multi-node GPU training infrastructure

- Strong Linux, containers, CUDA, and NVIDIA GPU stack knowledge

- Hands-on experience with NCCL and InfiniBand or RoCE/RDMA troubleshooting

- Deep experience with Kubernetes or Slurm

- Experience with infrastructure automation and observability

- Experience diagnosing issues across training workloads, networking, storage, and GPU hosts

- Strong incident leadership and provider-facing communication skills

- English – Upper-Intermediate or higher

 

Would be a plus:

 

- Experience in an AI lab, HPC environment, or specialist GPU cloud

- PyTorch, Megatron, DeepSpeed, or other distributed-training frameworks

- Experience with parallel storage and checkpoint optimization

- Experience working with multi-provider GPU platforms

 

Company offers:

 

- Long-term employment with possibilities for professional growth

- Fully remote work

- Reasonably flexible schedule

- 15 days of paid vacation

- Regular performance reviews