logo[мetahunt]
> Djinni

Machine Learning Engineer

format:Hybridtype:Contractcompany:Outstaff
pytorchcudakubernetesdockerawsgoogle cloudany one ofawsterraformtensorrtvllm
experience3+ years
domainAI
locationUnited States (New York, San Francisco)
> full description

Location: New York, NY or San Francisco, CA (Hybrid / On-site)

Employment Type: W-2 Contract or Full-Time (Must be authorized to work on W-2 in the US without sponsorship)

Experience: 3+ Years

 

Job Summary

We are seeking an AI / ML Infrastructure Engineer to join our client's high-performance engineering team. You will design, scale, and optimize the underlying infrastructure powering large-scale generative AI and machine learning models in production.

 

Key Responsibilities

  • Architect and scale high-performance GPU clusters (NVIDIA hardware) for model training and low-latency inference.
  • Optimize model serving pipelines using Triton Inference Server, vLLM, and TensorRT.
  • Collaborate with data scientists and MLOps engineers to streamline distributed training workflows.
  • Monitor, troubleshoot, and optimize cloud infrastructure costs and performance bottlenecks.

 

Requirements

  • Minimum 3 years of hands-on experience in ML infrastructure, systems engineering, or DevOps roles.
  • Strong proficiency in PyTorch, CUDA ecosystem, and container orchestration (Kubernetes, Docker).
  • Deep experience with cloud providers (AWS, GCP) and Infrastructure as Code (Terraform).
  • Valid W-2 work authorization in the United States (no C2C or cross-border arrangements).

    Sound like a fit?
    Apply now and our recruiting team will get back to you.