DevOps Engineer
Role overview
Svitla Systems Inc. is looking for a GPU & ML Infrastructure Engineer for a full-time position (40 hours per week) in Ukraine. Our client is a technology startup developing a physics-informed foundational model to monitor and predict GPU and compute health. Its technology uses multi-sensor telemetry, physics-based stress signals, semiconductor degradation estimates, and dynamical mathematical features to assess hardware health over time.
The team runs controlled experiments on GPU systems, executing standardized benchmarks and stress workloads while capturing detailed hardware telemetry (temperature, power, clocks, error counters, etc) from NVML, DCGM, BMC, Redfish, and lab-grade equipment. The test procedure is validated on a few GPU platforms. We are extending it across a broader fleet: multiple NVIDIA datacenter GPU generations and embedded/edge platforms, with more targets added over time. You will own the data generation side end-to-end: porting test procedures to new node types, automating deployment, ensuring telemetry quality and consistency across every target, and landing documented datasets in storage. You will own the workloads. The suite includes LLM inference and training workloads (Llama-family models at several sizes, quantized and full precision), extending to larger models as we move to higher-memory GPUs. Repeatability is crucial: identical load every run, so the variation comes from the hardware, not the workload. You are building the system that reliably produces the data the modeling team depends on.
Requirements
- Experience deploying LLM inference and training workloads on GPUs, including quantized models.
- Demonstrated ability to diagnose sensor and sampling problems in time-series hardware data independently.
- Understanding of how to make a GPU workload reproducible run-to-run: controlling sources of nondeterminism.
- Strong understanding of Linux systems, GPU driver stacks, process orchestration, scheduling, and timing behavior.
- Experience collecting hardware telemetry programmatically, in-band (NVML, DCGM) and out-of-band (BMC /IPMI / Redfish).
Nice to have
- Experience building data collection pipelines for hardware testing, qualification, or systems research.
- Familiarity with GPU benchmark and stress tooling (compute, memory bandwidth, multi-GPU communication) and benchmark methodology (warmup, steady-state, run-to-run variance).
- Solid understanding of GPU power and thermal management semantics, such as clock throttle reasons.
- Understanding of multi-GPU scaling for larger models (tensor/pipeline parallelism, NCCL).
- Experience building automated data-quality validation for time-series or sensor data.
Responsibilities
- Port test procedure to new hardware. Datacenter GPUs and edge devices differ in their driver stacks, power, thermal envelopes, and available sensors. Adapt the procedure, document what changed and why.
- Build and maintain the benchmark workload suite: synthetic kernels (GEMM, convolution), LLM inference and training, and new classes as target classes expand, including vision models for the edge. These are controlled load generators for hardware characterization.
- Own logger correctness and consistency. Make collection predictable across platforms and sources: maintain intended sampling rates, ensure trustworthy timestamps, use consistent field names, and align clocks between in-band and out-of-band sources. Diagnose sampling anomalies to root cause.
- Automate deployment. Repeatable, unattended process: stand up workloads and loggers on a new node, execute the run schedule, collect, and tear down cleanly. Runs are reproducible from configuration.
- Build automated data quality checks. Catch missing or irregular samples, clock mismatches between loggers, telemetry that doesn’t align with the workload, and sensors silently returning no data on new hardware.
- Own the storage handoff. Land data in a consistent, documented layout with full run metadata (hardware target, driver and firmware versions, procedure version, schedule).
- Document each hardware target.
Why join Svitla
Svitla Systems is a global digital solutions company headquartered in the U.S. and operating across the Americas, Europe, Asia, and APAC. Since 2003, we have served a wide range of clients — from innovative start-ups to Fortune 500 companies. Our success is built on partnership. By integrating seamlessly with clients’ teams, we create lasting collaborations that drive real results. We are strong advocates of workplace flexibility, remote culture, individual approach to professional and personal growth.
Our global mission is to build a business that contributes to wellbeing of our partners, personnel, and their families, improves our communities, and makes a lasting difference in the world. Together, we are coding a brighter tomorrow — and living it.
Join us!
Відгукнутись на вакансію