logo[мetahunt]
> Djinni
senior

DevOps Engineer

company:Outsource
prometheusopentelemetrygrafanaalertmanagerkubernetespythongo
mcpai agents
experience8+ years
domainCloud
> full description

About Our Client
 

Our client builds and operates GPU cloud infrastructure at the scale modern AI demands. They run dense accelerated-compute clusters on bare metal and Kubernetes, stitched together with high-performance Ethernet and InfiniBand fabrics, and deliver them to customers as reliable, secure, high-throughput platforms. They are building Cloud v2, their next-generation, clean-slate, Kubernetes-native GPU cloud, and scaling it across new data center sites. The infrastructure team sits at the center of that effort, building the systems that turn racks of hardware into a product customers depend on.
 

The role

We are hiring a Senior Observability Engineer to build the observability layer for Cloud v2. This role exists to close the single most damaging gap in our client’s current product: today their observability is reactive, so customers are their failure detector. Some of their worst incidents went undetected internally because they occurred below the layer they instrument (a corrupted forwarding entry, a flapping firewall sync process), and the customer reported the outage before their dashboards did. Cloud v2 is architected from day one against the industry’s Gold-tier bar for neocloud platforms (SemiAnalysis ClusterMAX), and proactive detection, where they see problems before the customer does, is the trust gap they are building this platform to fix.
 

You own the metrics, logs, alerting, dashboards, and detection stack that make that true. That means operator-facing observability that catches failures first and drives automated remediation, and tenant-facing observability that gives customers the usage, health, and status visibility a hyperscaler-grade cloud implies. This is the observability workstream the platform validates earliest, because detection is where the current product loses trust, and it maps directly to the ClusterMAX Monitoring and Reliability criteria.
 

This is a build-focused, hands-on senior IC role. You will spend your time engineering the telemetry pipelines, detection logic, dashboards, and alerting that turn raw signals from multi-vendor GPUs (AMD/NVIDIA), hosts, fabrics, and switches into problems we catch before customers feel them, and you will set the technical direction and standards for how we instrument the platform. You will partner closely with our client’s SRE, platform, network, and security engineers, who own the control-plane architecture, the fabric, and day-to-day operations; you own the observability that makes the whole platform legible and its failures visible.
 

What you’ll own

The OpenTelemetry foundation: own OpenTelemetry as the instrumentation standard the whole platform plugs into: the collector topology and pipelines, semantic conventions for metrics, logs, and traces, and consistent resource and tenant attribution across GPU, host, fabric, switch, and control-plane signals. Everything we instrument, internal and tenant-facing, lands on this common backbone so signals are portable, correlated, and not locked to any one backend.

Proactive detection: build the detection that fires before a customer reports, GPU failure-rate tracking, Xid detection, and the anomaly and threshold signals that feed the automated detect-drain-remediate loop, so mean time to detect stops being inverted and our client sees the problem first.

The metrics pipeline: own the metrics stack across GPU, CPU, network planes, thermal, and power (Prometheus and OpenTelemetry-based), including the DCGM profiling-metric suite (DCGM_FI_PROF_* for SM activity and tensor TFLOPs, plus PCIe AER, ECC, power and thermal, and NVLink and fabric throughput), with custom exporters and OpenTelemetry receivers where DCGM falls short.

Instrumenting below the layer we watch today: push instrumentation down into the sub-layers where our most damaging silent failures live (forwarding state, firewall and HA sync, fabric health), so the failure classes that used to reach us through the customer surface on our own dashboards instead.

Host and switch debug data: build reliable extraction of host and switch debug and telemetry data so responders have the signal they need without hand-collecting it during an incident.

Operator and tenant dashboards and alerting: build the operator dashboards and low-noise, actionable alerting the on-call team runs on, and a tenant-facing status page so customers are not our failure detector.

Tenant-scoped observability: deliver per-tenant usage and utilization visibility and per-tenant access to logs and metrics, cleanly isolated from internal ops telemetry, on a multi-tenant model where no tenant can see another’s data.

Proactive customer notifications: build automated, proactive maintenance and change notifications to affected tenants, replacing today’s manual and ad hoc comms.

Reliability and SLO instrumentation: build the telemetry behind the 99.9%+ node SLA measurement, the passive and active health-check battery, and MTTA and MTTR tracking, so reliability is measured, not asserted, and so the health automation has trustworthy signal to act on.

AI-leveraged observability: push our observability toward an AI-augmented model, agentic triage, anomaly detection, alert-noise reduction, and MCP-driven access to our operational telemetry, so a small team catches more, faster, than headcount alone would allow.

Technical direction and standards: set the instrumentation patterns, metric and label conventions, alerting standards, and SLO practice the rest of the platform adopts, and raise the bar through design review and mentorship.
 

What we’re looking for

Required

8+ years building production observability or monitoring systems at scale, ideally for cloud, hosting, or high-performance compute environments, with real platform-build experience (not solely operating dashboards someone else built).

Deep expertise with Prometheus and the metrics ecosystem: PromQL, exporters, recording and alerting rules, Alertmanager, and Grafana, used to build durable, reusable observability rather than one-off dashboards.

Hands-on experience designing and running OpenTelemetry as a foundational instrumentation layer: the Collector, receivers and exporters, semantic conventions, and vendor-neutral pipelines that multiple backends and teams plug into.

Experience running time-series at scale, including long-term and high-cardinality storage (Thanos, Mimir, Cortex, or VictoriaMetrics), and a real understanding of cardinality, retention, and query-performance tradeoffs.

A track record of building alerting that is actionable and low-noise, and of practicing SLO and error-budget-based reliability rather than alert on everything.

Strong programming skills in a real language (Python, Go, or similar) used to build maintainable exporters, collectors, and observability tooling, not just glue scripts.

Strong Linux systems fundamentals: you can reason about the OS, networking, storage, and the boot and provisioning path, and instrument at that level, not just at the application surface.

A track record of owning significant observability systems end to end and setting technical direction that other engineers adopt: defining standards and influencing architecture across a team.

A demonstrated bias toward building things right the first time: reproducible, code-defined observability and sound instrumentation design over short-term shortcuts.

Strong written communication: you document designs and decisions clearly, and you raise the team’s bar through design review and mentorship.
 

Strongly preferred

Multi-vendor GPU/HPC observability: DCGM & dcgm-exporter (NVIDIA) and RDC & AMD Device Metrics Exporter (AMD/ROCm); NVML and AMD SMI; Xid vs. amdgpu/RAS error reporting; DCGM_FI_PROF_* and ROCProfiler profiling metrics; GPU node health checks (DCGM diagnostics, AGFHC, ROCm Validation Suite).

Experience building multi-tenant, tenant-facing observability with per-tenant metric and log isolation.

Checkmk (our client runs parent and child sites today), and experience migrating or bridging a Checkmk estate into a modern metrics stack.

Network and fabric observability: switch telemetry, streaming telemetry (gNMI) and SNMP, BGP and fabric health, and RoCE Ethernet or InfiniBand/UFM fabric monitoring.

Log and trace pipelines and distributed tracing: Loki, Elastic, or Vector, and instrumenting traces across a distributed control plane.

Kubernetes observability: kube-state-metrics, cAdvisor, and GPU Operator metrics, in a multi-cluster environment.

Status-page and incident tooling, MTTA and MTTR instrumentation, and integration with an incident-response and paging workflow.

Building observability within a SOC 2, ISO 27001, or HIPAA compliance posture, including audit-log export.
 

How we work

Small, senior, high-trust teams. We hire people who can own and solve problems end to end and we give them the room to do it.

Code over clicks. The platform is defined, reviewed, and versioned. Manual changes are the exception and they get automated away.

Build it right. We invest in sound architecture and durable design so the platform scales cleanly instead of accumulating debt.

Customers are real. The platform exists to serve customer workloads, and we build to the reliability and self-service standard a hyperscaler-grade cloud implies.

We document. Designs, decisions, and architecture are written down so the team compounds knowledge instead of re-learning it.
 

AI as a force multiplier

We expect our engineers to use AI to move faster and operate at higher leverage, and we build our environment to make that the default rather than an afterthought. We maintain CLAUDE.md and AGENTS.md context files across our repositories, run MCP servers against our operational systems (Jira, NetBox, Checkmk), and use agentic coding tools such as Claude Code directly in our infrastructure and platform workflows.
 

For this role in particular, observability is one of the highest-leverage places to apply AI: anomaly detection, automated incident triage, alert-noise reduction, and AI-assisted root-cause analysis over the telemetry you build. You don’t need to be an AI expert, but you should be eager to fold these tools into how you design and operate observability, and to help us define what an AI-leveraged infrastructure team can do.

What success looks like
 

First 30 days

Ramped on the Cloud v2 architecture, the current observability stack (Prometheus, Checkmk parent and child sites, and the GPU and fabric telemetry), and the lab, and able to navigate the repos and design docs independently.

Shipped your first meaningful contribution to the metrics, alerting, or detection stack with the team’s review.

Aligned with SRE, platform, network, and security on where your work meets theirs and on the highest-value detection gaps to close first.
 

First 90 days

Owning a defined area of the v2 observability layer end to end, from design through delivery (for example the OpenTelemetry instrumentation foundation, the GPU and DCGM metrics pipeline, the alerting and SLO framework, or the tenant-facing dashboards and status page).

Delivered proactive detection for at least one failure class that the platform previously learned about from customers, so our client now detects it first.

Measurably reduced alert noise or improved mean time to detect for a meaningful part of the fleet.
 

First 12 months

The go-to engineer for the hardest observability and detection problems in Cloud v2, trusted across SRE, platform, network, and security, and setting the standards the team instruments by.

Designed and delivered a significant piece of the platform’s observability (proactive detection, tenant-facing observability, or reliability and SLO instrumentation) that raised the whole team’s leverage and moved the platform toward its Gold-grade Monitoring and Reliability bar.