logo[мetahunt]
> DOU

DevOps Engineer

Helpware
format:Remotecompany:Product
awsterraformaws ecsaws fargateaws lambdaaws rdspostgresqlredisaws s3cloudwatch
azuregoogle cloudlitellmvllmpythongoclickhousekafka
experience5+ years
domainSaaS
> full description

About the role

Helpware is building a multi-tenant B2B AI platform for customer support operations — a product portfolio spanning Agent Assist, AI-powered QA, document processing, support automation, workforce management, and others. We’re hiring a Sr. PlatformOps Engineer to own the reliability, security, and operability of this infrastructure, and to carry classic DevOps responsibilities: CI/CD, environments, and developer enablement.

This is a platform-ownership role. Your job is to help the development team make architectural designs operationally real, keep them healthy as tenant count grows, and enforce them structurally (in pipelines and shared tooling).

Working hours moving forward:

8/9 AM — 5/6 PM EST
14:00/15:00 — 23:00/24:00 CET

What you’ll do

Platform infrastructure (AWS)

  • Contribute with cloud architectural designs, focusing on Well-Architected Framework pillars: Security, Operational Excellence, Performance Efficiency, Reliability and Cost Optimization.
  • Own the AWS footprint: ECS Fargate services, Lambda and Batch workloads, RDS PostgreSQL, ElastiCache Redis, S3, CloudWatch, networking/IAM, Bedrock, etc.
  • Manage everything as code (Terraform preferred), with reviewable, repeatable environment provisioning for dev/test/staging/prod.
  • Operate our Observability and self-hosted Langfuse stack: deployment, upgrades, backup/restore, and access control (SSO-backed, role-scoped, itself audit-logged).

DevOps & developer enablement

  • Build and maintain CI/CD pipelines, including the platform’s conformance gates: contract tests for the logging standard and the logging conformance linter.
  • Own secrets management, artifact/image pipelines, and deployment strategies (blue/green or rolling) across services.

Monitoring and Support

  • Own our OpenTelemetry pipeline (SDK → Collector → CloudWatch for app/infra; Langfuse for LLM traces), keeping the three-layer separation between infra observability, LLM observability, and the business audit log intact.
  • Maintain per-tenant usage and cost infrastructure, platform-level dashboards, alerting, and SLOs.
  • Keep trace retention and masking aligned with the data governance standard (masked content only in LLM traces, defined retention windows per layer).
  • Support product teams day-to-day: environment issues, deployment questions, performance debugging.

Security & compliance

  • Maintain the security posture underpinning SOC 2 / HIPAA / GDPR commitments: least-privilege IAM, network isolation, encryption at rest and in transit, audit trails.
  • Ensure prompt/completion data, support the PII masking architecture at the infrastructure level.

What we’re looking for

  • 5+ years in platform, SRE, or DevOps roles running production workloads on AWS (and other Cloud Platforms, preferred)
  • Deep hands-on experience with ECS (or EKS), RDS, ElastiCache, S3, CloudWatch, IAM, and VPC networking.
  • Strong infrastructure-as-code practice (Terraform or CDK) and CI/CD ownership (GitHub Actions, GitLab CI, or similar).
  • Experience with Claude Code and AI Spec Driven Design
  • Experience operating stateful open-source systems in production — running ClickHouse, or comparable columnar/analytics stores (Druid, Pinot) or demonstrably transferable database operations depth (Postgres at scale, Elasticsearch, Kafka).
  • Working knowledge of OpenTelemetry: collectors, exporters, sampling, and instrumentation conventions.
  • A security-first habit: you treat IAM policies, secrets, and data boundaries as design work, not afterthoughts.

Nice to have

  • Experience with multiple Cloud providers (AWS, Azure, GCP).
  • Experience self-hosting LLM/GenAI tooling (Langfuse, LiteLLM, vLLM) or operating GenAI workloads (Bedrock, token cost management).
  • Scripting/automation fluency (Python or Go preferred; the platform’s services are Python/FastAPI).
  • Prior work in a compliance-driven environment (SOC 2 audits, HIPAA, GDPR data residency).
  • Experience with multi-tenant SaaS operations: tenant isolation, per-tenant metering, noisy-neighbor management.
  • FinOps experience — AWS cost allocation, tagging strategies, per-tenant cost attribution.
Відгукнутись на вакансію