DevOps Engineer
Role Overview
We are looking for a strong DevOps engineer to design, build and run the cloud infrastructure for a healthcare data platform that is moving from proof of concept to production. You will own the AWS and Kubernetes environment, the CI/CD pipelines and the operational practices that get the platform to 99.9% availability, 24×7.
There is no existing infrastructure to inherit and no established process. You will design a tight, cost-conscious platform from a defined architecture, build it, and keep it running — making the calls yourself and documenting them, with a hands-on CTO as your reviewer.
Project Description
The client is a US healthcare-data company building a platform for pharmacy benefit managers. The platform ingests historical patient and prescription data, runs it through an ETL pipeline into a unified data model, and serves machine-learning models that produce drug and treatment recommendations. Today the process is manual — data prepared on a desktop, models deployed by hand, no CI/CD. The goal is a production SaaS platform with a first MVP at the end of 2026.
The infrastructure is AWS, with Docker and Kubernetes as the runtime, and a strong preference for open-source tooling over managed proprietary services. The availability target is 99.9%, 24×7. Current usage is proof-of-concept scale, so the design has to reach that target without over-building, and cost must be a design input from day one: an environment that meets the checklist but whose bill doubles in six months is not acceptable.
This is a full-time, remote engagement. Working hours are 05:00–13:00 US Eastern, and development happens on client-provided virtual desktops.
Key Responsibilities
- Design the AWS and Kubernetes infrastructure for the platform from the target architecture, and document it.
- Build the environments — development, test and production — with clear separation and repeatable provisioning.
- Set up CI/CD from scratch for application code, data pipelines and model deployments.
- Define infrastructure as code and keep every environment reproducible from it.
- Design for 99.9% availability, 24×7: monitoring, alerting, backups, recovery and on-call practice.
- Own cost: choose instance types, storage and scaling policies deliberately, and report on spend.
- Put security and access controls in place appropriate to sensitive healthcare data.
- Work with the data engineer to run pipelines and model services reliably on Kubernetes.
- Replace manual deployment steps with automated, auditable ones.
- Propose an approach and defend it rather than waiting for a fully formed requirement.
- Write the runbooks and documentation the team needs to operate the platform without you in the room.
Required Qualifications
- Senior-level DevOps or platform engineering experience, including ownership of a production environment.
- Strong hands-on AWS experience.
- Strong hands-on Docker and Kubernetes experience in production.
- Experience building CI/CD pipelines from scratch.
- Infrastructure as code with a tool such as Terraform, and the discipline to keep environments reproducible.
- Experience designing for and operating to a defined availability target, including monitoring, alerting and incident response.
- Demonstrated cost-conscious design and cost management in a cloud environment.
- Ability to work independently and set process where none exists.
- English at a level that supports direct client meetings without an intermediary.
- Availability during 05:00–13:00 US Eastern.
Preferred Qualifications
- Experience running data pipelines and machine-learning workloads on Kubernetes.
- Experience with an open-source observability stack (Prometheus, Grafana, Loki or similar).
- Security and compliance experience with healthcare or other regulated data.
- Experience with GitOps tooling such as Argo CD or Flux.
- Experience with Python or scripting for automation.
- Experience taking a proof-of-concept environment to production.