
Site Reliability Engineer III
- GCP
- GKE
- Django
- Airflow
- Cloud SQL
- MySQL
- PostgreSQL
- Redis
- Firestore
- Terraform
- GitHub Actions
- Kubernetes
- Datadog
- Devops
- IAM
- Load Balancing
- Python
- HIPAA
- CI/CD
Vida has been operating and growing for years, and our infrastructure reflects that. We run on GCP with a production GKE cluster hosting around 50 workloads, from Django applications to scheduled Airflow jobs. Our data layer includes Cloud SQL (MySQL and PostgreSQL), Redis, and Firestore. Our infrastructure is defined in two Terraform repositories, one for core GCP infrastructure and one for our data platform, and both have grown across many contributors over time. Until now, our infrastructure has been managed by backend engineers with deep infrastructure experience, and this role adds our first dedicated SRE to that group.
You'll be Vida's first dedicated Site Reliability Engineer. You'll join the Enablement Team, which owns the platform and tooling our Engineering Teams build on. You'll report to the Engineering Manager and work closely with the team's Lead Engineer, who sets technical direction and will mentor you. This is a fully remote role with no time zone restrictions.
You'll modernize, consolidate, and scale our infrastructure as Vida takes on a wave of new enterprise contracts starting January 1. You'll also help shape what SRE looks like at Vida going forward.
Repsonsibilities:
- Consolidate our Terraform, which has grown into inconsistent patterns across our infrastructure and data repositories, into a clean, well-documented structure the whole team can work in. Establish conventions for state management, module structure, code review, and CI checks.
- Normalize environments, improve build and deploy automation in GitHub Actions, and add drift detection and alerting.
- Apply overdue patches and upgrades across our Cloud SQL databases and application runtimes.
- Right-size compute and database workloads for growth, including connection pooling and scaling improvements for high-traffic services.
- Evaluate our Kubernetes architecture as we grow, including whether and when to move to a multi-cluster setup.
- Improve monitoring and observability in Datadog and Cloud Monitoring so we catch issues before they become incidents.
- Design observability access for contractors and external partners that gives them the visibility they need while keeping protected health information out of view.
- Retire legacy infrastructure and tooling that has been replaced but not yet decommissioned.
- Build repeatable operational processes, including runbooks, an on-call rotation, and escalation documentation.
- Support infrastructure readiness for Vida's January 1 enterprise launches.
- Additional responsibilities as needed.
Qualifications:
- Bachelor's degree at a minimum.
- 5+ years of experience in SRE, DevOps, or infrastructure engineering, with real ownership of production systems.
- Deep hands-on Terraform experience, including structuring modules and managing state across environments.
- Strong working knowledge of GCP, including GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking and load balancing, and cost management.
- Production Kubernetes experience, including autoscaling, resource management, and judgment about what belongs in the cluster versus outside it.
- Hands-on experience building monitoring, alerting, and dashboards with tools like Datadog or Cloud Monitoring.
- Proficiency in Python for tooling and automation.
- Comfortable working across multiple teams and disciplines, and explaining infrastructure decisions to non-specialists.
Preferred:
- Experience as an early or first SRE hire.
- Experience refactoring or consolidating a large, organically grown Terraform codebase.
- Experience improving observability from a less mature baseline.
- Experience in a HIPAA-regulated or other compliance-driven environment.
- CI/CD experience with GitHub Actions.
- Experience running Django applications or Airflow in production on Kubernetes.
- Experience designing or migrating to multi-cluster Kubernetes architectures.
Site Reliability Engineer III · Vida Health