
Senior Site Reliability Engineer (US)
- Kubernetes
- Incident Response
- Azure
- Node.js
- Change Management
- Helm
- Disaster Recovery
- Devops
- AKS
- EKS
- IaC
- Terraform
- Ansible
- CI/CD
- GitHub Actions
- Datadog
- Prometheus
- Grafana
- Loki
- OpenTelemetry
- Jira
- Confluence
- Octopus Deploy
- SOC2
- Istio
- PostgreSQL
- RabbitMQ
- NATS
- Equity
- Pension
Senior Site Reliability Engineer
Remote | US
About Climavision
At Climavision, we’re rebuilding climate technology from the ground up and changing the way we see weather. We merge the power of our proprietary, high-resolution weather radar and satellite network with advanced weather prediction modelling and decades of industry expertise to reduce existing coverage gaps and drastically improve forecasting ability. Our revolutionary new approach to climate technology weather solutions is poised to help reduce the economic risks of climate change on companies, governments, and societies alike. We are backed by The Rise Fund, the world’s largest global impact platform committed to achieving measurable, positive social and environmental outcomes alongside competitive financial returns. Climavision is headquartered in Louisville, KY, with research and development operations in Raleigh, NC.
The Work
Are you an experienced Site Reliability Engineer who thrives at the intersection of software engineering and production operations? Do you take pride in keeping mission-critical customer systems reliable under real-world operational pressure? Are you looking for an opportunity to own production reliability for a modern hybrid infrastructure platform spanning cloud, colocation, and edge environments?
If so, we have an exceptional opportunity for you.
Climavision is seeking a Senior Site Reliability Engineer to contribute towards reliability, operational excellence, and production resilience across the company's platform and data services. This role sits on a shared SRE team that supports the full business rather than a single product line, covering both the radar network and the weather intelligence sides of the company as priorities shift. A central focus of this role is building the observability layer that puts the health of the full fleet in one place, and then automating recovery so that systems heal themselves. Multi-cluster and multi-replica high availability across our distributed edge fleet remains a core part of the work.
This is a hands-on engineering role for someone who is equally comfortable troubleshooting Kubernetes clusters, leading incident response, and improving operational maturity across the organization. The successful candidate will combine deep production operations expertise with a disciplined approach to reliability engineering and strong automation skills.
Climavision operates a hybrid infrastructure footprint spanning Microsoft Azure, colocation data centers, and edge Kubernetes clusters, deployed alongside weather radar systems. This role will drive production reliability across Azure, colocation, and edge environments. Right-sizing cluster resources and migrating workloads off Azure to reduce spend are active priorities for the team.
35% Kubernetes Platform Reliability and Operations
30% Production Reliability Engineering and Incident Response
20% Observability, Monitoring, and Alerting
15% Automation, Recovery, and Cost Optimization
Primary Responsibilities:
• Own production reliability for Climavision's customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
• Work as part of a shared SRE function supporting the whole company rather than a single product line, taking on work across both the radar network and the weather intelligence sides of the business as priorities shift.
• Contribute to the definition and improvement of SLIs, SLOs, alerting standards, and operational metrics used to measure platform reliability.
• Build and own the observability layer for the fleet. Today the underlying metrics exist but are only reachable from the command line inside each cluster. This role is responsible for surfacing that data in shared dashboards and building the alerting that tells the team something is going wrong before a customer does.
• Design and build automated recovery and self-healing for production systems, so that common failure modes are detected and remediated without human intervention.
• Optimize cluster resourcing and cost, including right-sizing workloads and nodes and supporting the migration of workloads off Azure to reduce spend.
• Support and coordinate production incident response efforts, including troubleshooting, mitigation, communication, and postmortem analysis.
• Diagnose and resolve complex production issues across application services, Kubernetes infrastructure, storage, and distributed systems.
• Drive multi-replica and multi-cluster high availability across Climavision's services, including workload placement, scheduling, and deployment patterns that allow services to run safely as multiple replicas across multiple clusters.
• Contribute to the multi-cluster high-availability strategy across Climavision's hybrid fleet, including active-active and active-passive failover behavior, traffic routing, data replication considerations, and graceful degradation when a cluster becomes unavailable.
• Operate and improve Climavision's self-managed Kubernetes platform spanning cloud-hosted, colocation, and edge clusters, with a focus on availability, resiliency, recovery, and operational performance.
• Ensure Kubernetes platform lifecycle activities including upgrades, patching, cluster health, node management, and production change management are executed in a manner that preserves service availability and minimizes customer-facing risk.
• Improve reliability and operational maturity of production platform services, including observability, autoscaling, ingress, and distributed storage. Partner with the teams responsible for the underlying networking and security primitives rather than owning those areas directly.
• Design and validate Kubernetes workloads for resiliency, scalability, and operational efficiency, including autoscaling behavior, workload placement, resource management, and graceful degradation strategies.
• Partner with software engineering teams across the company to improve production readiness, resiliency patterns, deployment safety, and operational visibility before services reach production.
• Maintain and improve deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation supporting safe and repeatable production releases.
• Support and evolve Climavision's observability platform, including metrics, logging, distributed tracing, dashboarding, and alerting.
• Conduct performance engineering and capacity-planning efforts for customer-facing services during peak weather-event demand.
• Help facilitate blameless postmortem reviews and drive operational follow-up items through completion.
• Improve disaster recovery, failover, and business continuity capabilities across cloud, colocation, and edge environments.
• Drive operational excellence initiatives, including automation, reduction of operational toil, game days, production readiness reviews, and reliability best practices.
• Contribute as a senior technical resource and mentor on reliability engineering and production operations practices.
On-Call Expectation:
Climavision operates customer-facing production systems under contractual SLAs that do not pause outside business hours. The Senior Site Reliability Engineer will participate in a rotating on-call schedule made up of two separate rotations:
• A weekday rotation. The engineer on a weekday shift is the first point of contact for production incidents and for engineering teams needing support during the business week.
• A separate weekend rotation, so that the engineer carrying weekday support is not also carrying the weekend.
At current and planned team size, engineers can expect a weekday shift roughly every five weeks and a weekend shift roughly every five weeks. The two are scheduled as far apart from each other as the rotation allows, so that a weekday shift and a weekend shift do not fall close together.
Qualifications
• A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience considered.
• Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
• Deep, hands-on experience operating native Kubernetes. Managed distributions such as AKS and EKS are acceptable, but experience running native or self-managed Kubernetes is strongly preferred and is the primary technical requirement for this role.
• Demonstrated experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and reducing infrastructure cost without sacrificing reliability. Be prepared to walk through a specific cluster optimization project you led.
• Demonstrated experience increasing operational visibility, including building dashboards, metrics pipelines, and alerting in an environment where little or none existed before.
• Experience designing and operating workloads for safe horizontal scaling across multiple replicas, including idempotency, concurrency, and state handling considerations.
• Experience designing or operating multi-cluster high-availability architectures, including failover behavior, traffic routing, and cross-cluster service deployment.
• Experience supporting customer-facing production systems with uptime, reliability, and incident-response responsibilities.
• Experience diagnosing and resolving production incidents across application, platform and Kubernetes infrastructure layers, including workload scheduling, storage, ingress, and cluster-level failures.
• Experience operating Kubernetes outside of strictly managed cloud environments, including bare-metal, colocation, edge, or hybrid infrastructure.
• Experience with Kubernetes operational tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.
• Strong understanding of infrastructure automation and Infrastructure as Code concepts using tools such as Terraform and Ansible.
• Experience supporting CI/CD and production deployment pipelines. GitHub Actions is used for CI/CD at Climavision.
• Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or comparable technologies.
• Experience operating distributed systems and microservice-based architectures in production environments.
• Working knowledge of Microsoft Azure infrastructure.
• Strong troubleshooting skills across infrastructure, application, and platform layers.
• Demonstrated experience participating in a structured production on-call rotation supporting business-critical systems.
• Working familiarity with Jira, Confluence, and Microsoft Entra, which the team uses day to day for ticketing, documentation, and authentication.
• Strong written and verbal communication skills, including incident documentation and postmortem authoring.
• Experience working in start-up, scale-up, or other fast-moving engineering environments, and comfort with the pace and ambiguity that comes with them.
Nice to have, but not required:
• Experience operating Kubernetes platforms using RKE2 and Rancher, which is how Climavision manages its clusters.
• Experience with Octopus Deploy.
• Experience with SOC 2 or comparable security auditing and compliance work.
• Experience supporting hybrid cloud and colocation infrastructure environments.
• Experience with service mesh technologies such as Istio.
• Experience with Kubernetes-native storage platforms such as Longhorn.
• Experience operating PostgreSQL or PostGIS in Kubernetes environments.
• Experience with distributed messaging systems such as RabbitMQ or NATS.
• Experience supporting GPU-enabled workloads in Kubernetes.
• Familiarity with reliability engineering practices, including SLIs, SLOs, error budgets, and operational maturity metrics.
Physical Demands & Work Environment:
• This is a full-time, exempt position
• Fully Remote - United States
• This job requires frequent use of a computer to complete tasks, attend meetings, and communicate via Microsoft Teams.
Once you land this position, you’ll get to enjoy:
• Benefits of a dynamic and growing organization
• A challenging, hands-on role that will have real impact on the business
• Competitive compensation
• Comprehensive benefits package
• 401(k) Savings Plan
• Medical/Dental/Vision Benefits
• Health Savings Account (HSA) and Flexible Spending Account (FSA)
• Unlimited Paid Time-off
• 11 Paid Holidays
• Paid Parental Leave
• Company Paid Short-term Disability (STD)
• Company Paid Long-term Disability (LTD)
• Company Paid Life Insurance
The salary range for this position is $130,000-170,000 annually, however Climavision considers several factors when extending an offer of employment including but not limited to, the applicant’s education, experience, the responsibilities of the role, training, knowledge, skills, and abilities, as well as internal equity and alignment with market data. Any offer of employment is contingent on completion of a background check to company standard. Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities, and activities may change at any time with or without notice.
Climavision is an equal opportunity employer. All aspects of employment including the decision to hire, promote, discipline, or discharge, will be based on merit, competence, performance, and business needs. We do not discriminate on the basis of race, color, religion, marital status, age, national origin, ancestry, physical or mental disability, medical condition, pregnancy, genetic information, gender, sexual orientation, gender identity or expression, veteran status, or any other status protected under federal, state, or local law.
Senior Site Reliability Engineer (US) · Climavision