
Senior Site Reliability Engineer
- AI
- CI/CD
- Incident Management
- Incident Response
- triage
- Disaster Recovery
- AI/ML
- Devops
- AWS
- Kubernetes
- IaC
- Grafana
- Datadog
- New Relic
- Python
- Bash
Role Summary:Â As a Site Reliability Engineer you will be embedded with a cross functional team who has key responsibilities for certain portions of our systems. Over the course of the first year you will gain the valuable context needed to be truly effective and move at speed in the Filevine environment. During your successive years you will be given specific mission critical objectives that help build out and improve our autonomous systems and simultaneously build out your personal brand as an exceptional engineer who has built and maintained amazing systems that can grow to internet scale.
Responsibilities
Observability & Alerting: Design and improve monitoring, logging, tracing, dashboards, and SLI/SLOs for production visibility.
Automation & CI/CD: Build internal tools and delivery pipelines to boost efficiency, eliminate toil, and ensure reliable deployments.
Reliability & System Quality: Drive continuous improvements in system performance, scalability, and security to mitigate customer impact.
Incident Management: Lead production incident response from triage to resolution, turning lessons into durable runbooks and preventatives.
Technical Leadership: Guide major technical initiatives, align cross-team engineering efforts, and manage technical risks.
Mentorship: Elevate SRE team capability through design reviews, paired problem-solving, and incident post-mortems.
Operations & On-Call: Join the on-call rotation while leading capacity planning and disaster recovery readiness.
AI/ML Operationalization: Leverage operational AI/ML tools to forecast capacity risks, detect patterns, and automate system health.
Qualifications
Experience: 8+ years in software/platform engineering or DevOps, including 5+ years dedicated to Site Reliability Engineering.
Infrastructure & Observability: Expertise in cloud platforms (AWS), Kubernetes, IaC, and full-stack observability (tracing, logging, SLI/SLOs).Grafana, Datadog, NewRelic
Automation & Scripting: Proficient in Python, Go, or Bash for building CI/CD pipelines, production tools, and toil-reducing automation.
Incident Leadership: Proven track record in root cause analysis, high-severity incident response, and long-term reliability engineering.
Leadership & Mentorship: Strong communication skills with a history of mentoring engineers, leading cross-functional projects, and setting technical strategy.
Operational AI/ML: Hands-on experience using AI/ML on telemetry data to predict capacity risks, spot anomalies, and optimize system workflows.
Work Location Expectation: Remote or option to be hybrid/in-office in one the following locations: San Francisco, New York, Chicago, or Salt Lake City  Cool Company Benefits:- A dynamic, rapidly growing company, focused on helping organizations thrive - Medical, Dental, & Vision Insurance (for full-time employees)- Competitive & Fair Pay- Maternity & paternity leave (for full-time employees)- Short & long-term disability- Opportunity to learn from a dedicated leadership team- Top-of-the-line company swag Privacy Policy NoticeFilevine will handle your personal information according to what’s outlined in our Privacy Policy. Communication about this opportunity, or any open role at Filevine, willonly come from representatives with email addresses using "filevine.com". Other addresses reaching out arenot affiliated with Filevine and should not be responded to.Â
Senior Site Reliability Engineer · Filevine