SRE Leader - GPS & YAVA Platform Reliability
- π¨π¦ Canada | πΊπΈ United States
- On-site
- 1 day ago
- Salesforce CRM
- Incident Management
- DNS
- F5
- DB2
- Change Management
- Network Security
- Devops
- Load Balancing
- TLS
- CRM
- Salesforce
- AI
- Five9
- Azure
- GCP
SRE Leader - GPS & YAVA Platform Reliability
Ld Director, Site Reliability Engineering
|
PRIMARY OUTCOME |
PLATFORM SCOPE |
ROLE TYPE |
Role Summary
The SRE Leader will build and lead the Site Reliability Engineering function for the GPS Salesforce CRM and YAVA -Self Service platforms. The primary mandate is to reduce P1/P2 incidents by establishing the people, practices, governance, and engineering capabilities required to identify risk early, control change-related failures, improve detection, and prevent recurring incidents.
What this role owns
The reliability operating model and the SRE team. This leader does not replace application, infrastructure, security, telephony, middleware, cloud, or vendor owners. The role creates shared standards, visibility, accountability, and preventive controls across those teams.
Core Responsibilities
1. Build the SRE Function
-
Define the SRE vision, charter, operating model, engagement model, and implementation roadmap.
-
Recruit, onboard, manage, coach, and retain SRE engineers with the right mix of software, infrastructure, observability, and production engineering expertise.
-
Define SRE competencies, career paths, team objectives, coverage expectations, and engineering standards.
-
Establish clear boundaries between SRE, application development, production support, infrastructure operations, incident management, and vendor teams.
2. Establish Reliability Standards and Governance
-
Define SLIs, SLOs, error budgets, and reliability policies for critical GPS and YAVA journeys.
-
Prioritize incident-oriented indicators such as GPS availability, member-search success, CTI integration success, YAVA transaction success, dependency error rates, and p95/p99 latency.
-
Establish production-readiness reviews and minimum requirements for monitoring, runbooks, capacity, rollback, and ownership.
-
Create a regular reliability review forum and an executive scorecard focused on material risks, incident trends, and corrective-action progress.
3. Lead the Incident Reduction Program
-
Analyze P1/P2 incidents across applications and shared dependencies to identify systemic failure patterns.
-
Create and maintain a prioritized reliability roadmap and backlog based on severity, recurrence, customer or agent impact, and preventability.
-
Ensure every P1/P2 has a defensible root cause, contributing factors, preventive actions, owners, and due dates.
-
Escalate overdue or ineffective preventive actions and drive elimination of repeat failure modes.
4. Strengthen Change and Upgrade Controls
-
Establish risk-based review and validation requirements for changes affecting GPS, YAVA, telephony, DNS, F5, Imperva, MQ, mainframe/DB2, API gateways, cloud, operating systems, databases, and vendors.
-
Require impact assessments, regression evidence, pre-change checks, post-change verification, and tested rollback for high-risk changes.
-
Partner with Change Management to identify changes requiring SRE review or sign-off.
-
Track change failure rate and use incident evidence to improve change controls without creating unnecessary bureaucracy.
5. Drive Observability and Preventive Engineering
-
Set the strategy for end-to-end telemetry, synthetic monitoring, dependency health, alert quality, and service-health correlation.
-
Ensure monitoring reflects actual GPS agent and YAVA customer experiences, not only component status.
-
Sponsor automation for configuration drift, certificates, DNS, routing, load balancer membership, queue health, and post-change validation.
-
Drive investment in resilience patterns, capacity testing, failover validation, and removal of single points of failure.
6. Lead Cross-Team Reliability Accountability
-
Establish working agreements with application, network, security, telephony, middleware, database, mainframe, cloud, and vendor teams.
-
Maintain clear dependency ownership, support contacts, escalation paths, and vendor communication expectations.
-
Use data to surface unresolved risks and secure decisions or investment from senior leadership.
-
Represent platform reliability in major change, readiness, incident, and operational risk forums.
Success Measures
|
Measure |
Expected Direction |
|
P1/P2 incidents |
Sustained reduction in count and business impact |
|
Change-induced incidents |
Lower percentage and severity of failures caused by changes |
|
Repeat incidents |
Reduction by dependency and failure category |
|
Early detection |
Higher percentage detected before customer or agent impact |
|
MTTD and MTTR |
Improvement in detection and restoration time |
|
Preventive actions |
Higher on-time completion and demonstrated effectiveness |
|
Change validation |
Higher coverage for in-scope high-risk changes |
|
SLO performance |
Improved attainment and disciplined error-budget use |
Required Qualifications
-
Demonstrated experience building or materially scaling an SRE function from the ground up.
-
Proven experience directly managing, coaching, and developing SRE engineers.
-
Leadership experience across SRE, production engineering, platform engineering, DevOps, or large-scale technology operations.
-
Strong command of SLIs, SLOs, error budgets, observability, incident management, resilience engineering, and production-readiness practices.
-
Experience driving reliability across organizational boundaries and influencing teams not in the direct reporting line.
-
Strong understanding of distributed systems, networking, DNS, load balancing, TLS, API gateways, cloud platforms, middleware, databases, and vendor dependencies.
-
Executive communication skills with the ability to translate technical risk into customer, agent, operational, and financial impact.
Preferred Experience
-
Enterprise CRM and Salesforce ecosystems.
-
Contact-center technology, CTI, telephony, real-time voice, or conversational AI platforms.
-
IBM MQ, DB2, mainframe integrations, Imperva, F5, Five9, Azure, or Google Cloud.
-
Regulated healthcare or another high-availability enterprise environment.
SRE Leader - GPS & YAVA Platform Reliability Β· Info Way Solutions LLC