SVP, Site Reliability Engineering Lead, SRE & Governance, Group Technology
- VPC
- OPC
- OpenShift
- Windows
- Kubernetes
- VMware
- Unix
- Linux
- Disaster Recovery
- Incident Response
- Incident Management
- ITIL
- Vulnerability Management
- Devops
- MariaDB
- PostgreSQL
- Redis
- DB2
- CI/CD
- IaC
Role Summary
The SVP, Site Reliability Engineering (SRE), will lead and oversee the24/7 infrastructure operations and reliability engineering function across critical platforms includingHypervisors (VPC, EPC, OPC), OpenShift , Windows, Databases, TWS, and Mainframe environments.
This role is responsible for drivingresilience, scalability, automation, and operational excellence across hybrid cloud and on-premises environments, while ensuring alignment with business, risk, and regulatory expectations.
Key Responsibilities
Leadership & Governance
- Lead and manage adistributed 24/7 SRE infrastructure team, including shift-based operations and command center functions
- Define and execute theSRE strategy aligned to enterprise technology and business priorities
- Establish strong governance acrossincident, problem, change, release, and capacity management
- DriveSLA/SLO/SLI frameworks to ensure service reliability and performance targets
Infrastructure & Platform Ownership
- Oversee end-to-end reliability of infrastructure platforms:
- Cloud & Container: VPC, OpenShift, Kubernetes
- Compute & Virtualization: Hypervisors (VMware/others), private cloud platforms
- Enterprise Platforms: Windows, Unix/Linux, TWS, Mainframe, Databases
- Ensure high availability, resilience, and disaster recovery readiness across all critical systems
- Own infrastructure lifecycle includingcapacity planning, patching, upgrades, and decommissioning
Reliability Engineering & Automation
- ChampionSRE principles including error budgets, toil reduction, and automation-first mindset
- Driveend-to-end observability strategy (monitoring, logging, tracing)
- Lead initiatives to reduceMTTR, incident volume, and manual operational effort
- Scale automation acrossdeployment, patching, incident resolution, and self-healing capabilities
Operational Excellence
- Ensure24/7 monitoring, incident response, and recovery processes are robust and continuously improved
- Lead major incident management andcommand bridge coordination for critical outages
- ConductRCA, trend analysis, and preventive engineering improvements
- EmbedITIL best practices across service management processes
Risk, Compliance & Security
- Identify infrastructure risks and driveproactive mitigation strategies
- Ensure compliance withregulatory, audit, and internal security requirements
- Partner with security teams onhardening, vulnerability management, and access controls
Stakeholder & Cross-Functional Collaboration
- Collaborate withapplication, DevOps, security, architecture, and business teams to improve system reliability
- Provide leadership inlarge-scale transformation programs (cloud adoption, infra modernization, SRE maturity)
- Act as a key interface withsenior management and external stakeholders
People & Talent Development
- Build and develop ahigh-performing SRE organization across L1/L2/L3 layers
- Drivefungibility, cross-skilling, and leadership development within the team
- Mentor senior leaders and establish clearcareer progression frameworks
Requirements
Experience
- 18+ years of experience in IT infrastructure, SRE, or production operations
- Proven leadership in managinglarge-scale 24/7 infrastructure teams in banking/financial services
- Strong experience inhybrid cloud, data center, and enterprise platforms
Technical Expertise
- Deep expertise in:
- Cloud platforms (private/public cloud architectures)
- Container platforms (OpenShift/Kubernetes)
- Hypervisors & virtualization technologies
- Operating systems (Windows, Linux/Unix)
- Databases(MariaDB, Postgres, MSSQL, Redis, DB2)
- Enterprise scheduling & legacy systems (TWS, Mainframe)
- Strong understanding ofDevOps, CI/CD, and infrastructure as code
Leadership & Functional Skills
- Strong strategic thinking with ability totranslate business goals into technology outcomes
- Excellentincident leadership and crisis management skills
- Proven track record of drivingautomation and operational transformation
- Strong stakeholder management and executive communication skills
Other Skills
- Expertise inITIL / Service Management frameworks
- Stronganalytical, problem-solving, and decision-making capabilities
- Ability to managehigh-pressure situations and multiple priorities
Key Success Metrics (Optional for your slide/JD refinement)
- Infrastructure availability (SLA/SLO adherence)
- Reduction inMTTR / incident volume
- Automation coverage & reduction in manual toil
- Capacity utilization and cost optimization
- Audit and compliance adherence
Location:
DBS Asia HubJob:
TechnologySchedule:
RegularEmployee Status:
Full timeSVP, Site Reliability Engineering Lead, SRE & Governance, Group Technology Β· Dbs