SVP, Site Reliability Engineering Lead, SRE & Governance, Group Technology

DBS BankSingaporeOn-siteFull-timePrincipal, 12–15+ yearsListed 1 week ago

Apply now

About this role

Role Summary

The SVP, Site Reliability Engineering (SRE), will lead and oversee the 24/7 infrastructure operations and reliability engineering function across critical platforms including Hypervisors (VPC, EPC, OPC), OpenShift , Windows, Databases, TWS, and Mainframe environments .

This role is responsible for driving resilience, scalability, automation, and operational excellence across hybrid cloud and on-premises environments, while ensuring alignment with business, risk, and regulatory expectations.

Key Responsibilities

Leadership & Governance

- Lead and manage a distributed 24/7 SRE infrastructure team , including shift-based operations and command center functions
- Define and execute the SRE strategy aligned to enterprise technology and business priorities
- Establish strong governance across incident, problem, change, release, and capacity management
- Drive SLA/SLO/SLI frameworks to ensure service reliability and performance targets

Infrastructure & Platform Ownership

- Oversee end-to-end reliability of infrastructure platforms: Cloud & Container : VPC, OpenShift, Kubernetes
- Compute & Virtualization : Hypervisors (VMware/others), private cloud platforms
- Enterprise Platforms : Windows, Unix/Linux, TWS, Mainframe, Databases

- Ensure high availability, resilience, and disaster recovery readiness across all critical systems
- Own infrastructure lifecycle including capacity planning, patching, upgrades, and decommissioning

Reliability Engineering & Automation

- Champion SRE principles including error budgets, toil reduction, and automation-first mindset
- Drive end-to-end observability strategy (monitoring, logging, tracing)
- Lead initiatives to reduce MTTR, incident volume, and manual operational effort
- Scale automation across deployment, patching, incident resolution, and self-healing capabilities

Operational Excellence

- Ensure 24/7 monitoring, incident response, and recovery processes are robust and continuously improved
- Lead major incident management and command bridge coordination for critical outages
- Conduct RCA, trend analysis, and preventive engineering improvements
- Embed ITIL best practices across service management processes

Risk, Compliance & Security

- Identify infrastructure risks and drive proactive mitigation strategies
- Ensure compliance with regulatory, audit, and internal security requirements
- Partner with security teams on hardening, vulnerability management, and access controls

Stakeholder & Cross-Functional Collaboration

- Collaborate with application, DevOps, security, architecture, and business teams to improve system reliability
- Provide leadership in large-scale transformation programs (cloud adoption, infra modernization, SRE maturity)
- Act as a key interface with senior management and external stakeholders

People & Talent Development

- Build and develop a high-performing SRE organization across L1/L2/L3 layers
- Drive fungibility, cross-skilling, and leadership development within the team
- Mentor senior leaders and establish clear career progression frameworks

Requirements

Experience

- 18+ years of experience in IT infrastructure, SRE, or production operations
- Proven leadership in managing large-scale 24/7 infrastructure teams in banking/financial services
- Strong experience in hybrid cloud, data center, and enterprise platforms

Technical Expertise

- Deep expertise in: Cloud platforms (private/public cloud architectures)
- Container platforms (OpenShift/Kubernetes)
- Hypervisors & virtualization technologies
- Operating systems (Windows, Linux/Unix)
- Databases (MariaDB, Postgres, MSSQL, Redis, DB2)
- Enterprise scheduling & legacy systems (TWS, Mainframe)

- Strong understanding of DevOps, CI/CD, and infrastructure as code

Leadership & Functional Skills

- Strong strategic thinking with ability to translate business goals into technology outcomes
- Excellent incident leadership and crisis management skills
- Proven track record of driving automation and operational transformation
- Strong stakeholder management and executive communication skills

Other Skills

- Expertise in ITIL / Service Management frameworks
- Strong analytical, problem-solving, and decision-making capabilities
- Ability to manage high-pressure situations and multiple priorities

Key Success Metrics (Optional for your slide/JD refinement)

- Infrastructure availability (SLA/SLO adherence)
- Reduction in MTTR / incident volume
- Automation coverage & reduction in manual toil
- Capacity utilization and cost optimization
- Audit and compliance adherence

Location:
DBS Asia Hub

Job:
Technology

Schedule:
Regular

Employee Status:
Full time