Skip to sign up

Job matched to your search

D

SVP, Site Reliability Engineering Lead, SRE & Governance, Group Technology

DBS Bank · Singapore

Singapore · HybridFull-TimePosted Sep 9, 2026

Free · Join 5,000+ job seekers using Qarera

How well do you match this role?

Tap the skills you already have — then see your real match score, what’s missing, and your resume fixed for this job.

↑ tap the skills you have
Loading sign-in…
Free · no credit card · 30 seconds

Job description

Role Summary

The SVP, Site Reliability Engineering (SRE), will lead and oversee the 24/7 infrastructure operations and reliability engineering function across critical platforms including Hypervisors (VPC, EPC, OPC), OpenShift , Windows, Databases, TWS, and Mainframe environments.

This role is responsible for driving resilience, scalability, automation, and operational excellence across hybrid cloud and on-premises environments, while ensuring alignment with business, risk, and regulatory expectations.

Key Responsibilities

Leadership & Governance

  • Lead and manage a distributed 24/7 SRE infrastructure team, including shift-based operations and command center functions

  • Define and execute the SRE strategy aligned to enterprise technology and business priorities

  • Establish strong governance across incident, problem, change, release, and capacity management

  • Drive SLA/SLO/SLI frameworks to ensure service reliability and performance targets

Infrastructure & Platform Ownership

  • Oversee end-to-end reliability of infrastructure platforms:

  • Cloud & Container: VPC, OpenShift, Kubernetes

  • Compute & Virtualization: Hypervisors (VMware/others), private cloud platforms

  • Enterprise Platforms: Windows, Unix/Linux, TWS, Mainframe, Databases

  • Ensure high availability, resilience, and disaster recovery readiness across all critical systems

  • Own infrastructure lifecycle including capacity planning, patching, upgrades, and decommissioning

Reliability Engineering & Automation

  • Champion SRE principles including error budgets, toil reduction, and automation-first mindset

  • Drive end-to-end observability strategy (monitoring, logging, tracing)

  • Lead initiatives to reduce MTTR, incident volume, and manual operational effort

  • Scale automation across deployment, patching, incident resolution, and self-healing capabilities

Operational Excellence

  • Ensure 24/7 monitoring, incident response, and recovery processes are robust and continuously improved

  • Lead major incident management and command bridge coordination for critical outages

  • Conduct RCA, trend analysis, and preventive engineering improvements

  • Embed ITIL best practices across service management processes

Risk, Compliance & Security

  • Identify infrastructure risks and drive proactive mitigation strategies

  • Ensure compliance with regulatory, audit, and internal security requirements

  • Partner with security teams on hardening, vulnerability management, and access controls

Stakeholder & Cross-Functional Collaboration

  • Collaborate with application, DevOps, security, architecture, and business teams to improve system reliability

  • Provide leadership in large-scale transformation programs (cloud adoption, infra modernization, SRE maturity)

  • Act as a key interface with senior management and external stakeholders

People & Talent Development

  • Build and develop a high-performing SRE organization across L1/L2/L3 layers

  • Drive fungibility, cross-skilling, and leadership development within the team

  • Mentor senior leaders and establish clear career progression frameworks

Requirements

Experience

  • 18+ years of experience in IT infrastructure, SRE, or production operations

  • Proven leadership in managing large-scale 24/7 infrastructure teams in banking/financial services

  • Strong experience in hybrid cloud, data center, and enterprise platforms

Technical Expertise

  • Deep expertise in:

  • Cloud platforms (private/public cloud architectures)

  • Container platforms (OpenShift/Kubernetes)

  • Hypervisors & virtualization technologies

  • Operating systems (Windows, Linux/Unix)

  • Databases (MariaDB, Postgres, MSSQL, Redis, DB2)

  • Enterprise scheduling & legacy systems (TWS, Mainframe)

  • Strong understanding of DevOps, CI/CD, and infrastructure as code

Leadership & Functional Skills

  • Strong strategic thinking with ability to translate business goals into technology outcomes

  • Excellent incident leadership and crisis management skills

  • Proven track record of driving automation and operational transformation

  • Strong stakeholder management and executive communication skills

Other Skills

  • Expertise in ITIL / Service Management frameworks

  • Strong analytical, problem-solving, and decision-making capabilities

  • Ability to manage high-pressure situations and multiple priorities

Key Success Metrics (Optional for your slide/JD refinement)

  • Infrastructure availability (SLA/SLO adherence)

  • Reduction in MTTR / incident volume

  • Automation coverage & reduction in manual toil

  • Capacity utilization and cost optimization

  • Audit and compliance adherence

Location: DBS Asia Hub

Job: Technology

Schedule: Regular

Employee Status: Full time

Don’t just read the job — see if you’ll get it.

Get your match score, a resume tailored to this exact role, and jobs like it — free.

Check my fit for this job
Loading sign-in…
Apply →

Hiring for a role like this? Join the employer waitlist.