Job matched to your search
Associate, Site Reliability Engineer (Platform Support), SRE and Governance, Group Technology
DBS Bank · Singapore
Free · Join 5,000+ job seekers using Qarera
How well do you match this role?
Tap the skills you already have — then see your real match score, what’s missing, and your resume fixed for this job.
Job description
Responsibilities:
-
Administrate the full Kubernetes platform life cycle to ensure the platform remains secure, reliable and highly available.
-
Work with other engineers to support the infrastructure of OpenShift as well as assist to resolve infrastructure queries/issues of OpenShift tenant applications.
-
Automation of platform operations and management.
-
Configure and Manage Monitoring, alerts for the OpenShift environments.
-
Manage capacity for Kubernetes environments.
-
Hands-on experience with Docker containerization, Kubernetes orchestration and tools such as Jira and Jenkins.
-
Hands-on experience in handling Linux OS such as Patching, Filesystem management, User management, SELINUX etc.
-
Administration and support Windows Servers and Linux server environments.
-
Perform server hardening, configuration, patching, upgrades, and maintenance.
-
Troubleshoot IIS, SSL certificate, OpenSSH, tectia and IBM CD file transfer
-
Monitor server health, CPU, memory, disk, services, and system availability.
-
Troubleshoot OS, application, network-connectivity, and performance issues.
-
Handle Linux services, packages, file permissions, SSH, cron jobs, and system logs.
-
Respond to incidents, alerts, service requests, and production changes.
-
Perform vulnerability remediation and OS hardening.
-
Maintain operational documentation, SOPs, and troubleshooting guides.
Requirements:
- 3 years of work experience with a bachelor’s degree in computer science or related
Preferred Qualifications. At least 1 year’s hands-on experience with containers in Production Environments - Docker, OpenShift, Kubernetes, Linux and Window Servers preferred
-
Provide operational support for OpenShift Container Platform (OCP), Red Hat Linux, Windows Server, and infrastructure platforms across production and non-production environments.
-
Deliver 24x7 operational support for infrastructure and platform services, ensuring timely incident resolution and service restoration.
-
Perform system administration, monitoring, troubleshooting, performance tuning, and capacity management for Linux, Windows, and containerized environments.
-
Participate in incident management, problem management, root cause analysis (RCA), and post-incident reviews to improve operational stability.
-
Maintain operational documentation, standard operating procedures (SOPs), knowledge articles, and support runbooks.
-
Support security and compliance requirements by executing vulnerability remediation, patch management, access control, and platform hardening activities.
-
Experience with configuration management tools (Chef, Ansible, terraform etc.).
-
Hardening, securing the Kubernetes cluster with monitoring and auditing dashboards
-
Knowledge in infrastructure technologies such as HP and DELL hardware (Blades and Rack servers)
-
Excellent verbal, written, skills; in particular, demonstrated ability to effectively communicate technical and business issues and solutions to multiple organizational levels internally and externally.
-
Candidate must have demonstrated and be prepared to exhibit initiative and ownership of consistent delivery success
-
Be scheduled On-Call to support the infrastructure and systems
Location: DBS Asia Hub
Job: Technology
Schedule: Regular
Employee Status: Full time
More jobs in Singapore
Browse related jobs
Don’t just read the job — see if you’ll get it.
Get your match score, a resume tailored to this exact role, and jobs like it — free.
Check my fit for this job