- Provide L2/L3 support for business-critical production applications
- Troubleshoot and resolve application, infrastructure, and platform incidents
- Perform root cause analysis and implement preventative improvements
We're hiring a Site Reliability Engineer (SRE) to join a global engineering team supporting mission-critical financial markets platforms undergoing a large-scale cloud, cyber security, and platform modernisation programme.
This is a hybrid Production Engineering and Cloud Operations role where you'll help maintain the stability of business-critical systems while driving automation, operational improvements, and cloud transformation initiatives. It's an excellent opportunity for engineers who enjoy troubleshooting complex production environments, improving platform reliability, and reducing manual operational effort through automation within a highly regulated, enterprise-scale environment.
Key Responsibilities
- Provide L2/L3 support for business-critical production applications
- Troubleshoot and resolve application, infrastructure, and platform incidents
- Perform root cause analysis and implement preventative improvements
- Build and maintain CI/CD pipelines and deployment automation
- Develop automation solutions using Python and Shell scripting
- Support AWS-based cloud environments and platform modernisation initiatives
- Improve monitoring, observability, and operational processes
- Partner with Engineering, Security, and Operations teams to deliver platform improvements
- Support vulnerability remediation, patching, and security compliance activities
- Contribute to platform reliability, resilience, and continuous improvement initiatives
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Production Engineering, Cloud Operations, or Infrastructure Engineering
- Strong experience supporting production-critical environments
- Hands-on AWS experience
- Strong Linux administration experience
- Python and Shell scripting experience
- Experience with CI/CD pipelines and deployment automation
- Strong incident management and root cause analysis skills
- Experience with monitoring and observability tools such as Datadog, BigPanda, or Splunk
- Excellent communication and stakeholder management skills
Preferred Skills
- Red Hat Enterprise Linux (RHEL)
- Ansible
- Terraform
- HashiCorp Vault
- Artifactory
- Security remediation or cyber security programme experience
- Financial Services experience
- Experience supporting global platforms across multiple regions
EA registration number : ANDREW JONAS MATTHEW, R21103843 Allegis Group Singapore Pte Ltd, Company Reg No. 200909448N, EA Licence No. 10C4544
Please Stay Alert to Potential Scams
We would like to remind you that eFinancialCareers is a job board and does not conduct hiring or ask for payment or any financial details as part of the job application process.
If you receive any suspicious messages claiming to be from us or a hiring company, we urge you not to click on any links and not to reply to the message itself.
Instead, please report the message to our support team at support@efinancialcareers.com.
It is advisable to always verify job offers directly with the hiring company.
We’re partners in transformation. We help clients activate ideas and solutions to take advantage of a new world of opportunity. We are a committed team working with over 6,000 clients across North America, Europe and Asia Pacific.
As an industry leader in talent services, we work with progressive leaders to drive change. That’s the power of true partnership.
TEKsystems is an Allegis Group company.