
Ready to take your career global?
Make your mark at one of the biggest names in payments. We’re looking for an Associate Site Reliability Engineer to join our ever-evolving Experience & Portals Application Operations team and help shape the future of global commerce.
What you will own
We are seeking a highly motivated Associate SRE Specialist to join the Experience & Portals Application Operations team. In this role, you will be responsible for ensuring the reliability, availability, and performance of production applications and cloud infrastructure. You will collaborate with cross-functional teams to support critical business services, troubleshoot production issues, and drive operational excellence through automation and continuous improvement.
Lead a team of 6-8 Site Reliability Engineers/Production Support Analysts responsible for ensuring the availability, stability, and performance of business-critical applications and infrastructure.
Hands-on technical expertise in leading production support for Google Cloud Platform (GCP), Java-based applications, Cloud SQL, Service Management, and Change Management.
Lead critical incident management activities and coordinate resolution efforts across multiple teams/applications
Drive Root Cause Analysis (RCA) and implement preventive measures to reduce recurring incidents
Define and track operational KPIs, SLAs, SLOs, and service health metrics.
Improve Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR)
Drive automation initiatives to reduce manual operational effort.
Implement proactive monitoring, alerting, and observability solutions.
Identify opportunities for process improvements and operational efficiency.
Support CI/CD and DevOps initiatives to enhance deployment reliability.
Create and maintain operational runbooks, knowledge articles, and standard operating procedures
What you will bring
8 to 10 years of experience in a Production Support, Application Support, SRE or similar IT operations role.
Minimum 2-3 years of experience leading support or SRE teams.
Strong knowledge of Java/J2EE applications.
Experience supporting Spring Boot and Microservices architectures.
Understanding REST APIs and application of troubleshooting techniques.
Experience analyzing application logs and performance bottlenecks
Experience working with IT Service Management (ITSM) tools such as ServiceNow.
Working knowledge of Google Cloud Platform (GCP) and Java-based applications.
Hands-on experience with databases such as Cloud SQL, SQL Server, Big Query, Fire-store and MongoDB.
Hands-on experience on supporting applications in the Payments domain
Knowledge of monitoring and observability tools like Splunk, Grafana, Logic monitor
Strong analytical and troubleshooting skills with the ability to diagnose and resolve production issues effectively.
Strong verbal and written communication skills, including the ability to communicate effectively with both technical and non-technical stakeholders.
Self-motivated, proactive, and accountable, with a strong sense of ownership and commitment to outcomes.
It’s a bonus if you have
Experience with release deployment and release management using Harness.
Knowledge of data analytics and reporting tools such as Power BI.
Automation and scripting experience using Python and Microsoft Power Automate.
About the team
Our inclusive and global teams win together every day. We’re proud to have the best minds in the industry, who you can learn from as you grow your career. The people, the energy, the connections
