
Ready to take your career global?
Make your mark at one of the biggest names in payments. We’re looking for an Associate Site Reliability Engineer to join our ever-evolving Experience & Portals Application Operations team and help shape the future of global commerce.
What you will own
We are seeking a highly motivated Associate Site Reliability Engineer (SRE) to join the Experience & Portals Application Operations team. In this role, you will be responsible for ensuring the reliability, availability, and performance of production applications and cloud infrastructure. You will collaborate with cross-functional teams to support critical business services, troubleshoot production issues, and drive operational excellence through automation and continuous improvement.
Monitor applications and services in the production environment to ensure optimal performance and availability.
Monitor and support infrastructure hosted on Google Cloud Platform (GCP).
Perform database monitoring and troubleshoot production issues using application, service, and cloud logs.
Act as a liaison between customers, operations teams, and development teams by applying a software engineering mindset to operational challenges.
Manage incidents, service requests, and changes in accordance with established IT service management processes.
Participate in root cause analysis and implement corrective and preventive actions for recurring issues.
Balance operational support and on-call responsibilities with the development of systems, tools, and automation that improve reliability, scalability, and performance.
Contribute to continuous service improvement initiatives and operational excellence programs.
Participate in a 24x5 support model, including rotational weekend on-call support.
What you will bring
1 to 3 years of experience in a Production Support Engineer, Application Support, or similar IT operations role.
Experience working with IT Service Management (ITSM) tools such as ServiceNow.
Working knowledge of Google Cloud Platform (GCP) and Java-based applications.
Hands-on experience with databases such as Cloud SQL, SQL Server, Big query, Firestore and MongoDB.
Knowledge of monitoring and observability tools like Splunk, Grafana, Logic monitor
Strong analytical and troubleshooting skills with the ability to diagnose and resolve production issues effectively.
Strong verbal and written communication skills, including the ability to communicate effectively with both technical and non-technical stakeholders.
Self-motivated, proactive, and accountable, with a strong sense of ownership and commitment to outcomes.
It’s a bonus if you have
Experience with release deployment and release management using Harness.
Knowledge of data analytics and reporting tools such as Power BI.
Automation and scripting experience using Python and Microsoft Power Automate.
Understanding Site Reliability Engineering (SRE) principles, observability, monitoring, and automation practices.
About the team
Our inclusive and global teams win together every day. We’re proud to have the best minds in the industry, who you can learn from as you grow your career. The people, the energy, the connections
