The Enterprise Technology Services organization partners with every part of the American Express business to power the company’s growth and innovation with trust and efficiency, and drive competitive differentiation with speed. We support the delivery and operations of technology, digital, and data capabilities, platforms, and services globally. Specifically, our team is responsible for the company’s technology engineering, architecture, and infrastructure, providing 24x7 support to ensure an uninterrupted, high-quality experience for customers and colleagues. We also provide product management for core enterprise platforms, and lead technology risk and information security, enterprise data governance and platforms, digital product and design, and enterprise AI platforms on behalf of the company.
Manager, Site Reliability Engineering leads and mentors Site Reliability Engineering (SRE) teams, fostering a culture of continuous improvement and inclusivity, while collaborating across the organization to enhance system resilience, scalability, and alignment with business objectives.
- Bachelor’s degree in computer science, Information Technology, Engineering, and/or comparable experience; advance degree preferred
- Knowledge of modern observability stack – Splunk, Elastic Search, Prometheus, Grafana
- Knowledge of containerization technologies (e.g., Kubernetes, Docker) and microservices architecture
- Knowledge of observability tools and methodologies, including experience with logging, monitoring, tracing, and performance analysis platforms
- Knowledge of cloud-based Site Reliability Engineering (SRE) practices and experience with public cloud platforms such as AWS, Azure, or Google Cloud.
Knowledge of Jira, confluence, rally and project management tools including MS office suit.
-
Work Experience:
- Must have Domain knowledge of Cards Payments systems. Understanding of E2E workflows of Authorization Approval and clearing & reconciliation processes.
- Have SRE experience with knowledge of SRE functions
- Have knowledge on Splunk, ELF/Kibana and Prometheus/Grafana and experience to use these tools to automate and configure alerting and build dashboards.
- Have experience in application support (Must have).
- Application support in cloud-based environment.
- Conceptual/support knowledge of microservices in cloud environment and deployment process
- Incident management system knowledge (service now)
- Good communication skills to run production bridges.
- Experience in source control using tools such as Git with DevOps and IT automation concepts.
- Basic UNIX knowledge and any programing language (preferably java, go lang or UI stack).
- Some knowledge in Redhat Open Shift 3.9/3.11 or Kubernetes 1.9/1.11 and above.
- Perform day today support activities to track incidents, respond timely on incidents and review and analyze issues at level 2.
- Create automation dashboards using Splunk, Grafana and Kibana
- Flexible to work shifts (only day shift, start may be little late than usual time). And ready to provide weekend support as per roster.
- Review current issues and work with Engineering team to get code fixed and deployed.
- Familiar with Agile or other rapid application development methods
- Experience with design and coding across one or more platforms and languages as appropriate
- Experience with distributed (multi-tiered) systems, algorithms, and relational databases
- A proactive approach to spotting problems, areas for improvement, and performance bottlenecks.
- Good Knowledge of Networking & Services like TCP/UDP, HTTPs, rest APIs.
- Able to understand and use complex data structures and associated components
- Designs, codes, tests, maintains, and documents applications
- Takes part in reviews of own work and reviews of colleagues' work
Defines test conditions based on the requirements and specifications provided
-
Depending on factors such as business unit requirements, the nature of the position, cost and applicable laws, American Express may provide visa sponsorship for certain positions.