Birmingham District (b), England
Job Summary
Key responsibilities
Provide technical support by handling and consulting on BAU, Incidents/emails/alerts for the respective applications.
Perform post-mortem, root cause analysis using ITIL standards of Incident Management, Service Request fulfilment, Change Management, Knowledge Management, and Problem Management.
Manage regional L2 team and vendor teams supporting the application. Ensure the team is up to speed and picks up the support duties.
Build up technical subject matter expertise on the applications being supported including business flows, application architecture, and hardware configuration.
Define and track KPIs, SLAs and operational metrics to measure and improve application stability and performance.
Conduct real time monitoring to ensure application SLAs are achieved and maximum application availability (up time) using an array of monitoring tools.
Build and maintain effective and productive relationships with the stakeholders in business, development, infrastructure, and third-party systems / data providers & vendors.
Assist in the process to approve application code releases as well as tasks assigned to support to perform. Keep key stakeholders informed using communication templates.
Approach support with a proactive attitude, desire to seek root cause, in-depth analysis, and strive to reduce inefficiencies and manual efforts.
Mentor and guide junior team members, fostering technical upskill and knowledge sharing.
Provide strategic input into disaster recovery planning, failover strategies and business continuity procedures
Collaborate and deliver on initiatives and install these initiatives to drive stability in the environment.
Perform reviews of all open production items with the development team and push for updates and resolutions to outstanding tasks and reoccurring issues.
Drive service resilience by implementing SRE(site reliability engineering) principles, ensuring proactive monitoring, automation and operational efficiency.
Ensure regulatory and compliance adherence, managing audits, access reviews, and security controls in line with organizational policies.
The candidate will have to work in shifts as part of a Rota covering APAC and EMEA & USA hours providing 24x7 support. In the event of major outages or issues we may ask for flexibility to help provide appropriate cover.
Your skills and experience
7-12 years of experience in providing hands on IT application support.
Experience in managing vendor team’s providing 24x7 support.
Preferred : Team lead role experience, Experience in an investment bank, financial institution.
Bachelor’s degree from an accredited college or university with a concentration in Computer Science or IT-related discipline (or equivalent work experience/diploma/certification).
Preferred : ITIL v3 foundation certification or higher.
Knowledgeable in cloud products like Google Cloud Platform (GCP) and Kubernetes and hybrid applications.
Strong understanding of ITIL /SRE/ DEVOPS best practices for supporting a production environment.
Understanding of production KPIs, SLO, SLA and SLI.
Monitoring Tools: Knowledge of Elastic Search, Control M, Grafana, Geneos, OpenShift, Prometheus, Google Cloud Monitoring, Airflow, Splunk.
Working Knowledge of creation of Reporting Dashboards and reports for senior management
Key Responsibilities
Key responsibilities
Provide technical support by handling and consulting on BAU, Incidents/emails/alerts for the respective applications.
Perform post-mortem, root cause analysis using ITIL standards of Incident Management, Service Request fulfilment, Change Management, Knowledge Management, and Problem Management.
Manage regional L2 team and vendor teams supporting the application. Ensure the team is up to speed and picks up the support duties.
Build up technical subject matter expertise on the applications being supported including business flows, application architecture, and hardware configuration.
Define and track KPIs, SLAs and operational metrics to measure and improve application stability and performance.
Conduct real time monitoring to ensure application SLAs are achieved and maximum application availability (up time) using an array of monitoring tools.
Build and maintain effective and productive relationships with the stakeholders in business, development, infrastructure, and third-party systems / data providers & vendors.
Assist in the process to approve application code releases as well as tasks assigned to support to perform. Keep key stakeholders informed using communication templates.
Approach support with a proactive attitude, desire to seek root cause, in-depth analysis, and strive to reduce inefficiencies and manual efforts.
Mentor and guide junior team members, fostering technical upskill and knowledge sharing.
Provide strategic input into disaster recovery planning, failover strategies and business continuity procedures
Collaborate and deliver on initiatives and install these initiatives to drive stability in the environment.
Perform reviews of all open production items with the development team and push for updates and resolutions to outstanding tasks and reoccurring issues.
Drive service resilience by implementing SRE(site reliability engineering) principles, ensuring proactive monitoring, automation and operational efficiency.
Ensure regulatory and compliance adherence, managing audits, access reviews, and security controls in line with organizational policies.
The candidate will have to work in shifts as part of a Rota covering APAC and EMEA & USA hours providing 24x7 support. In the event of major outages or issues we may ask for flexibility to help provide appropriate cover.
Skill Requirements
Other Requirements
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-