We're looking for a Senior Site Reliability Engineer to join our engineering team and play a key role in maintaining the reliability, performance and scalability of our cloud infrastructure. This is a hands-on position focused on building resilient platforms, improving operational efficiency through automation and ensuring production services remain secure, stable and highly available.
Working closely with Software Engineers, Platform Engineers and Security teams, you'll help shape the technical direction of our infrastructure, drive operational excellence and implement engineering practices that support reliable software delivery.
Key Responsibilities
- Design, implement and maintain highly available, scalable cloud infrastructure supporting business-critical applications.
- Develop and manage Infrastructure as Code (IaC) using Terraform, CloudFormation or equivalent technologies.
- Build, optimise and maintain CI/CD pipelines to enable secure, reliable and efficient software deployments.
- Improve platform reliability through automation, proactive monitoring and continuous optimisation.
- Implement and maintain observability solutions, including monitoring, logging, alerting and distributed tracing across production environments.
- Define, monitor and improve Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets.
- Lead technical investigations into production incidents, perform root cause analysis and implement permanent corrective actions.
- Work alongside development teams to improve application reliability, deployment strategies and production readiness.
- Optimise Kubernetes clusters and containerised workloads to maximise performance, resilience and resource efficiency.
- Strengthen platform security by implementing infrastructure best practices, identity management and secure deployment processes.
- Identify opportunities to reduce operational complexity through automation and engineering improvements.
- Produce technical documentation, operational runbooks and engineering standards to support platform consistency.
- Mentor engineers, share technical knowledge and contribute to the continuous development of Site Reliability Engineering practices across the organisation.
Requirements
- Proven commercial experience as a Senior Site Reliability Engineer, Platform Engineer, DevOps Engineer or Infrastructure Engineer.
- Strong experience managing production environments within AWS, Microsoft Azure or Google Cloud Platform.
- Excellent knowledge of Linux systems administration and production infrastructure.
- Hands-on experience with Kubernetes, Docker and container orchestration technologies.
- Strong expertise in Infrastructure as Code using Terraform, CloudFormation or similar tools.
- Experience designing and maintaining CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins or Azure DevOps.
- Proficiency in Python, Go or Bash for automation, tooling and operational scripting.
- Practical experience implementing monitoring, logging and distributed tracing using technologies such as Prometheus, Grafana, OpenTelemetry, Datadog, ELK or Splunk.
- Strong understanding of networking, distributed systems, cloud security and high-availability architectures.
- Experience managing production incidents, conducting root cause analysis and driving service reliability improvements.
- Excellent analytical, communication and stakeholder management skills with the ability to work effectively across multidisciplinary engineering teams.
The following would be advantageous:
- Experience with service mesh technologies, including Istio or Linkerd.
- Knowledge of PostgreSQL, MySQL, MongoDB, Redis or other production database technologies.
- Experience with configuration management tools such as Ansible or Puppet.
- Professional certifications in AWS, Azure, Google Cloud or Kubernetes.
What We Offer
- Competitive salary and comprehensive employee benefits package.
- Company pension scheme.
- Private medical insurance.
- Life assurance.
- Generous annual leave entitlement.
- Professional development and technical training opportunities.
- Support for relevant certifications and continuous learning.
- Opportunities to lead strategic engineering initiatives and influence platform architecture.
- Flexible working arrangements, subject to business requirements.
- Employee wellbeing initiatives and recognition programmes.
- A collaborative, inclusive and supportive engineering culture.
Apply
If you're an experienced Site Reliability Engineer who enjoys solving complex technical challenges, building reliable cloud platforms and driving operational excellence through automation, we'd like to hear from you.
Pay: £94,000.00-£98,000.00 per year
Benefits:
- Bereavement leave
- Canteen
- Cycle to work scheme
- Free parking
- Life insurance
- Sick pay
Work Location: In person