Aventra Systems is looking for a Site Reliability Engineer to join our growing team!
Job Purpose
The Site Reliability Engineer will be responsible for improving the reliability, availability, scalability, performance, and operational efficiency of the organisation’s production applications, systems, and supporting infrastructure.
The role combines software engineering, infrastructure engineering, automation, monitoring, incident management, and production operations. The Site Reliability Engineer will work closely with software development and infrastructure teams to ensure systems are designed, deployed, operated, and continuously improved in a reliable and sustainable manner.
The role will also focus on reducing repetitive manual operational work through automation and implementing resilient and self-healing systems.
Key Duties and Responsibilities
- Design, develop, test, implement, and maintain automation for applications, systems, infrastructure, deployment processes, and recurring operational tasks.
- Support and maintain production applications and infrastructure to achieve required levels of availability, reliability, scalability, and performance.
- Investigate, diagnose, and resolve complex and high-priority production incidents involving applications, infrastructure, operating systems, networks, and cloud services.
- Participate in and lead incident response activities, including coordinating technical investigation and service restoration.
- Conduct root cause analysis and facilitate blameless post-incident reviews to identify technical and process improvements.
- Implement corrective and preventive measures following production incidents to reduce recurrence and improve service reliability.
- Work with software development teams throughout the software development lifecycle to ensure applications are designed and released with appropriate reliability, scalability, monitoring, and operational requirements.
- Develop and maintain automated deployment, configuration, infrastructure, and operational processes using scripting and infrastructure automation technologies.
- Analyse historical incidents, system metrics, logs, alerts, capacity information, and usage patterns to identify trends and predict potential reliability or performance issues.
- Develop proactive solutions to identified operational risks before they result in production incidents.
- Design and implement resilient, fault-tolerant, highly available, and self-healing infrastructure and application patterns.
- Develop and improve monitoring, logging, alerting, dashboards, and observability capabilities for production services.
- Define and monitor appropriate service reliability measures, including service-level indicators and service-level objectives where applicable.
- Lead and participate in performance, load, stress, scalability, and reliability testing of applications and infrastructure.
- Analyse performance-test results and production metrics to identify bottlenecks and opportunities for optimisation.
- Undertake capacity analysis and forecasting to identify future application, compute, storage, network, and infrastructure requirements.
- Support and improve continuous integration and continuous deployment processes to enable reliable and repeatable software releases.
- Collaborate with software engineers, infrastructure engineers, security teams, and other technical stakeholders to resolve technical issues and improve production readiness.
- Produce and maintain technical documentation, operational procedures, runbooks, incident records, and system reliability documentation.
- Identify opportunities to reduce manual operational effort and improve engineering productivity through automation and standardisation.
- Contribute to ongoing improvements in system architecture, infrastructure, deployment processes, operational practices, and production reliability.
Essential Experience
Candidates must have a minimum of 5 years of relevant professional experience in Site Reliability Engineering, DevOps Engineering, Production Engineering, Cloud Infrastructure, Systems Engineering, Platform Engineering, or a closely related technical role.
Candidates should demonstrate professional experience in:
- Supporting and troubleshooting business-critical production applications and infrastructure.
- Developing automation and scripts using technologies such as Python, Bash, PowerShell, or equivalent.
- Linux/Unix operating systems and production systems administration.
- Cloud computing platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
- Infrastructure automation and Infrastructure as Code.
- CI/CD pipelines and automated software deployment.
- Containers and container orchestration technologies such as Docker and Kubernetes.
- Production monitoring, logging, alerting, and observability.
- Incident response, root cause analysis, and post-incident reviews.
- Performance analysis, troubleshooting, optimisation, and capacity planning.
- Highly available, scalable, fault-tolerant, and resilient systems.
- Working collaboratively with software development and infrastructure teams.
Essential Qualification
A Master’s degree in Computer Science, Information Technology, Software Engineering, Engineering, or a closely related discipline is required.
The successful candidate must be able to demonstrate the technical knowledge, professional experience, and practical capability required to perform the responsibilities of the position.
Technical Skills
Relevant technical knowledge may include Linux/Unix, Python, Bash, PowerShell, AWS, Azure, Google Cloud Platform, Docker, Kubernetes, Terraform, Ansible, CI/CD technologies, Git, Prometheus, Grafana, Datadog, Splunk, ELK/OpenSearch, monitoring and observability technologies, networking, distributed systems, and infrastructure automation.
Equivalent technologies may be considered where they provide substantially similar functionality.
Role Requirements
The Site Reliability Engineer is expected to exercise independent technical judgement when investigating production issues, designing automation, improving infrastructure, analysing system performance, and implementing reliability improvements.
Job Type: Full-time
Pay: £42,000.00-£56,000.00 per year
Benefits:
Experience:
- CRM software: 1 year (preferred)
- Reliability Engineer: 1 year (preferred)
- Salesforce: 1 year (preferred)
- Veeva’s cloud-based solution: 1 year (preferred)
- Database management: 1 year (preferred)
Work Location: In person