An exciting global technology organisation is looking for a Site Reliability Engineer (SRE) to join its growing engineering team. This is a fully remote position, offering the opportunity to work on large-scale, business-critical platforms used by customers around the world.
The role would suit an experienced Site Reliability, DevOps, Platform or Cloud Engineer with strong hands-on experience across Kubernetes and Microsoft Azure who enjoys solving complex production problems, improving reliability and automating manual processes.
You will work closely with Development and DevOps teams, helping to design, build, operate and scale highly available production environments while ensuring services remain reliable, secure and performant. Responsibilities:
Supporting and improving highly available, business-critical production services
Monitoring production environments to ensure availability, scalability, performance and security
Responding to production incidents, diagnosing complex technical issues and restoring services
Carrying out root cause analysis and contributing to post-incident reviews
Building automation to reduce manual and repetitive operational tasks
Using Infrastructure as Code to improve the consistency and scalability of environments
Working closely with Development and DevOps teams throughout the application release process
Balancing the delivery of new functionality with platform reliability and service-level objectives
Identifying bottlenecks and proposing improvements across infrastructure and applications
Improving the reliability, quality and time-to-market of software solutions
Supporting the continued development and scaling of a global product platform
Participating in an out-of-hours on-call rota
Strong hands-on Kubernetes experience
Strong hands-on Microsoft Azure experience
A solid background within Site Reliability Engineering, DevOps, Platform Engineering or a similar environment
Terraform and Infrastructure as Code experience
Strong production troubleshooting and incident response experience
Experience supporting highly available production environments
A strong understanding of reliability, scalability and automation
Experience with scripting or programming, ideally PowerShell or Python or Go
Knowledge of database technologies such as MySQL
A security-first approach to designing and operating infrastructure
Neutral 2–4 sentence summary of what working at this company is like, drawn from public reviews and press coverage. Tone, collaboration style, pace, benefits highlights.