Site Reliability Engineers (SRE)
The Role: You will join a team working with Observability, Escalations, Post-mortems, Correction of Errors, and other practices that will c…
Job Title: Site Reliability Engineer -Cloud Kubernetes Platform
Location: Leeds/Bristol, UK-hybrid
Hybrid/Remote: 3 days /week onsite
Mode of employment: Inside IR 35 Contract
We are looking for a passionate and experienced Site Reliability Engineer (SRE) to join our Cloud Platform team. The ideal candidate will have hands-on experience managing large-scale Kubernetes clusters on public cloud environments (AKS, EKS, or GKE) and a strong understanding of modern SRE and DevOps practices. You will be responsible for ensuring high availability, reliability, scalability, and performance of our cloud-native infrastructure and CI/CD systems.
• Manage, monitor, and optimize large-scale Kubernetes clusters hosted on public cloud platforms (Azure AKS, AWS EKS, or Google GKE).
• Implement and maintain infrastructure as code using tools such as Terraform.
• Collaborate with development and operations teams to improve system reliability and deployment automation.
• Build and maintain CI/CD pipelines using Jenkins or similar tools.
• Troubleshoot production issues, conduct root cause analysis, and implement preventive measures.
• Automate operational tasks using Python or other scripting languages.
• Contribute to observability and monitoring improvements using modern tools and best practices.
• Participate in on-call rotations and incident response processes.
• 5–9 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles.
• Strong hands-on experience managing Kubernetes clusters in production (AKS/EKS/GKE).
• Proficiency with Terraform and cloud infrastructure automation.
• Practical experience with Jenkins and CI/CD pipeline management.
• Sound understanding of SRE principles (incident management, blameless postmortems, capacity planning, error budgets, etc.).
• Good programming or scripting skills in Python (preferred) or similar languages.
• Strong analytical, troubleshooting, and problem-solving abilities.
• Excellent written and verbal communication skills.
• Experience with Prometheus, Grafana, or OpenTelemetry for observability.
• Exposure to GitOps practices and tools (e.g. Flux).
Neutral 2–4 sentence summary of what working at this company is like, drawn from public reviews and press coverage. Tone, collaboration style, pace, benefits highlights.
The Role: You will join a team working with Observability, Escalations, Post-mortems, Correction of Errors, and other practices that will c…
Are you looking for a new challenge? Fancy helping us shape the future of motor insurance? Prima could be the place for you. Since 2015, we…
eleQtron develops and operates full-stack quantum computers based on trapped-ion technology. Quantum computing is a new computing paradigm r…
IntroductionAt a glance Location & work model: Berlin, hybridTech stack: Kubernetes on our own servers, Harvester (KubeVirt), Argo CD/Flux,…
<h2 class="s-pb-2 s-pt-4 s-heading-xl s-text-foreground dark:s-text-foreground-night"><strong class="s-font-semibold s-text-foreground dark:…
<div class="content-intro"><p>At NiCE, we don’t limit our challenges. We challenge our limits. Always. We’re ambitious. We’re game changers.…