We have partnered with a high-availability B2B SaaS company delivering mission-critical software to enterprise clients.
As customer adoption increases, they are expanding their SRE function to improve reliability, scalability, and performance across their cloud-native environment. This is a core growth function for the business.
What you'll be doing:
- Defining SLIs, SLOs, and error budgets
- Improving system reliability across AWS environments
- Driving observability improvements
- Automating infrastructure recovery and scaling
- Leading incident response and postmortem processes
- Optimising performance across distributed systems
About you:
- 5+ years in SRE, DevOps or cloud infrastructure
- Strong AWS or GCP experience
- Kubernetes and container orchestration
- Experience in high-availability SaaS environments