We are seeking an experienced Site Reliability Engineering (SRE) Subject Matter Expert to help establish and mature reliability practices within an AWS environment. This individual will create scalable SRE frameworks, define service-level objectives (SLOs), and help engineering teams adopt consistent standards for system availability, performance, monitoring, and incident response.
This is a long-term, 12-month engagement with a high probability of extension. The ideal candidate will combine hands‑on AWS experience with the ability to establish strategy, frameworks, and processes across multiple technical teams.
Key Responsibilities
- Assess the organization’s current reliability practices and identify areas for improvement.
- Develop reusable SRE frameworks, standards, and operating models.
- Partner with application and infrastructure teams to define service-level indicators (SLIs), SLOs, and error budgets.
- Establish consistent monitoring, observability, alerting, and incident-management practices.
- Improve system availability, performance, scalability, and operational resilience.
- Create dashboards and reporting that provide visibility into service health and reliability.
- Help teams reduce manual operational work through automation.
- Facilitate root-cause analysis and translate findings into long-term improvements.
- Develop documentation, playbooks, and governance standards that can be adopted across teams.
- Coach engineering teams on SRE principles and help build sustainable internal capabilities.
Required Qualifications
- Extensive experience in Site Reliability Engineering, DevOps, cloud infrastructure, or a related discipline.
- Demonstrated experience building or maturing an SRE practice.
- Strong understanding of SLIs, SLOs, SLAs, error budgets, and reliability measurement.
- Hands‑on experience supporting applications and infrastructure in AWS.
- Experience with observability, monitoring, alerting, and incident-management tools.
- Strong understanding of infrastructure automation, CI/CD, and cloud operational practices.
- Experience creating technical frameworks and driving adoption across multiple teams.
- Ability to communicate effectively with engineering teams, technical leaders, and business stakeholders.
- Strong documentation, facilitation, and technical leadership skills.
Preferred Qualifications
- Experience working in a large, complex enterprise environment.
- Familiarity with AWS services such as CloudWatch, EC2, ECS/EKS, Lambda, and related cloud-native tooling.
- Experience with infrastructure-as-code tools such as Terraform or CloudFormation.
- Experience with observability platforms such as Datadog, Dynatrace, Splunk, Grafana, or New Relic.
- Previous experience serving as an SRE lead, architect, or enterprise-level subject matter expert.
- Extension potential: High
- Primary focus: SRE framework development, SLO implementation, reliability maturity, and long-term organizational adoption