Unknown Company

Site Reliability Engineer – Lead

az • Posted Yesterday
Onsite Full Time IT & Technology

  • Build and lead a team to deliver technology products and services
  • Develop a technology strategy and ensure technology solutions comply with standards
  • Promote design, engineering, and organizational practices
  • Advocate and advance modern, Agile solution delivery practices
  • Define and implement SRE frameworks, including SLIs/SLOs/SLAs, error budgets, and incident response protocols
  • Establish governance models for reliability engineering across distributed teams
  • Champion a culture of observability and proactive monitoring
  • Lead root cause analysis (RCA) and post-incident reviews
  • Implement proactive problem detection using telemetry
  • Develop and maintain capacity models and monitor performance trends
  • Drive automation of operational tasks including deployments and scaling
  • Oversee major incident response and communication processes
  • Serve as a senior technical advisor and thought leader in SRE and platform engineering
  • Mentor SRE teams and partner with engineering leaders across the enterprise

Requirements

  • 10+ years of experience in systems engineering, DevOps, or SRE roles in large-scale environments
  • Deep understanding of Linux/Unix & Windows systems, networking, and distributed computing
  • Proven experience with observability stacks (e.g., Dynatrace, Grafana, Splunk, OpenTelemetry)
  • Expertise in infrastructure-as-code and automation tools (e.g., Terraform, Ansible, Python)
  • Strong knowledge of cloud platforms and container orchestration (Kubernetes)
  • Demonstrated success in leading incident response and driving systemic improvements
  • Experience with capacity planning, performance tuning, and cost optimization
  • Excellent communication and stakeholder management skills, including executive engagement.

Core Competencies

Demonstrates extensive expertise in Site Reliability Engineering (SRE) and platform engineering, with a strong focus on developing technology strategies, implementing observability practices, and leading incident response efforts. Proven ability to mentor teams and drive automation in large-scale environments while ensuring compliance with industry standards.

Highest-signal resume keywords

  • Site Reliability Engineering (SRE)
  • Infrastructure-as-Code
  • Observability Stacks
  • Cloud Platforms
  • Incident Response Leadership

Hard Skills

  • Systems Engineering
  • DevOps
  • Linux/Unix Systems
  • Windows Systems
  • Networking
  • Distributed Computing
  • Capacity Planning
  • Performance Tuning
  • Cost Optimization
  • Automation Tools

Soft Skills

  • Excellent Communication
  • Stakeholder Management
  • Executive Engagement

Industry Keywords

  • Agile Solution Delivery
  • Governance Models
  • Proactive Monitoring
  • Root Cause Analysis
  • Telemetry

Tools & Technologies

  • Terraform
  • Ansible
  • Python
  • Kubernetes
  • Dynatrace
  • Grafana
  • Splunk
  • OpenTelemetry

#J-18808-Ljbffr
Back to Job Search