- Build and lead a team to deliver technology products and services
- Develop a technology strategy and ensure technology solutions comply with standards
- Promote design, engineering, and organizational practices
- Advocate and advance modern, Agile solution delivery practices
- Define and implement SRE frameworks, including SLIs/SLOs/SLAs, error budgets, and incident response protocols
- Establish governance models for reliability engineering across distributed teams
- Champion a culture of observability and proactive monitoring
- Lead root cause analysis (RCA) and post-incident reviews
- Implement proactive problem detection using telemetry
- Develop and maintain capacity models and monitor performance trends
- Drive automation of operational tasks including deployments and scaling
- Oversee major incident response and communication processes
- Serve as a senior technical advisor and thought leader in SRE and platform engineering
- Mentor SRE teams and partner with engineering leaders across the enterprise
Requirements
- 10+ years of experience in systems engineering, DevOps, or SRE roles in large-scale environments
- Deep understanding of Linux/Unix & Windows systems, networking, and distributed computing
- Proven experience with observability stacks (e.g., Dynatrace, Grafana, Splunk, OpenTelemetry)
- Expertise in infrastructure-as-code and automation tools (e.g., Terraform, Ansible, Python)
- Strong knowledge of cloud platforms and container orchestration (Kubernetes)
- Demonstrated success in leading incident response and driving systemic improvements
- Experience with capacity planning, performance tuning, and cost optimization
- Excellent communication and stakeholder management skills, including executive engagement.
Core Competencies
Demonstrates extensive expertise in Site Reliability Engineering (SRE) and platform engineering, with a strong focus on developing technology strategies, implementing observability practices, and leading incident response efforts. Proven ability to mentor teams and drive automation in large-scale environments while ensuring compliance with industry standards.
Highest-signal resume keywords
- Site Reliability Engineering (SRE)
- Infrastructure-as-Code
- Observability Stacks
- Cloud Platforms
- Incident Response Leadership
Hard Skills
- Systems Engineering
- DevOps
- Linux/Unix Systems
- Windows Systems
- Networking
- Distributed Computing
- Capacity Planning
- Performance Tuning
- Cost Optimization
- Automation Tools
Soft Skills
- Excellent Communication
- Stakeholder Management
- Executive Engagement
Industry Keywords
- Agile Solution Delivery
- Governance Models
- Proactive Monitoring
- Root Cause Analysis
- Telemetry
Tools & Technologies
- Terraform
- Ansible
- Python
- Kubernetes
- Dynatrace
- Grafana
- Splunk
- OpenTelemetry