Unknown Company

Director, Site Reliability Engineering

california, mo • Posted 2 days ago
Onsite Full Time Manufacturing & Production

  • Define and execute the long-term Site Reliability Engineering strategy and roadmap
  • Establish the SRE operating model, team scope, engagement models, ownership boundaries, and success measures
  • Build and develop a high-performing team of site reliability and operations engineers
  • Modernize SRE through automation, AI-assisted operations, self-service capabilities, and engineering-first practices
  • Translate business priorities and customer impact into reliability investments and engineering outcomes
  • Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs
  • Establish service-level indicators, service-level objectives, error budgets, and reliability standards
  • Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems
  • Define production operational and observability readiness, including readiness reviews and certification practices
  • Drive improvements in availability, performance, resiliency, and recovery
  • Define enterprise observability strategy across metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry
  • Establish instrumentation, telemetry, dashboards, alerting, and service-health standards
  • Improve incident detection, response, mitigation, communication, learning, and on-call practices
  • Lead automated detection, diagnosis, remediation, and incident creation
  • Establish blameless post-incident reviews and eliminate recurring operational toil
  • Develop intelligent-operations roadmaps covering anomaly detection, event correlation, automated triage, root-cause analysis, and remediation
  • Evaluate AI-assisted workflows across observability, incident response, capacity planning, and operational support
  • Build safe, measurable, auditable automation with appropriate human oversight
  • Partner with engineering, DevOps, and business stakeholders to build shared accountability for production outcomes
  • Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field
  • 10+ years of progressive engineering experience
  • 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teams
  • Proven experience building or transforming a reliability or operational engineering organization
  • Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery
  • Experience establishing observability, incident-management, service-level objective, and production-readiness practices
  • Demonstrated ability to improve reliability through engineering and automation
  • Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems
  • Strong understanding of metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring
  • Demonstrated success driving alignment across cross-functional engineering teams and executive stakeholders
  • Ability to balance immediate operational needs with long-term engineering transformation
  • Strong written, verbal, and executive communication skills
  • Preferred: Master's degree or MBA
  • Preferred: Experience operating large-scale systems in AWS or another major cloud environment
  • Preferred: Experience with New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry
  • Preferred: Experience implementing OpenTelemetry or common instrumentation standards
  • Preferred: Experience building internal developer platforms, paved roads, or self-service reliability capabilities
  • Preferred: Experience applying AI, machine learning, or agent-based automation to operational workflows
  • Preferred: Experience with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering
  • Preferred: Software engineering experience and ability to engage in architecture and design discussions
  • Preferred: Experience supporting high-profile launches, events, or systems with significant customer and business impact

Core Competencies

Demonstrates expertise in Site Reliability Engineering, focusing on automation, observability, and operational excellence. Proven ability to lead high-performing teams and drive engineering transformations that enhance system reliability and performance.

#J-18808-Ljbffr
Back to Job Search