Unknown Company

Site Reliability Engineer Lead

plano, tx • Posted 3 days ago
Onsite Full Time IT & Technology

Overview

Seeking a seasoned Site Reliability Engineering (SRE) Leader to drive the reliability, scalability, and performance of critical Infrastructure Automation platforms. This role will lead the design and implementation of SRE practices across a federated technology ecosystem, ensuring operational excellence through automation, observability, and resilient architecture.

The ideal candidate will bring deep expertise in distributed systems, cloud‑native infrastructure, SaaS application support and DevOps/SRE principles, along with strong leadership and collaboration skills to influence cross‑functional engineering and Production management teams and drive continuous improvement in service reliability.

Responsibilities

SRE Strategy & Governance:
  • Define and implement SRE frameworks, including SLIs/SLOs/SLAs, error budgets, and incident response protocols.
  • Establish governance models for reliability engineering across distributed teams.
  • Champion a culture of observability, proactive monitoring, and continuous feedback loops.
Reactive & Proactive Problem Management:
  • Lead root cause analysis (RCA) and post‑incident reviews to identify systemic issues and prevent recurrence.
  • Implement proactive problem detection using telemetry, anomaly detection, and trend analysis.
  • Collaborate with engineering and operations teams to eliminate toil and reduce incident frequency and impact.
Capacity & Performance Management:
  • Develop and maintain capacity models to ensure systems scale efficiently with business demand.
  • Monitor performance trends and lead optimization efforts across infrastructure and applications.
  • Partner with finance and engineering teams to align capacity planning with cost and growth objectives.
Platform Reliability & Automation:
  • Drive automation of operational tasks including deployments, scaling, and recovery.
  • Integrate reliability tooling with CI/CD pipelines, ITSM platforms (ServiceNow), and observability systems.
Incident Management & Operational Excellence:
  • Oversee major incident response, escalation, and communication processes.
  • Develop and maintain runbooks, playbooks, and escalation protocols.
  • Drive continuous improvement through blameless retrospectives and operational reviews.
Technical Leadership:
  • Serve as a senior technical advisor and thought leader in SRE and platform engineering.
  • Mentor and guide SRE teams and partner with engineering leaders across the enterprise.
  • Provide input on staffing, tooling strategy, and budget planning for reliability initiatives.

Managerial Responsibilities

  • Opportunity & Inclusion Champion: Models an inclusive environment for employees and clients, aligned to company Great Place to Work goals.
  • Manager of Process & Data: Demonstrates deep process knowledge, operational excellence and innovation through a focus on simplicity, data based decision making and continuous improvement.
  • Enterprise Advocate & Communicator: Communicates enterprise decisions, purpose, and results, and connects to team strategy, priorities and contributions.
  • Risk Manager: Ensures proper risk discipline, controls and culture are in place to identify, elevate and debate issues.
  • People Manager & Coach: Provides inspection, coaching and feedback to motivate, differentiate and improve performance.
  • Financial Steward: Actively manages expenses and budgets in alignment with objectives, making sound financial decisions.
  • Enterprise Talent Leader: Assesses talent and builds bench strength for roles across the organization.
  • Driver of Business Outcomes: Delivers results by effectively prioritizing, inspecting and appropriately delegating team work.

Required Qualifications

  • 10+ years of experience in systems engineering, DevOps, or SRE roles in large‑scale environments.
  • Deep understanding of Linux/Unix & Windows systems, networking, and distributed computing.
  • Proven experience with observability stacks (Dynatrace, Grafana, Splunk, OpenTelemetry).
  • Expertise in infrastructure‑as‑code and automation tools (Terraform, Ansible, Python).
  • Strong knowledge of cloud platforms and container orchestration (Kubernetes).
  • Demonstrated success in leading incident response and driving systemic improvements.
  • Experience with capacity planning, performance tuning, and cost optimization.
  • Excellent communication and stakeholder management skills, including executive engagement.

Desired Qualifications

  • Experience with ITIL/ITSM processes and integration with platforms like ServiceNow.
  • Familiarity with security and compliance in regulated industries (financial services).
  • Background in performance engineering and infrastructure analytics.
  • Experience developing dashboards and metrics for operational health and reliability.

Skills

  • Influence
  • Risk Management
  • Solution Design
  • Stakeholder Management
  • Technical Strategy Development
  • Analytical Thinking
  • Application Development
  • Collaboration
  • Result Orientation
  • Solution Delivery Process
  • Agile Practices
  • Architecture
  • Automation
  • Data Management
  • DevOps Practices

Shift

1st shift (United States of America)

Hours Per Week

40

#J-18808-Ljbffr
Back to Job Search