Unknown Company

Site Reliability Engineer (SRE)

chicago, il • Posted 1 weeks ago
Onsite Full Time IT & Technology

We are looking for highly capable Site Reliability Engineers to build, operate, and continuously improve the cloud platform that powers our products. Every SRE serves as an engineering first-responder for infrastructure incidents, restoring service quickly, identifying root causes, and engineering permanent improvements that strengthen the reliability, scalability, and operational excellence of our platform.

You’ll own the operational health of infrastructure spanning cloud platforms, Kubernetes and container orchestration, service meshes, networking, storage, databases, Infrastructure as Code, CI/CD and GitOps pipelines, and the core platform services that power our applications. Through automation, observability, and continuous improvement, you’ll reduce operational risk, eliminate repetitive work, and help engineering teams deliver software with confidence.

As you progress from Site Reliability Engineer to Senior Site Reliability Engineer and Lead Site Reliability Engineer, your responsibilities expand from independently operating production systems to leading complex incident response, mentoring engineers, improving operational practices, and helping elevate the effectiveness of the SRE team.

This is a high-trust, high-impact role for engineers who enjoy solving difficult operational problems, automating away operational toil, and building resilient platforms that enable the business to scale.

Key Responsibilities

Primary: Reliability Engineering and Incident Response

  • Serve as a primary engineering responder for service incidents.
  • Diagnose issues across cloud infrastructure, Kubernetes, networking, storage, databases, CI/CD systems, and platform services.
  • Restore service quickly while balancing immediate mitigation with long-term reliability.
  • Perform root cause analysis and drive corrective actions through completion.
  • Clearly communicate incident status, customer impact, and recovery progress during production events.
  • Participate in post-incident reviews and continuously improve operational practices.
  • Ensure services are observable through metrics, logs, traces, health checks, and actionable alerting.
  • Contribute to the implementation of telemetry across our backend systems to improve observability and operational insight.
  • Define service level indicators (SLIs), service level objectives (SLOs), and error budgets that align platform reliability with business priorities.
  • Communicate complex technical issues clearly to both technical and non-technical stakeholders.
  • Design, deploy, and maintain resilient cloud infrastructure.
  • Build and support Kubernetes clusters and platform services.
  • Develop and maintain Infrastructure as Code using Terraform or similar tools.
  • Engineer improvements that increase platform reliability, scalability, resiliency, and operational efficiency.
  • Strengthen disaster recovery, backup, and business continuity capabilities.
  • Continuously improve monitoring, alerting, logging, and observability.
  • Perform capacity planning and infrastructure forecasting to support future growth and maintain service reliability.
  • Continuously identify and eliminate operational toil through automation, tooling, and platform improvements.

Supporting: Automation & Platform Operations

  • Automate operational processes to reduce manual effort and operational risk.
  • Improve deployment pipelines and GitOps workflows.
  • Build internal tooling that improves engineering productivity.
  • Partner with Product Engineering, Production Engineering, and Security to improve production readiness and platform reliability.
  • Improve deployment reliability, release processes, and the developer experience through automation and platform engineering.

#J-18808-Ljbffr

Site Reliability Engineer (SRE) in chicago at Unknown Company

This position is listed as full time and onsite.

Back to Job Search