Senior Site Reliability Engineer (SRE)
Job Description Required Education • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent work experience Required Experience • 7+ years in Backup Engineering, Infrastructure Engineering, or Site Reliability Engineering • 5+ years designing enterprise backup solutions • 3+ years supporting cyber recovery architectures • Experience implementing SRE principles within enterprise infrastructure environments • Strong understanding of distributed systems and high availability architectures • Cohesity • Dell PowerProtect Data Manager • Dell Data Domain • Dell Cyber Recovery • Rubrik • Commvault • Veritas NetBackup • Veeam • Air-gapped vaults • Immutable backups • Clean Rooms • Isolated Recovery Environments (IRE) • Recovery orchestration • Cyber resilience testing • Ransomware recovery • Recovery validation • Microsoft Azure • AWS • Google Cloud Platform • Cloud-native backup • Cross-region recovery • Hybrid cloud resiliency • VMware • Hyper-V • Kubernetes • OpenShift • Linux • Windows Server • Active Directory • Enterprise storage platforms • Ansible • Terraform • Python • PowerShell • Bash • GitHub • GitHub Actions • CI/CD pipelines • Dynatrace • Grafana • Prometheus • Splunk • ELK Stack • ServiceNow • Zero Trust architecture • NIST Cybersecurity Framework • CIS Controls • Encryption and key management • Identity and Access Management (IAM) • Multi-factor authentication (MFA) • Secure recovery processes Preferred Qualifications • Experience in financial services or another highly regulated industry • Experience supporting GSIB cyber resiliency programs • Knowledge of regulatory expectations from agencies such as the Federal Reserve, OCC, or FFIEC • Experience with chaos engineering and resilience testing • Familiarity with SRE tooling and reliability metrics • Experience implementing AI-assisted operations (AIOps) and predictive analytics • Strong systems thinking and engineering mindset • Excellent troubleshooting and root cause analysis skills • Ability to lead cross-functional technical recovery efforts • Strong communication and executive presentation skills • Proven ability to influence engineering standards and drive operational excellence • Commitment to continuous improvement through automation and reliability engineering Job Description – Project Overview • Seeking a highly technical Senior Site Reliability Engineer (SRE) with deep expertise in enterprise backup engineering, cyber recovery, and platform resiliency • Responsible for engineering highly available, secure, and automated recovery capabilities that protect against operational failures, ransomware, and other cyber threats • Combines traditional SRE principles (automation, observability, reliability engineering, and resilience) with experience designing and operating enterprise backup platforms, immutable storage, air-gapped cyber vaults, isolated recovery environments (IREs), and recovery orchestration • Partners closely with Infrastructure, Cyber Security, Cloud Engineering, Application Development, and Disaster Recovery teams to ensure critical services remain recoverable, resilient, and continuously validated • Engineer and maintain highly available, resilient enterprise platforms using SRE principles • Define and measure Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for backup and recovery services • Develop automation to reduce operational toil and improve reliability • Perform root cause analysis (RCA) and implement permanent corrective actions • Continuously improve platform reliability, scalability, performance, and recoverability • Establish proactive monitoring, alerting, and observability for backup and cyber recovery platforms • Participate in incident response and major incident recovery activities • Design, implement, and administer enterprise backup and recovery solutions across on-premises, cloud, and SaaS platforms • Engineer immutable backup architectures that support ransomware resilience • Design backup strategies for virtual environments, physical servers, databases, Kubernetes/OpenShift, cloud-native workloads, NAS/Object Storage, and enterprise applications • Optimize backup performance, retention, replication, encryption, and recovery objectives • Implement policy-based backup automation and lifecycle management • Ensure compliance with enterprise RPO and RTO requirements • Design and implement enterprise cyber recovery solutions including air-gapped recovery vaults, clean rooms, Isolated Recovery Environments (IRE), and immutable storage architectures • Develop secure recovery workflows following cyberattack scenarios • Engineer automated malware scanning and recovery validation processes • Design and test recovery orchestration for severe-but-plausible cyber events • Support recovery point validation and promotion into production recovery environments • Collaborate with Cyber Security teams on ransomware resilience strategies • Develop Infrastructure as Code (IaC) and Recovery as Code automation • Build automated recovery runbooks using Ansible, Terraform, PowerShell, Python, and GitHub Actions • Automate recovery validation, reporting, and compliance evidence generation • Eliminate manual recovery processes wherever possible • Implement monitoring for backup success rates, replication health, recovery readiness, storage utilization, cyber vault health, and infrastructure dependencies • Build dashboards for executive and operational visibility • Integrate with enterprise observability platforms (Dynatrace, Grafana, Splunk, Prometheus) • Plan and execute cyber recovery exercises, clean room validation, air-gap recovery testing, full isolated recovery environment exercises, Bare Metal Recovery (BMR) testing, and Disaster Recovery testing • Validate application recoverability against defined RTO/RPO objectives • Produce executive reporting on recovery readiness and testing outcomes
Senior Site Reliability Engineer and Backup Engineer in chicago at Unknown Company
This position is listed as full time and hybrid.