Job DescriptionJob Descriptionn
nSr. Site Reliability Engineer
nnnWe are seeking an experienced Sr. Site Reliability Engineer to join a small, high-impact infrastructure team. This role blends software engineering and systems automation to scale reliable cloud and hybrid systems. You will own critical projects from design through deployment, specifically driving a pivotal infrastructure modernization and hosting migration initiative in your first year.
nnResponsibilities:
n- n
- Infrastructure & Migration: Lead modernization efforts and hosting migrations while maintaining hybrid infrastructure (Azure, AWS, on-prem). n
- Automation &IaC: Streamline provisioning using Infrastructure as Code (Terraform, Ansible, PowerShell DSC) and enhance CI/CD pipelines(GitHub Actions, Jenkins)for rapid delivery. n
- Observability: Introduce and manage comprehensive monitoring platforms (e.g. Prometheus, Grafana, Datadog)to establish operational standards. n
- Reliability Engineering: Manage containerized work loads in Docker and implement SRE principles including SLOs and error budgets. n
- Operational Excellence: Lead incident response, conduct root cause analysis, and automate manual processes through scripting and runbooks. n
- Security & Continuity: Support DevSecOps initiatives and ensure robust backup and disaster recovery strategies. n
- Emerging Tech & R&D: Continuously evaluate and pilot emerging technologies, tools, and industry trends to ensure our infrastructure stack remains modern, efficient, and scalable. n
Qualifications:
n- n
- Minimum 5 years in SRE, DevOps, or Cloud Engineering with production experience in Azure or AWS. n
- Strong Linux/Unix administration and networking troubleshooting skills. n
- Expertise in Infrastructure as Code (Terraform) and CI/CD pipeline design. n
- Proficiency in scripting or programming (Python, Go, Bash, or PowerShell). n
- Expertise with Docker in production environments; Kubernetes experience is a strong plus. n
- Self-starter with strong communication and documentation skills; able to take ownership in a small-team environment. n
- Excellent oral and written communication skills. Able to communicate effectively with a diverse group of individuals with varying levels of technical understanding and varying skillsets. n
- Ability to work collaboratively with application development and data engineering areas to define standards and manage change. n
- Must reside in the Greater Cincinnati Metropolitan Area (Hybrid Schedule 3 days a week in office) n
Qualifications:
n- n
- Practical knowledge of observability tools (Prometheus, Grafana, ELK, or similar). n
- Windows Server/Active Directory administration. n
- Experience with legacy Unix (AIX/Solaris). n
- Database (Oracle/MS-SQL) or BI platform experience (Snowflake/Azure Fabric). n
- Relevant industry certifications (Azure, AWS, or CKA). n
Hybrid Schedule three days a week in office required
Sr. Site Reliability Engineer in Cincinnati in cincinnati at Unknown Company
This position is listed as full time and hybrid.