Job Title: Site Reliability Engineer (SRE)
Position Overview
We are seeking a highly skilled Site Reliability Engineer (SRE) to ensure the reliability, scalability, performance, and operational excellence of mission-critical production environments. The ideal candidate will combine strong software engineering principles with infrastructure expertise to build resilient platforms, automate operations, enhance observability, and minimize operational toil.
As an SRE, you will collaborate closely with Platform Engineering, DevOps, Infrastructure, Security, and Development teams to deliver highly available, secure, and scalable services across hybrid cloud and on-premises environments.
Key Responsibilities
Reliability & System Performance
- Ensure high availability, scalability, and performance of production environments.
- Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets .
- Perform proactive capacity planning, performance tuning, and system optimization.
- Lead Root Cause Analysis (RCA) for production incidents and implement preventive solutions.
- Design and develop automation to eliminate manual operational tasks.
- Build reusable Infrastructure as Code (IaC) using Terraform, Ansible, Pulumi , or similar tools.
- Automate deployments, provisioning, patching, configuration management, and operational workflows.
- Improve developer productivity through self-service platform capabilities.
CI/CD & Release Engineering
- Design, build, and optimize CI/CD pipelines.
- Support automated deployments using:
- GitHub Actions
- Jenkins
- Ensure safe, repeatable, and low-risk software releases.
- Implement deployment strategies such as Blue/Green and Canary deployments.
Monitoring & Observability
- Design enterprise-grade monitoring and alerting solutions.
- Build dashboards and telemetry using:
- Grafana
- ELK Stack
- CloudWatch
- Splunk (preferred)
- Improve application and infrastructure observability using logs, metrics, and distributed tracing.
Incident Response & Reliability Engineering
- Participate in 24x7 on-call rotations.
- Lead production incident response and service restoration.
- Drive post-incident reviews and continuous improvement initiatives.
- Reduce Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
- Manage cloud infrastructure across:
- AWS
- Support hybrid environments including:
- VMware
- Nutanix
- Deploy and manage highly available infrastructure services.
- Deploy and manage Kubernetes clusters.
- Support:
- Amazon EKS
- Amazon ECS
- Troubleshoot containerized workloads and optimize cluster performance.