About The Role
The role owns the reliability, scalability, and performance of mission-critical production systems that serve millions of users daily.
The team works closely with software engineering squads to architect resilient infrastructure, automate operational toil, and maintain high availability standards across distributed cloud environments.
Key Responsibilities
- Design, build, and maintain production infrastructure using Terraform and Kubernetes on AWS or GCP
- Implement comprehensive monitoring, logging, and alerting systems using Prometheus, Grafana, and Datadog to ensure rapid incident response
- Automate deployment pipelines and CI/CD workflows using GitHub Actions and ArgoCD to enable fast, reliable software releases
- Conduct root cause analysis for production incidents, establish blameless post-mortems, and drive systemic architectural improvements
- Optimize cloud infrastructure costs, resource utilization, and network performance without compromising system reliability
- Participate in an on-call rotation to respond to production alerts and maintain service level objectives (SLOs)
What We Are Looking For
- 3–6 years of experience in site reliability engineering, DevOps, or systems engineering within high-scale production environments
- Strong proficiency in infrastructure-as-code tools such as Terraform, CloudFormation, or Pulumi
- Hands-on expertise with Kubernetes container orchestration, Docker, and Linux kernel administration
- Solid programming skills in Python, Go, or Bash for automation and tooling development
- Deep understanding of distributed systems, networking protocols, and cloud architecture principles
- Bonus: Experience with service mesh technologies (Istio, Linkerd), chaos engineering practices, and achieving SOC2 or ISO compliance