About The Role
The role owns the reliability, scalability, and security of distributed infrastructure powering high-traffic production systems serving millions of users globally.
You will work alongside platform architects and software engineering teams to design resilient cloud architectures, automate deployment pipelines, and minimize system downtime.
Key Responsibilities
- Design, build, and maintain cloud infrastructure on AWS and Kubernetes using Terraform and Infrastructure as Code best practices
- Define and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and comprehensive monitoring metrics using Prometheus, Grafana, and Datadog
- Lead incident response and root cause analysis (RCA) for production outages, implementing automated remediation to prevent recurrence
- Optimize cloud resource utilization, compute performance, and infrastructure security posture across all environments
- Build and scale CI/CD deployment pipelines using GitHub Actions or ArgoCD to support rapid, safe software releases
What We Are Looking For
- 3–6 years of experience in Site Reliability Engineering, DevOps, or systems engineering in cloud-native environments
- Strong hands-on experience with Kubernetes, Docker, and container orchestration at scale
- Proficiency in infrastructure automation and configuration management using Terraform and Ansible
- Deep understanding of Linux systems internals, networking fundamentals (TCP/IP, DNS, TLS), and load balancing
- Bonus: Experience with service meshes like Istio, chaos engineering practices, and software development proficiency in Go or Python
Site Reliability Engineer in minneapolis at Unknown Company
This position is listed as full time and onsite.