Join the Site Reliability Engineering (SRE) team to support high-engagement multimodal applications. You will be responsible for managing infrastructure, observability solutions, and platform automation while ensuring the reliability and security of hosted applications.
RESPONSIBILITIES
- Automate infrastructure and application rollouts using IaC (Infrastructure as Code)
- Manage and maintain Kubernetes-based hosting environments.
- Implement observability solutions and handle security-related infrastructure incidents.
- Support the transition toward "agentic" and AI-driven infrastructure solutions.
SKILLS
- 3–5 years of professional DevOps experience.
- High proficiency in Terraform and Ansible for environment automation, Infrastructure as Code (IaC)
- Strong hands-on experience with AWS services (EKS, DocumentDB, and Object Storage).
- Expert knowledge of Kubernetes and Docker for application deployment.
- Practical understanding of GitOps workflows and integration.
NICE TO HAVE
- Basic understanding of security groups and application exposure management.
- Familiarity with GCP or Azure is a plus.