Responsibilities
- Own the deployment and operation of critical collaboration services across cloud and hybrid environments
- Design, evolve, and optimize CI/CD pipelines and automation
- Lead incident response for complex production issues, perform root cause analysis
- Use observability data to guide capacity planning, scaling strategies, and resource optimization
- Define and champion operational best practices, documentation standards, and a culture of reliability and operational excellence
Requirements
- Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience)
- 5+ years in Site Reliability Engineering, Cloud Operations, or Systems Engineering
- Strong hands‑on experience operating production services using Docker and Kubernetes in cloud or hybrid environments
- Proficiency in one or more programming or scripting languages (e.g., Python, Go, Bash)
- Experience with monitoring, observability, and incident response in production environments
- Working knowledge of Linux systems, networking, distributed systems, CI/CD pipelines, infrastructure‑as‑code, and Git‑based workflows
Core Competencies
Demonstrates expertise in deploying and operating critical collaboration services in cloud and hybrid environments, with a strong focus on CI/CD pipeline optimization and incident response. Proficient in using observability data for capacity planning and resource optimization while championing operational best practices.
#J-18808-Ljbffr