Responsibilities
- Ensure reliability and performance of Plaud.ai’s AI products at scale
- Design and operate highly available, scalable cloud‑native systems for AI workloads
- Own production reliability, incident response, and on‑call practices
- Build observability (metrics, logs, tracing) and reliability automation
- Define and manage SLOs, SLIs, and error budgets with engineering teams
- Drive postmortems and reliability improvements across the platform
- Lead incident response and continuous reliability improvement
- Partner with product and engineering teams on reliability design
- Improve observability and operational maturity
Requirements
- 8+ years in SRE, Infra, or Platform Engineering roles
- Strong experience with cloud platforms (AWS/GCP/Azure)
- Hands‑on with Kubernetes and distributed systems
- Experience in on‑call rotation and incident management
- Proficient in at least one programming language (Go, Python, Java)
#J-18808-Ljbffr