Job Title
Design and maintain highly available, scalable, and fault-tolerant systems
Implement and manage monitoring, logging, and observability tools Grafana, Prometheus, etc.
Participate in incident management, RCA, and postmortems
Automate operational tasks using scripting and tools Python, Bash, etc.
Collaborate with
Site Reliability Engineer (SRE) in plano at Unknown Company
This position is listed as full time and onsite.