Jobtailor is seeking an experienced Site Reliability Engineer to design, operate, and scale reliable cloud-native systems for AI workloads. You will own production reliability, on-call practices, and incident response, while driving observability and postmortem learning across our platform.
The role requires hands-on Kubernetes and distributed systems expertise, plus strong programming skills (Go/Python/Java) and 8+ years in SRE/Infra/Platform Engineering.
#J-18808-Ljbffr