Fluidstack is seeking a Production Engineer to own compute fleet health end-to-end, build the observability and automation that scales GPU infrastructure, and drive reliability across Kubernetes-managed and bare-metal environments. You will define failure modes, implement triage automation, and own the GPU qualification and firmware tooling.
The role emphasizes end-to-end ownership, rapid incident response, and fluency with AI tooling and modern automation.
#J-18808-LjbffrSRE - Compute & Hyperscale GPU Fleet Reliability in new york at Unknown Company
This position is listed as full time and onsite.