Unknown Company

SRE - Compute & Hyperscale GPU Fleet Reliability

new york, new york • Posted 1 weeks ago
Onsite Full Time General

Fluidstack is seeking a Production Engineer to own compute fleet health end-to-end, build the observability and automation that scales GPU infrastructure, and drive reliability across Kubernetes-managed and bare-metal environments. You will define failure modes, implement triage automation, and own the GPU qualification and firmware tooling.

The role emphasizes end-to-end ownership, rapid incident response, and fluency with AI tooling and modern automation.

#J-18808-Ljbffr
Back to Job Search