Unknown Company

GPU Fleet SRE for AI Inference Platform

san francisco, ca • Posted 1 weeks ago
Onsite Full Time Engineering

Beam is an ultrafast AI inference platform building a serverless runtime that launches GPU-backed containers in under a second and scales to thousands of GPUs.

The role centers on owning compute fleet health end to end, automating failure triage, and designing a GPU qualification platform with burn-in and baselining for new hardware. You will also own firmware telemetry and low-level log collection critical to repair automation.

#J-18808-Ljbffr
Back to Job Search