Unknown Company

Senior SRE: GPU Fleet Orchestration & Auto-Scaling

menlo park, ca • Posted 1 weeks ago
Onsite Full Time IT & Technology

Hippocratic AI is seeking a Senior Site Reliability Engineer to own the GPU management and scheduling platform that runs a fleet of ~30 models on heterogeneous hardware.

You will design metrics, admission control, and autoscaling, build infrastructure automation with Terraform and CI/CD, and operate secure production systems on AWS, GCP, or Azure.

Join a team building healthcare AI at scale, mentoring engineers and collaborating with researchers to ensure reliability, performance, and safety.

#J-18808-Ljbffr

Senior SRE: GPU Fleet Orchestration & Auto-Scaling in menlo park at Unknown Company

This position is listed as full time and onsite.

Back to Job Search