Hippocratic AI is seeking a Senior Site Reliability Engineer to own a complex GPU management platform that schedules inference calls across a fleet of ~30 models on heterogeneous hardware. You will build the metrics, admission control, and autoscaling systems, while ensuring secure, scalable, and compliant production environments on cloud providers.
A decade in SRE/DevOps with strong software skills is required.
#J-18808-LjbffrSenior GPU SRE: GPU Scheduling & Autoscaling Platform in menlo park at Unknown Company
This position is listed as full time and onsite.