Physical Intelligence in San Francisco is building the core ML infrastructure to scale training from prototype to production-grade runs. The ML Infrastructure team owns training/inference systems, scheduling, checkpointing, and metrics collection.
You will scale JAX-based training across TPU and GPU clusters, profile memory usage, improve throughput, and create abstractions for launching and monitoring experiments in close collaboration with researchers.
#J-18808-LjbffrML Infrastructure Engineer: Scale & Performance in san francisco at Unknown Company
This position is listed as full time and onsite.