River AI in Palo Alto, CA, seeks exceptional systems engineers to build the distributed training engines behind the River API. Your work will focus on making fine-tuning fast, numerically correct, and reliable across large GPU clusters.
You will own execution of training workloads, including gradient computation, optimizer updates, rollout coordination, and checkpoint recovery. Collaborate with researchers and inference engineers to deploy new learning methods and improve compute efficiency.
#J-18808-LjbffrDistributed Training Engineer — High-Perf GPU Scale in palo alto at Unknown Company
This position is listed as full time and onsite.