NVIDIA is building an RL post-training infrastructure team to scale experimental workflows from a single GPU to thousands of nodes, delivering reliable, high-performance runtimes for researchers. You will collaborate with researchers and labs, optimize PyTorch-based RL loops, and improve fault tolerance, elastic scaling, and portability across CPU, GPU, and LPUs.
This role offers the chance to contribute to VeRL, Miles, TorchTitan and related ecosystems while shaping production-grade distributed
#J-18808-LjbffrRL Post-Training Systems Engineer (Distributed Infra) in new york at Unknown Company
This position is listed as full time and onsite.