ML Systems EngineerYou'll work alongside some of the world's leading ML systems engineers, including leaders behind Megatron-LM, SGLang, Liger Kernel, TorchRec, CleanRL, TorchRL, and JAX-MD.We're looking for exceptional ML Systems Engineers to build the agentic infrastructure powering our large-scale training, inference, and reinforcement learning. You'll own critical pieces of the ML systems stack to maximize performance, scalability, reliability, and productivity for both engineers and AI agents.What You'll DoBuild and optimize large-scale training and reinforcement learning infrastructure while ensuring its correctnessDevelop high-performance inference and serving systemsDesign distributed runtimes and scheduling systems for complex ML workloadsBuild secure and large-scale sandboxing and execution environmentsOptimize memory, GPU kernels and communication for maximum throughput and end-to-end efficiencyImprove scalability, reliability, and efficiency across the ML systems stackWhat We're Looking ForStrong systems programming and performance engineering skillsExperience building high-performance ML infrastructure at scaleAbility to own complex technical problems end-to-endStrong coding ability and engineering judgment, including the ability to work effectively with AI agents to design, implement, test, and debug complex systemsHigh ownership, fast execution, and a passion for pushing the frontier of AI systems and accelerating scientific discoveryYou should have deep expertise in at least one of the following:Training: Strong experience building, debugging and optimizing large-scale training systems with Megatron-LM. Familiarity with TorchTitan, FSDP, veRL, Slime, or other distributed training systems is a plus.Distributed Runtime: Strong experience with Ray.
Familiarity with Monarch or other distributed execution frameworks is a plus.Inference: Strong experience with SGLang. Familiarity with vLLM, TensorRT-LLM, or production LLM serving systems is a plus.Sandboxing: Strong experience with secure execution environments, containers, virtualization, or code sandboxing.GPU Kernels: Strong experience with CUDA, Triton, CUTLASS, CuTe, or custom GPU kernel development.GPU Communication: Strong experience with NCCL, NVLink, InfiniBand, RDMA, GPUDirect RDMA, or large-scale communication optimization.MechanicsMinimum education: Bachelor's degree or similar experienceLocation: Menlo Park, CA (Soon: San Francisco, too)Compensation: $250,000-$350,000 base + equityVisa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process with our legal support.We're building a team of the world's best — the scientists, engineers, and problem-solvers who don't just follow the frontier, they define it. If you're driven to bring AI to life in the physical world and make discoveries that have never been made before, you belong here.