Luma AI is seeking an engineer to design and optimize distributed training systems for large multimodal models. You will work on multi-GPU clusters, implement advanced parallelization, and build tools to monitor and debug massive training runs.
Requires extensive distributed PyTorch experience, deep GPU cluster knowledge, and familiarity with NCCL and MPI. Experience with containerization and orchestration is beneficial. This role is a high-depth engineering position at Luma AI.
#J-18808-Ljbffr