Unknown Company

Senior AI Distributed Systems Engineer

Remote • Posted 4 days ago
Onsite Full Time IT & Technology

Cerence Inc. is seeking an ML Infrastructure Engineer to design and operate distributed training systems for large neural networks across GPU clusters.

You will optimize multi-node, multi-GPU execution and diagnose bottlenecks in compute, memory, and networking to maximize throughput and training stability. You will partner with research and applied ML teams to productionize large-model training pipelines using PyTorch Distributed, Megatron-LM, and DeepSpeed, with emphasis on scalable,

#J-18808-Ljbffr
Back to Job Search