The Institute of Foundation Models (IFM) is seeking a deeply technical engineer to co‑design and optimize the communication stack for large‑scale distributed training, including hybrid parallelism and MoE workloads.
This role focuses on performance engineering, distributed debugging, and cross‑layer optimization across thousands of GPUs, with a strong emphasis on fault‑tolerant execution and topology‑aware design.
#J-18808-Ljbffr