About the Role
We are seeking a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide.
Responsibilities
- Architect, build, and scale low-latency distributed inference serving systems for massive generative models.
- Optimize GPU utilization, memory management, and kernel execution for state-of-the-art transformer architectures.
- Collaborate with research teams to ensure smooth transition of new model architectures into production.
- Monitor system performance, troubleshoot bottlenecks, and implement robust reliability measures.
Requirements
- BS, MS, or Ph.D. in Computer Science or related technical field.
- 3+ years of industry experience building large-scale distributed systems or ML infrastructure.
- Deep proficiency in C++ and Python.
- Extensive experience with CUDA, Triton, or deep learning hardware accelerators.
- Familiarity with distributed training and inference frameworks (vLLM, TensorRT-LLM, Megatron).
Benefits
- Top-tier compensation including equity.
- Full medical, dental, and vision coverage with zero employee contribution options.
- Unlimited paid time off and flexible hybrid work policy.
- Catered daily lunches and wellness stipends.
Machine Learning Engineer, Inference Infrastructure in northern at Unknown Company
This position is listed as full time and hybrid.