Join one of the most sophisticated trading firms in the world as they build a new machine learning initiative from the ground up. The team is currently just three people, including engineers from two competitors and a researcher from a top FAANG AI lab, and they're looking to add a few key hires who will help define the future of ML infrastructure at the firm.
This particular hire will focus on model serving, inference optimization, and deploying large-scale models onto GPU infrastructure. The team is already operating at scale with over 1,000 GPUs today and is rapidly expanding toward 10,000+ GPUs, creating unique engineering challenges around performance, efficiency, and hardware utilization.
Rather than joining a large, established ML organization, you'll have significant ownership over architecture, technical direction, and platform decisions from day one.
What You Will Do
- Design and build large-scale serving and inference infrastructure for state-of-the-art ML models.
- Optimize model performance, latency, throughput, and GPU utilization in production environments.
- Deploy and scale models across thousands of GPUs while improving efficiency and reliability.
- Develop model compression, quantization, distillation, and other optimization techniques to maximize hardware performance.
- Work on taking cutting-edge models from research into production by bringing them closer to the hardware.
- Partner closely with researchers and engineers to accelerate experimentation and deployment.
- Help define the architecture and roadmap for a rapidly growing ML platform.
What You Bring
- 8+ years of software engineering, systems engineering, or machine learning infrastructure experience.
- Deep experience with model serving, inference, and performance optimization at scale.
- Strong understanding of model compression, quantization, kernel optimization, and GPU acceleration.
- Strong knowledge of CUDA, GPU programming, and hardware-aware optimization.
- Expert-level programming skills in C++ and Python.
- Experience with modern ML frameworks such as PyTorch, JAX, or TensorFlow.
- BS, MS, or PhD in Computer Science, Engineering, Mathematics, or a related field.
Why Consider It
- Ground-floor opportunity to build a new ML platform inside one of the world's leading trading firms.
- Work alongside engineers and researchers from top AI labs and elite quantitative trading firms.
- Ownership over critical infrastructure supporting next-generation ML workloads.
- Massive scale, with GPU infrastructure growing from the low thousands to 10s of thousands
- Solve challenging problems at the intersection of distributed systems, machine learning, and high-performance computing.
This role can sit out of NYC or Chicago
#J-18808-LjbffrSenior ML Engineer, Serving & Optimization in new york at Unknown Company
This position is listed as full time and onsite.