Unknown Company

Senior ML Engineer, Serving & Optimization

new york, ny • Posted 4 days ago
Onsite Full Time IT & Technology

Join one of the most sophisticated trading firms in the world as they build a new machine learning initiative from the ground up. The team is currently just three people, including engineers from two competitors and a researcher from a top FAANG AI lab, and they're looking to add a few key hires who will help define the future of ML infrastructure at the firm.

This particular hire will focus on model serving, inference optimization, and deploying large-scale models onto GPU infrastructure. The team is already operating at scale with over 1,000 GPUs today and is rapidly expanding toward 10,000+ GPUs, creating unique engineering challenges around performance, efficiency, and hardware utilization.

Rather than joining a large, established ML organization, you'll have significant ownership over architecture, technical direction, and platform decisions from day one.

What You Will Do

  • Design and build large-scale serving and inference infrastructure for state-of-the-art ML models.
  • Optimize model performance, latency, throughput, and GPU utilization in production environments.
  • Deploy and scale models across thousands of GPUs while improving efficiency and reliability.
  • Develop model compression, quantization, distillation, and other optimization techniques to maximize hardware performance.
  • Work on taking cutting-edge models from research into production by bringing them closer to the hardware.
  • Partner closely with researchers and engineers to accelerate experimentation and deployment.
  • Help define the architecture and roadmap for a rapidly growing ML platform.

What You Bring

  • 8+ years of software engineering, systems engineering, or machine learning infrastructure experience.
  • Deep experience with model serving, inference, and performance optimization at scale.
  • Strong understanding of model compression, quantization, kernel optimization, and GPU acceleration.
  • Strong knowledge of CUDA, GPU programming, and hardware-aware optimization.
  • Expert-level programming skills in C++ and Python.
  • Experience with modern ML frameworks such as PyTorch, JAX, or TensorFlow.
  • BS, MS, or PhD in Computer Science, Engineering, Mathematics, or a related field.

Why Consider It

  • Ground-floor opportunity to build a new ML platform inside one of the world's leading trading firms.
  • Work alongside engineers and researchers from top AI labs and elite quantitative trading firms.
  • Ownership over critical infrastructure supporting next-generation ML workloads.
  • Massive scale, with GPU infrastructure growing from the low thousands to 10s of thousands
  • Solve challenging problems at the intersection of distributed systems, machine learning, and high-performance computing.

This role can sit out of NYC or Chicago

#J-18808-Ljbffr

Senior ML Engineer, Serving & Optimization in new york at Unknown Company

This position is listed as full time and onsite.

Back to Job Search