Unknown Company

Machine Learning Engineer - Inference

new york, ny • Posted 1 weeks ago
Onsite Full Time Electrical & Energy Engineering

A leading quantitative trading firm is seeking a Machine Learning Engineer specializing in inference to join a highly technical AI research team developing and deploying large-scale machine learning models for financial markets.

The team builds powerful foundation models for markets, trained on vast quantities of market and alternative data to predict future market behavior. These models are deployed directly into live trading environments, making inference speed, efficiency, and reliability critical to the business.

As a Machine Learning Engineer, you will work at the intersection of machine learning, high-performance computing, and systems engineering, with a broad mandate to improve large-scale model inference.

What You’ll Work On

  • Design and optimize high-performance inference systems for large-scale deep learning models
  • Develop and optimize GPU kernels using CUDA, Triton, Pallas, CuTe DSL, and related technologies
  • Improve lower-level performance across PyTorch, JAX, XLA, and CUDA Graphs
  • Explore and develop inference solutions across GPUs, ASICs, and FPGAs
  • Optimize data streaming and model-serving infrastructure for demanding real-time environments
  • Work closely with ML researchers to co-design efficient inference architectures
  • Improve latency, throughput, hardware utilization, and overall inference efficiency
  • Help shape the team’s broader machine learning and systems research agenda

The inference environment spans multiple platforms deployed globally and supports a variety of model architectures and trading strategies. The work is highly performance-sensitive, technically challenging, and has a direct impact on live trading performance.

Qualifications

  • 2+ years of professional experience building deep learning or machine learning systems
  • Strong software engineering and systems fundamentals
  • Experience building deep learning systems in computationally intensive domains such as robotics, recommendation systems, biology, chemistry, physics, audio, video, or similar areas
  • Ability to translate techniques and approaches across different machine learning domains

Plus experience with one or more of the following:

  • Lower-level PyTorch, JAX, or XLA development
  • CUDA Graphs
  • GPU performance optimization
  • FPGA or ASIC development
  • High-performance ML inference systems

Nice to Have

  • Experience with large-scale or low-latency model inference
  • LLM or foundation model experience
  • Experience optimizing GPU kernels or distributed ML workloads

Prior finance or trading experience is not required.

#J-18808-Ljbffr

Machine Learning Engineer - Inference in new york at Unknown Company

This position is listed as full time and onsite.

Back to Job Search