Unknown Company

Scale Low-Latency ML Inference Engineer (GPU/CUDA)

northern, ky • Posted 1 weeks ago
Hybrid Full Time Electrical & Energy Engineering

AI Breaking Wire seeks a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide.

The role emphasizes C++ and Python proficiency, GPU optimization with CUDA and Triton, and collaboration with research teams to productionize new architectures. This hybrid position is based in San Francisco.

#J-18808-Ljbffr

Scale Low-Latency ML Inference Engineer (GPU/CUDA) in northern at Unknown Company

This position is listed as full time and hybrid.

Back to Job Search