AI Breaking Wire seeks a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide.
The role emphasizes C++ and Python proficiency, GPU optimization with CUDA and Triton, and collaboration with research teams to productionize new architectures. This hybrid position is based in San Francisco.
#J-18808-LjbffrScale Low-Latency ML Inference Engineer (GPU/CUDA) in northern at Unknown Company
This position is listed as full time and hybrid.