Unknown Company

LLM Inference Engineer: GPU-Optimized Serving

san francisco, ca • Posted 2 days ago
Onsite Full Time Electrical & Energy Engineering

GMI Cloud in San Francisco is building the leading inference optimization solution and the most advanced token platform in the global token market — and we are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability.

GMI Cloud invites engineers to work on frontier research across quantization, speculative decoding, KV cache & memory management, and PD disaggregation, partnering with platform,

#J-18808-Ljbffr

LLM Inference Engineer: GPU-Optimized Serving in san francisco at Unknown Company

This position is listed as full time and onsite.

Back to Job Search