Unknown Company

Inference Performance Engineer: Optimize Model Serving

san francisco, ca • Posted 1 weeks ago
Onsite Full Time IT & Technology

Adaption in San Francisco Bay Area seeks a senior ML systems engineer to own cost and performance of the inference stack. You will optimize caching, batching, quantization, and decoding, while collaborating with the serving fleet to deliver scalable, low-latency model serving.

Required are 5+ years in ML systems with deep knowledge of model serving, GPU performance, and proficiency in Python and C++/Rust. You’ll work with engines like vLLM and TensorRT-LLM and contribute to profiling and

#J-18808-Ljbffr

Inference Performance Engineer: Optimize Model Serving in san francisco at Unknown Company

This position is listed as full time and onsite.

Back to Job Search