Adaption Labs in San Francisco is looking for a senior ML systems engineer to own the cost and performance of the inference stack. You’ll optimize caching, batching, quantization, and kernel-level performance to improve throughput and reduce latency without sacrificing model quality.
You’ll collaborate with the serving fleet engineers and tune routing, profiling, and measurement systems across production workloads. A strong background in Python/C++/Rust and GPU optimization is required.
#J-18808-LjbffrML Inference Performance Engineer — Optimize Cost & Latency in san francisco at Unknown Company
This position is listed as full time and onsite.