Unknown Company

Remote ML Inference Engineer - Serving & Performance

Remote • Posted 3 days ago
Remote Full Time IT & Technology

Yobitel Communications is seeking an engineer to own the model serving layer, optimizing latency and throughput for scale. You will manage vLLM, TensorRT-LLM, and Triton deployments, shaping the benchmarking discipline that keeps performance honest.

You will drive workload-specific serving strategies, quantization choices, and rigorous benchmarking to stay ahead in a fast-moving field. This role is remote-friendly and full-time, with a focus on high-performance inference at scale.

#J-18808-Ljbffr
Back to Job Search