United States Digital Space LLC is seeking an experienced engineer to optimize model serving pipelines using vLLM, TensorRT, and Triton. You will own the inference layer and drive benchmarking to ensure low latency and high throughput at scale.
You'll implement quantization strategies (FP8, AWQ, GPTQ) and build eval harnesses and InferenceBench-style tests to compare against industry standards. This is a remote, full-time role.
#J-18808-Ljbffr