NVIDIA in Seattle is seeking a Senior Research Engineer focused on Generative AI inference. You will design routing policies for LLM traffic to balance accuracy and latency, and build agentic benchmarks to calibrate models across NVIDIA’s serving stack.
Collaborate with research and engineering teams to contribute design docs, code reviews, and open-source contributions, while mentoring junior engineers and driving experimental validation.
#J-18808-Ljbffr