NVIDIA is seeking a Sr. Inference Engineer to push LLM inference performance through GPU kernel optimization.
You will help develop silicon-measured benchmarking, model-level performance projection tooling, and agentic optimization systems, collaborating across compiler, hardware, and framework teams to surface bottlenecks and deliver measurable gains. You will work on GPU kernel microbenchmarking, end-to-end model performance analysis, and agentic optimization, shaping production inference
#J-18808-Ljbffr