NVIDIA is seeking a Senior Inference Engineer to push GPU kernel optimization for LLM inference across silicon-level measurements. You will help build benchmarking infrastructure, model-level performance tooling, and optimization policies for production deployments.
The role collaborates with compiler, hardware, kernel, and framework teams to surface bottlenecks and deliver measurable gains, in a full-time, on-site capacity in Santa Clara.
#J-18808-Ljbffr