NVIDIA Corporation in Santa Clara, CA, seeks a Sr. Inference Engineer to accelerate LLM inference through GPU kernel optimization. You will lead kernel benchmarking, model-level performance analysis, and AI-driven optimization workflows across silicon and software stacks.
The role requires strong Python and C++, hands-on GPU profiling (CUPTI/NSYS/NCU), and experience with TRT-LLM, SGLang, or vLLM. Collaboration with compiler, hardware, kernel, and framework teams is essential to deliver
#J-18808-LjbffrSenior Inference Engineer: AI-Driven GPU Kernel Optimization in santa clara at Unknown Company
This position is listed as full time and onsite.