NVIDIA is seeking a Sr. Inference Engineer to push LLM inference performance through GPU kernel optimization and silicon-backed benchmarking. The role collaborates with compiler, kernel, hardware, and framework teams to surface bottlenecks and deliver measurable gains.
You will drive three streams: microbenchmarking real silicon kernels, end-to-end performance analysis, and agentic optimization using AI-driven methods. A strong background in Python/C++ and GPU profiling is essential.
#J-18808-Ljbffr