NVIDIA in New York seeks a Senior Inference Engineer to push GPU kernel optimization for LLM inference. You will drive microbenchmarking, connect performance evidence to model economics, and develop agentic optimization policies with teams across compiler, hardware, kernel, and frameworks.
Requires an advanced degree, 6+ years' experience, and expertise in Python/C++, GPU profiling (CUPTI/NSYS/NCU), and LLM frameworks (TRT-LLM, SGLang, vLLM). Equity and benefits accompany the role.
#J-18808-Ljbffr