Unknown Company

Senior Inference Engineer: GPU Kernel Optimization for LLMs

new york, ny • Posted 4 days ago
Onsite Full Time IT & Technology

NVIDIA in New York seeks a Senior Inference Engineer to push GPU kernel optimization for LLM inference. You will drive microbenchmarking, connect performance evidence to model economics, and develop agentic optimization policies with teams across compiler, hardware, kernel, and frameworks.

Requires an advanced degree, 6+ years' experience, and expertise in Python/C++, GPU profiling (CUPTI/NSYS/NCU), and LLM frameworks (TRT-LLM, SGLang, vLLM). Equity and benefits accompany the role.

#J-18808-Ljbffr
Back to Job Search