NVIDIA AI in Santa Clara is seeking a highly capable software engineer to advance an advanced inference framework using modern C++. The role focuses on extending TensorRT with autoregressive model serving capabilities and requires collaboration across CUDA, kernel libraries, compilers, and robotics teams to deliver high-performance, production-ready solutions.
The candidate should hold a BS/MS/PhD (or equivalent) and have at least four years of software development experience with a deep
#J-18808-LjbffrEdge AI LLM Inference Architect (C++, TensorRT) in santa clara at Unknown Company
This position is listed as full time and onsite.