Unknown Company

Distributed AI Inference Performance Engineer

santa clara, ca • Posted 4 days ago
Onsite Full Time IT & Technology

NVIDIA Gruppe in Santa Clara is looking for an engineer to optimize and benchmark GenAI inference on cutting-edge accelerators. This role involves owning end-to-end optimization pipelines and defining next-generation inference benchmarks across multiple platforms.

The ideal candidate will have a strong background in software development, particularly in Python or C++, and a deep understanding of LLM/VLM architectures. Join our team to influence the ecosystem and contribute to impactful open-source projects.

#J-18808-Ljbffr
Back to Job Search