Optimizing AI inference configurations, the full-time Inference Performance Engineer will enhance performance on large-scale benchmarks by developing reusable workflows and methodologies, profiling workloads, and collaborating with cross-functional teams, with opportunities for both onsite and remote work. Key responsibilities Develop reusable skills and workflows for autonomous AI agents to optimize performance in AI inference workloads Measure and enhance throughput-per-GPU and user interactivity through various optimization techniques and configurations Collaborate with multiple teams to translate profiling insights into effective performance improvements across NVIDIA's GPU platforms Required qualifications BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Math, or a related field, or equivalent experience 3+ years of relevant engineering experience in AI model execution optimization Extensive knowledge of performance benchmarking and profiling GPU workloads using tools like Nsight Systems and PyTorch profiler Strong Python engineering skills and experience with large C++/CUDA codebases Rigorous experimental methodology for controlled comparisons and reproducible benchmarks