NVIDIA AI in Redmond seeks a recent graduate with an MS/PhD in CS to collaborate with AI/ML research teams, identify infrastructure gaps, and implement scalable solutions on GPU clusters.
You will monitor performance, optimize for high availability, and ensure researchers have efficient resources using PyTorch, Kubernetes, Slurm, and Docker.
Role focuses on distributed training, data processing, and model inference across HPC workloads with equity and comprehensive benefits.
#J-18808-LjbffrAI/ML Infra Engineer — GPU Clusters & HPC in redmond at Unknown Company
This position is listed as full time and onsite.