Joining a small research infrastructure team, the full-time remote HPC Infrastructure Engineer will operate and improve NVIDIA GPU clusters, focusing on automation, performance tuning, and hardware management to enhance research throughput and efficiency. Key responsibilities Operate and enhance the GPU fleet, including provisioning, scheduling, monitoring, and upgrades Build automation to maintain fleet health and streamline operations without human intervention Evaluate and benchmark rented GPU capacity while ensuring security and access control for clusters Required qualifications Experience managing large-scale Linux server or GPU environments in production Proficient with the NVIDIA stack, including drivers, CUDA, and NCCL Strong skills in automation using Python and/or Bash, with familiarity in IaC tools like Ansible or Terraform Ability to analyze metrics and logs to troubleshoot performance issues effectively Comfortable with hands-on hardware work and datacenter operations
HPC Infrastructure Engineer in workfromhome at Unknown Company
This position is listed as full time and able to be worked remotely.