Operating the Kubernetes GPU fleet and managing multi-tenancy, the full-time Senior Software Engineer, GPU Infrastructure will focus on maintaining large-scale experiments, ensuring fault tolerance, and enhancing platform security in a remote work environment. Key responsibilities Handle day-to-day operations of the Kubernetes GPU fleet, including capacity planning and node lifecycle management Design and implement batch scheduling and storage solutions for the GPU cluster, ensuring efficient resource allocation Collaborate directly with research teams to address infrastructure challenges and improve platform functionality Required qualifications 3+ years of experience in systems or infrastructure engineering on production Linux, specifically with GPU or large-scale batch platforms Proven experience managing production Kubernetes for GPU workloads, including batch layers and resource management Familiarity with infrastructure as code tools such as Terraform or Ansible, and monitoring solutions like Prometheus Strong programming skills in languages commonly used for infrastructure, such as Python, Go, Rust, or C++ Ability to write clear documentation for engineers, researchers, and providers, including design documents and incident reports
Senior Software Engineer, GPU Infrastructure in workfromhome at Unknown Company
This position is listed as full time and able to be worked remotely.