Leading a distributed engineering team, the full-time Technical Lead - GPU Infrastructure will own the architecture and delivery of a GPU compute platform, focusing on building and operating a managed Slurm service and Kubernetes control plane, all while working remotely. Key responsibilities Own the end-to-end platform architecture, including design proposals and current baseline maintenance Lead and manage a distributed team across various engineering disciplines, ensuring adherence to engineering standards and performance growth Design, build, and operate a managed Slurm scheduling layer for research users, ensuring efficient resource allocation and system health Required qualifications Eight or more years of hands-on engineering experience, with at least three years in a leadership role building infrastructure platforms Proven experience with Slurm at scale, including operational management of HPC or GPU training clusters Deep knowledge of GPU fleet operation on bare metal, including NVIDIA driver and CUDA lifecycle management Strong background in production Kubernetes operations, including control plane management and multi-tenancy design Working fluency in JavaScript and Node.js for architecture decision-making and code review
Technical Lead - GPU Infrastructure in workfromhome at Unknown Company
This position is listed as full time and able to be worked remotely.