Leading the infrastructure sub-team, the full-time salaried Tech Lead Manager, GPU Infrastructure will set the technical direction and roadmap for a remote team, manage platform architecture, and oversee the hiring and growth of senior engineers while ensuring the performance and fault tolerance of large-scale experiments. Key responsibilities Set the platform's technical direction and own its roadmap, including system operations and performance metrics Manage architecture, scheduling, and storage design while debugging failures across multiple layers Hire and mentor a small team of senior engineers, establishing project priorities and providing feedback Required qualifications 5+ years of experience in systems or infrastructure engineering, particularly with production Linux and GPU platforms Proven experience running production Kubernetes for GPU workloads, including batch processing Expertise in infrastructure as code and observability, utilizing tools like Terraform, Ansible, and Prometheus Strong programming skills in languages commonly used for infrastructure, such as Python, Go, Rust, or C++ Demonstrated leadership experience in managing engineers or technical teams, with a focus on setting direction and scoping projects
Tech Lead Manager, GPU Infrastructure in workfromhome at Unknown Company
This position is listed as full time and able to be worked remotely.