Lambda is hiring a Senior Site Reliability Engineer to build and operate monitoring for AI cloud infrastructure, deploy large-scale HPC clusters remotely, and automate lifecycle tasks using Ansible and Terraform.
You will lead incident response, collaborate with on-site deployment teams, and help define SOPs while continuously improving reliability and performance for GPU-focused workloads.
#J-18808-LjbffrAI Cloud HPC Engineer — Build & Automate Massive Clusters in Location not specified at Unknown Company
This position is listed as full time and able to be worked remotely.