Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. We are seeking an experienced SRE/DevOps engineer to build and operate monitoring, alerting, and automation for large-scale AI workloads.
You will deploy HPC clusters remotely, automate lifecycle tasks as code, and drive reliability through runbooks and incident response. You will collaborate with on-site teams, optimize Linux-based distributed systems, and contribute to standard
#J-18808-LjbffrSenior HPC Cloud Engineer for AI Infrastructure in san jose at Unknown Company
This position is listed as full time and able to be worked remotely.