United States Digital Space LLC is seeking a Site Reliability Engineer (STARSHIELD) to design, operate, and scale GPU/CPU infrastructure for Top Secret datacenters and AI clusters. The role focuses on automation, on-prem Kubernetes, and collaboration with AI teams to deliver scalable, reliable software products.
The position requires SRE/DevOps experience, Linux proficiency, and familiarity with Terraform/Ansible. A TS/SCI-like clearance and willingness to travel are important for success.
#J-18808-LjbffrSite Reliability Engineer — AI GPU Infrastructure in Location not specified at Unknown Company
This position is listed as full time and onsite.