Serving as the operational linchpin of a global fleet, the full-time remote Datacenter Infrastructure Specialist will manage the technical lifecycle and operational health of a high-density GPU fleet, ensuring uptime and performance for demanding AI workloads. Key Responsibilities Assist in validating new hardware and ensuring partner deployments meet specifications for distributed AI/ML workloads Monitor fleet health to identify performance degradation and provide technical data to protect customer SLAs Coordinate technical incident communications and automate network triage using AI-driven operations Required Qualifications 3-5 years of experience in infrastructure operations, systems reliability, or datacenter engineering Strong proficiency in datacenter networking and performance troubleshooting, with exposure to RDMA, InfiniBand, or RoCE preferred Hands-on experience with the NVIDIA Software Stack and understanding of multi-node performance tuning Solid Linux system administration skills and experience with containerization (Docker) Clear written and verbal communication skills for explaining technical issues to partners and leadership