Nscale seeks a Principal Infrastructure Engineer to own AI cluster performance and validation across thousands of GPUs. You’ll define healthy-at-scale criteria, set the architecture, and partner with senior leaders to production-readiness goals.
You’ll run real AI workloads at scale, diagnose performance regressions, design scalable validation systems, and optimize fabric, storage, and schedulers to deliver reliable, cost-efficient AI clusters.
#J-18808-LjbffrSenior AI Cluster Performance & Validation Engineer in seattle at Unknown Company
This position is listed as full time and onsite.