Our client is a high-growth GPU infrastructure provider delivering factory-scale GPU-as-a-Service to organisations running large-scale training and high-throughput inference workloads.
They operate full-stack GPU environments end-to-end — including bare metal, Kubernetes control planes, storage, networking, observability, and day-two operations — enabling customers to deploy AI models with predictable performance and reliability.
The Role
This is a hands-on Technical Account Manager role supporting customers running production AI workloads on large-scale GPU clusters.
You will:
- Act as the primary technical contact for a portfolio of customers
- Troubleshoot production issues across GPU infrastructure and Kubernetes
- Lead incident triage, mitigation, and communication
- Support distributed training and inference workloads
- Help customers optimise performance, reliability, and utilisation
- Create documentation, runbooks, and best practices
- Run regular technical check-ins and health reviews
This is an engineering-led TAM position — not a ticket-routing role.
Required Experience
- 7+ years in customer-facing infrastructure roles
- Strong hands-on Kubernetes experience (networking, storage, scaling, troubleshooting)
- Experience operating bare metal GPU clusters
- Linux systems and networking fundamentals
- GPU drivers and CUDA stack familiarity
- Experience supporting distributed training and/or inference workloads in production
- Ability to lead incidents and communicate clearly under pressure