Graphcore seeks an experienced Site Reliability Engineering leader to build and lead a new SRE organization responsible for the production operation of a rapidly scaling AI supercomputing platform. This role blends hands-on engineering with organizational leadership to ensure 24x7x365 availability and reliability across compute, networking, storage, and orchestration layers.
The SRE Manager will hire, mentor, and develop the team, define operating models, and drive automation and reliability
#J-18808-LjbffrSRE Lead for 24x7 AI Platform Reliability in austin at Unknown Company
This position is listed as full time and onsite.