Thinking Machines is hiring an engineer to ensure the reliability of our GPU supercomputing fleet, owning the seam between hardware, firmware, and operating system. You will diagnose hardware anomalies, track root causes to the hardware, and coordinate fixes with vendors so researchers can run at scale.
Based in San Francisco, this full-time role requires owning drivers, kernel surfaces, and diagnostics, plus automating fleet monitoring and reliability improvements across multi-disciplinary
#J-18808-LjbffrGPU Supercomputing Reliability Engineer — Unlimited PTO in san francisco at Unknown Company
This position is listed as full time and onsite.