Thinking Machines is hiring a Site Reliability Engineer in San Francisco to own the health of post-training and RL pipelines, partnering with researchers to unblock training and harden infrastructure. Youll implement recovery, observability, and tooling to keep runs fast and reliable.
Youll operate in a high-scale ML environment, collaborate across teams, and participate in on-call rotations while driving permanent fixes and reducing toil for researchers.
#J-18808-LjbffrSRE for AI Training Pipelines & RL Runs in san francisco at Unknown Company
This position is listed as full time and onsite.