Unknown Company

AI Safety Scientist (Post-Training & Interpretability)

san francisco, ca • Posted 2 weeks ago
Onsite Full Time Bio & Pharmacology & Health

Scale Labs in New York is seeking a Research Scientist focused on Safety Post-Training to develop methods and interpretability techniques that make frontier AI systems safer and better understood by researchers and policymakers.

You will design post-training pipelines, develop interpretability-informed evaluations, and collaborate with policymakers, engineers, and researchers to translate findings into tangible safety standards and benchmarks.

#J-18808-Ljbffr

AI Safety Scientist (Post-Training & Interpretability) in san francisco at Unknown Company

This position is listed as full time and onsite.

Back to Job Search