Scale Labs in New York is seeking a Research Scientist focused on Safety Post-Training to develop methods and interpretability techniques that make frontier AI systems safer and better understood by researchers and policymakers.
You will design post-training pipelines, develop interpretability-informed evaluations, and collaborate with policymakers, engineers, and researchers to translate findings into tangible safety standards and benchmarks.
#J-18808-LjbffrAI Safety Scientist (Post-Training & Interpretability) in san francisco at Unknown Company
This position is listed as full time and onsite.