Scale AI is seeking Research Scientists and Research Engineers to advance evaluation for LLMs and multimodal models. You will analyze failures, design rigorous benchmarks, and connect findings to data and training interventions.
You’ll publish results at top AI conferences and collaborate with leading labs to influence future model development. The role emphasizes strong communication, deep learning expertise, and experience with post-training techniques like RLHF and preference modeling in
#J-18808-LjbffrGenAI Evaluation Scientist: LLM Benchmarks & Diagnostics in san francisco at Unknown Company
This position is listed as full time and onsite.