Unknown Company

GenAI Evaluation Scientist - LLM Benchmarks & Failures

seattle, wa • Posted 1 weeks ago
Onsite Full Time Bio & Pharmacology & Health

Scale is seeking a Machine Learning Research Scientist, Evaluations to join the GenAI Research Organization in San Francisco. You will develop rigorous evaluations, diagnose failure modes in frontier LLMs and agents, and design benchmarks for text and multimodal modalities.

Collaboration with researchers and engineers will shape evaluation-driven AI development. The role emphasizes post-training techniques like SFT and RLHF, with opportunities to publish findings at top conferences and influence

#J-18808-Ljbffr

GenAI Evaluation Scientist - LLM Benchmarks & Failures in seattle at Unknown Company

This position is listed as full time and onsite.

Back to Job Search