Scale AI, Inc. seeks Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation.
The role focuses on building benchmarks and diagnosing model failure modes in text and multimodal modalities within the GenAI Research Organization. You will develop rigorous evaluations, collaborate with researchers and engineers, and translate failure analysis into input for next-generation generative AI models.
#J-18808-LjbffrGenAI Evaluation Scientist: Benchmarks & Failures in san francisco at Unknown Company
This position is listed as full time and onsite.