Mercor is seeking PhD and Master's scientists to author AI evaluation tasks (Sci Code) and contribute to a new benchmark for scientific computing in collaboration with leading AI labs. You will craft original, executable research problems that current frontier models struggle to solve.
You will source materials, write prompts, and define grading criteria, calibrating against models to ensure robust evaluation during a 6-week, part-time (20+ hours/week) engagement with immediate start.
#J-18808-LjbffrAI Benchmark Architect: Computational Mathematician in san francisco at Unknown Company
This position is listed as part time and onsite.