Mercor is hiring PhD and Master’s scientists to author AI evaluation tasks for Sci Code. You will design original, executable research problems that today’s frontier models cannot solve.
Source your own material, write scientific prompts, build grading criteria, and calibrate against frontier models. Tasks ship only when strong models fail it more often than they succeed, with Docker runs and PR-based quality checks.
#J-18808-LjbffrBiochemist for AI Benchmark Tasks & Prompt Design in san francisco at Unknown Company
This position is listed as full time and onsite.