Mercor is seeking researchers to author AI evaluation tasks and original, executable problems for frontier models. You will source material, write prompts, and define grading criteria across subdomains with a focus on two areas in mathematics. Engagement is six weeks, part-time, with start date immediate.
Experience with Python or R for scientific computing and familiarity with Git/GitHub and Docker workflows are required to ensure reproducible runs and automated quality checks.
#J-18808-LjbffrAI Benchmark Architect for Scientific Computing in san francisco at Unknown Company
This position is listed as part time and onsite.