1. Role Overview
Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real‑world data science work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task‑specific grading criteria and scoring completed work samples with rigorous, well‑reasoned written justifications.
2. Key Responsibilities
- Design precise, task‑specific grading criteria for real‑world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations)
- Score AI-generated and human work samples against those criteria, with detailed written justifications for every score
- Apply consistent, evidence‑based judgment so that scores are reproducible and defensible
- Incorporate structured feedback from senior reviewers and iterate quickly on your work
3. Ideal Qualifications
- 5+ years of professional data science experience in industry
- Background in business operations, product, or growth data science at top‑tier technology companies
- Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders
- Exceptionally strong written communication
- Detail‑oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers
- Prior experience with AI training, evaluation, or human‑data projects is a strong plus
4. Application Process
- Submit your resume or relevant technical background to get started
- Qualified applicants may be asked to complete a brief technical assessment or submit additional information