Role Overview
Define what excellent data science work looks like, and judge whether AI systems and humans meet that standard. Rather than creating analyses or models yourself, you will design task-specific grading criteria for real-world data science outputs and score completed work samples with rigorous, evidence-based written justifications. The work supports an evaluation of how well AI systems perform practical data science tasks.
Key Responsibilities- Design precise, task-specific grading rubrics for deliverables such as analyses, models, dashboards, experiment readouts, and written recommendations
- Evaluate AI-generated and human work samples against those criteria, assigning scores and providing detailed written justification for every score
- Apply consistent, evidence-based judgment so scores are reproducible and defensible
- Receive and incorporate structured feedback from senior reviewers, iterating quickly on grading criteria and scoring approach
- At least 5 years of professional data science experience in industry
- Experience in business operations, product, or growth data science at top-tier technology companies preferred
- Deep fluency in experiment design and A/B testing, metric definition, SQL and Python analysis, and presenting findings to executive stakeholders
- Exceptionally strong written communication skills, able to produce clear, well-reasoned justifications for scores
- Detail-oriented and consistent, comfortable having your judgment reviewed and calibrated against peers
- Prior experience with AI training, evaluation, or human-data projects is a strong plus
- Location: Remote
- Employment type: Hourly
- $120 to $170 per hour
To apply, submit your resume or a summary of relevant technical experience. Qualified candidates may be asked to complete a brief technical assessment or to provide additional information for evaluation.
EligibilityNo additional work-authorization or eligibility details were specified. The role is fully remote.