Unknown Company

Research Engineer, Benchmarks

san francisco, ca • Posted 3 days ago
Remote Full Time Architecture and Engineering Occupations
About the Role
Join a small, technically elite team - including International Olympiad medalists and published AI researchers - building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. As a Research Engineer, Benchmarks , you'll own the design and implementation of evaluations that frontier labs and enterprise customers trust. This role is central to ensuring our benchmarks are rigorous, credible, and tightly aligned with real-world agent performance.
This is an on-site role based in San Francisco, CA . Visa sponsorship is available.
What You'll Do
  • Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
  • Partner with subject-matter experts to define realistic workflows and tasks for domain-specific evaluations.
  • Build reliable infrastructure to run models and agents against benchmark tasks at scale.
  • Develop metrics and analyses that measure benchmark difficulty, reliability, and failure modes.
  • Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.
  • Write clear documentation and benchmark reports that make results legible and credible to technical audiences.
What We're Looking For
Required
  • 2-4 years of experience in software engineering, ML engineering, or research roles.
  • Strong proficiency in Python, Docker, and Linux environments.
  • Experience building environments, evaluations, or benchmarks for AI systems.
  • Published research or technical writing on topics such as public benchmarks, model failure modes, or evaluation methodology.
  • Deep understanding of what makes a benchmark realistic, reliable, and practically useful.
  • Curiosity and genuine ability to understand how real-world workflows operate across diverse domains.
  • Strong attention to detail - a habit of spotting subtle inconsistencies and edge cases in task design.
  • Ability to reason from first principles about task design, scoring, and failure modes.
  • Comfort thriving in unstructured problem spaces and working independently in fast-paced, early-stage environments.
  • Excellent communication skills for collaborating across time zones and with technical teams.
Compensation & Benefits
  • Salary: $150,000 - $250,000 USD annually, depending on experience.
  • Visa sponsorship available.
  • Opportunity for significant early-stage equity and career growth within a high-impact, research-driven team.
Location
This is a full-time, on-site position in San Francisco, CA . Candidates must be willing and able to work in-office.
Back to Job Search