About the Role
We are looking for an AI Research Evaluator to design difficult, verifiable research problems for frontier AI browsing agents.
This is an investigative research rolenot a conventional subject-matter expert or content-writing position. You will begin with a verifiable fact, work backwards to create a challenging research question, and build a complete evidence trail proving the answer.
Key Responsibilities
- Design complex, natural-language research questions with objectively verifiable answers.
- Develop interconnected clues involving people, dates, places, organisations, works, events, records, and quantities.
- Research government databases, archives, registries, institutional websites, PDFs, and primary records.
- Validate information using reliable and independently checkable sources.
- Document exact pages, tables, sections, URLs, and supporting evidence.
- Record obvious searches attempted and explain why they did or did not reveal the answer.
- Prepare structured research outputs for AI benchmark and evaluation datasets.
- Identify weaknesses, search limitations, and failure points in AI-generated research.
Required Qualifications
- 3+ years of professional experience in open-web, investigative, archival, factual, or records-based research.
- Demonstrated ability to locate and verify primary sources.
- Strong written English communication skills.
- Excellent attention to detail and source precision.
- Ability to research unfamiliar topics independently.
- Experience with LLM evaluation, AI red-teaming, benchmark creation, or fact-checking.
- Comfort creating detailed, structured evidence documentation.
- Familiarity with JSON or structured data formats is desirable.
Contract Details
- Remote contractor assignment
- 40 hours per week
- Minimum 4 hours of PST overlap required
- Contract duration: 8 weeks
- No medical benefits or paid leave under the contractor arrangement
Web Research Specialist in Remote at Unknown Company
This position is listed as contract and able to be worked remotely.