Obsidian is seeking a candid evaluator to assess the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train frontier AI labs. You will review repository-level tasks, reference patches, test harnesses, and grading integrity, and provide rubric-based written feedback.
Responsibilities include identifying gaps in tooling, ensuring task reproducibility across environments, and delivering concrete, rubric-based assessments to guide model training and
#J-18808-LjbffrAI Benchmark Quality Engineer for Task Evaluation in san francisco at Unknown Company
This position is listed as full time and onsite.