Anthropic is seeking Research Engineers to design evaluations that quantify Claude's capabilities, reasoning, safety properties, and alignment with leadership expectations.
You will build scalable eval infrastructure, run experiments across live checkpoints, and present clear results to researchers and decision-makers to push Anthropic toward leadership in well-characterized AI systems.
#J-18808-LjbffrResearch Engineer: Model Evaluations & Metrics in Remote at Unknown Company
This position is listed as full time and onsite.