Mindrift is creating a dataset to evaluate AI coding agents by simulating real-world developer tasks. You will guide and evaluate the agent-written code, not write code from scratch, and help design tests that challenge Frontier models in coding scenarios.
This project-based role offers flexible scheduling and compensation up to $50/hr, depending on level and pace, with roughly 20 hours per task.
#J-18808-LjbffrSenior AI Agent Evaluation Engineer in Remote at Unknown Company
This position is listed as full time and onsite.