24-MAG LLC is seeking an experienced QA/test engineer to design robust benchmark test cases, review complex tasks, and debug Python environments for frontier AI evaluation work. The role is fully remote within the United States and requires strong attention to detail and independent problem-solving.
Your work will help ensure accurate, reproducible, and defensible evaluation results across multi-step AI tasks, with close collaboration to researchers and task authors.
#J-18808-LjbffrRemote AI Benchmark QA Engineer in workfromhome at Unknown Company
This position is listed as full time and able to be worked remotely.