To support the evaluation of AI coding agents, the part-time Software Engineer - AI Evaluation will create realistic developer environments, design tasks, and write tests to assess agent solutions in a remote capacity. Key responsibilities Build realistic developer environments that simulate a virtual company with a codebase and context Design tasks and evaluation criteria to assess AI agent performance and ensure tasks are solvable Write and iterate on tests that verify agent solutions based on QA feedback and analysis of failures Required qualifications 5+ years of experience in software development Proficiency in core technologies: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience in writing functional and integration tests English proficiency at B2 level or higher
Software Engineer - AI Evaluation in workfromhome at Unknown Company
This position is listed as part time and able to be worked remotely.