Role Overview
Design and execute evaluations that measure code quality and system behavior across major open-source languages and repositories, build reliable test suites, analyze real-world usage patterns, and deliver clear findings to research teams to drive improvements at scale.
Key Responsibilities- Design and oversee the creation of evaluations for a wide range of coding tasks across multiple languages, including JavaScript, TypeScript, Python, Java, and C.
- Develop test cases and test suites that accurately assess system performance in diverse engineering scenarios.
- Analyze system behavior on real-world user use cases to identify strengths, failure modes, and areas for improvement.
- Document and communicate evaluation results clearly to the research team to inform development and optimization.
- Strong public open-source presence, for example on GitHub or a similar platform, with frequent, high-quality contributions to leading projects in the last 12 months.
- Expertise in one or more of the following languages: Python, Java, C, JavaScript, or TypeScript.
- Deep familiarity with widely used libraries, frameworks, and tools in your language or languages of choice.
- Excellent understanding of software architecture, performance tuning, and scalable code patterns.
- Experience collaborating in distributed, asynchronous teams and contributing to open-source governance workflows.
- Comfortable using Git and CI/CD systems.
- Ability to identify opportunities for contribution and to execute improvements with minimal oversight.
Remote role, engaged on an hourly basis.
Compensation100 - 150 hourly
EligibilityCandidates must demonstrate recent, high-quality open-source contributions as part of their qualifications. No other specific work authorization or sponsorship details are provided.