Unknown Company

Remote | QA Test Engineer — $55–$85/hour

new york, ny • Posted 4 days ago
Remote Contract General

Specialised Full-Time Consulting Opportunity for QA and Test EngineersWe are sharing a specialised full-time consulting opportunity for experienced QA and test engineers with strong expertise in test-case design, end-to-end debugging, quality assurance, Python, and complex technical evaluation workflows.This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will review complex multi-step tasks, test reference solutions, identify ambiguity and grading gaps, debug technical environments, and develop repeatable quality processes that keep benchmark results accurate and trustworthy.Key ResponsibilitiesCreate comprehensive test cases confirming that benchmark tasks function as intendedDesign positive, negative, boundary, and edge-case testsValidate task requirements, expected outputs, reference solutions, and grading logicIdentify scenarios that may produce incorrect or misleading evaluation resultsEnsure tests measure the intended technical capability accuratelyReview complex multi-step tasks and reference solutions before finalisationIdentify ambiguous instructions, inconsistent requirements, missing assumptions, and incomplete acceptance criteriaRun tasks independently to confirm reproducibility and expected behaviourAssess whether grading standards are clear, fair, and technically defensibleProvide actionable feedback to task authors and researchersInvestigate failures across Python scripts, test harnesses, repositories, and task environmentsDiagnose unexpected behaviour within unfamiliar codebasesReproduce reported issues and isolate their underlying causesCorrect or document environment, dependency, logic, and validation problemsUse Git-based workflows to support structured review and collaborationDevelop practical checklists and repeatable review procedures for benchmark qualityImprove consistency across task validation, testing, and approval workflowsDocument findings clearly so authors can resolve issues efficientlyTrack recurring defects and recommend preventive quality measuresCollaborate closely with researchers, task authors, and other technical reviewersExamine AI agent runs for unintended shortcuts, loopholes, and grading weaknessesIdentify cases where models can receive credit without completing the intended reasoning or technical workTest whether benchmark tasks remain robust across alternative approachesStrengthen evaluation criteria to maintain reliable and meaningful benchmark scoresDistinguish valid solution diversity from unintended task exploitationIdeal ProfileStrong candidates may have:At least 1 year of experience in test engineering, quality assurance, software engineering, research engineering, or a related technical roleDemonstrated experience designing test cases and quality-review processesStrong end-to-end debugging skills across complex technical systemsWorking proficiency in Python and GitComfort navigating unfamiliar codebases, repositories, and execution environmentsExceptional attention to detail and strong written documentation habitsAbility to identify ambiguity, edge cases, hidden assumptions, and quality gapsCapacity to work independently through open-ended technical problemsReliable availability for approximately 35 hours per weekEducational BackgroundA master's degree or PhD in a STEM field is highly relevantEquivalent practical experience in an engineering-intensive or research-intensive domain may also be consideredAcademic or professional experience involving computer science, software engineering, machine learning, mathematics, statistics, or a related technical field may strengthen an applicationTechnical research, open-source contributions, testing projects, or substantial engineering work may also be valuableNice to HaveExperience with AI training, model evaluation, or quality review of AI-generated workFamiliarity with agentic systems and multi-step AI benchmarksBackground testing machine learning, research, or data-processing workflowsExperience developing automated test suites or validation scriptsFamiliarity with CI/CD systems, test harnesses, containers, or reproducible environmentsExperience reviewing reference solutions, grading logic, or technical rubricsKnowledge of adversarial testing, failure-mode analysis, or benchmark designPrior collaboration with AI research or evaluation teamsWhy This OpportunityServe as the quality backbone for advanced agentic AI benchmarksEnsure complex evaluation tasks are accurate, reproducible, and resistant to shortcutsApply testing and debugging expertise to technically challenging AI research workflowsWork closely with researchers and task authors on benchmark improvementHelp protect the reliability of evaluation results for frontier AI systemsParticipate in a structured full-time remote role with competitive hourly compensationContract DetailsFull-time W-2 contingent employment opportunityFully remote within the United StatesExpected commitment of approximately 35 hours per weekCompetitive rates between $55–$85 per hour depending on expertise and project scopeIndividual task reviews may require one to two days of focused technical workWork may include test design, benchmark review, Python debugging, quality-process development, and shortcut detectionClose collaboration with research and task-development teamsEngagement scope and duration may evolve according to project requirements and performanceAbout the PlatformThis opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy:

Back to Job Search