Unknown Company

QA / Evaluation Lead

washington, dc • Posted 4 days ago
Hybrid Full Time General

QA / Evaluation LeadHybrid - Washington D.CInnodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale.

We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.About the Program:Innodata's Federal Practice builds the trusted data layer for critical infrastructure Trust & Safety work. Partnering with a leading systems integrator, we're delivering a modern, governed data services platform in a secure federal (IL4) environment.

Over an intensive 20-week phase, you'll help stand up a data services storefront, a DataCard governance framework, synthetic data integration, and Databricks write-back capabilities.About the Role:As the QA/Evaluation Lead, you'll own quality and evaluation across the platform. You'll design the evaluation framework that measures whether our data services and outputs meet the bar, build repeatable test and validation processes, and give the team an objective read on readiness at each milestone. Partnering with the Delivery Owner and engineering leads, you'll turn quality from an afterthought into a measurable, demonstrable strength.

It's a role for someone who thinks rigorously about evaluation and takes pride in evidence-backed quality.Key Responsibilities:Design and own the inter-annotator agreement (IAA) methodology for the Phase 1 demonstration corpus — metric selection (Cohen's kappa, Fleiss, Krippendorff's alpha), sampling design, adjudication workflow, and agreement thresholdsDefine evaluation framework architecture: test and evaluation plans, IAA targets, drift detection gates, and model performance metrics per SOW Section 2.9Configure and operate sampling-based quality control across the self-service and white-glove annotation paths during Phase D corpus productionDesign and implement confidence-threshold escalation routing from automated annotation to senior-annotator adjudicationValidate quality scoring and IAA computation within the Innodata data layerSupport AI Solutions Engineer on evaluation design for SAM 2 and Frontier model API validation — define what 'good enough' looks like quantitativelyProduce evaluation framework documentation for the Phase 1 NPP closeout package, including per-DataCard documentation with the SAMust-Have Qualifications:Bachelor's degree in Statistics, Data Science, Computer Science, or related quantitative field required; Master's degree preferred. Equivalent experience may substitute for degree on a 2-for-1 basis.6+ years total professional experience, 4+ years in data quality, evaluation methodology, or QA on AI/ML programsIAA methodology expertise — Cohen's kappa, Fleiss' kappa, Krippendorff's alpha: hands-on, not theoreticalEvaluation framework design for AI/ML training data programsQC process design: sampling methodology, escalation workflows, adjudication protocolsPython for QC tooling, metric computation, and statistical analysisActive Secret clearance with TS/SCI eligibilityNice-to-Have Qualifications:Prior DoD or IC data quality program experienceCVAT or equivalent annotation platform QC workflow configurationDrift detection and model monitoring methodologyExperience with FMV / video annotation quality standardsThe expected hourly salary range for this position is $45 to $50 p/hour, based on experience, skills, and qualifications.

Back to Job Search