AI Evaluation EngineerLocations: Charlotte, North Carolina, United States; Denver, Colorado, United States; New York, New York, United StatesAbout Judi HealthJudi Health is an enterprise health technology company providing a comprehensive suite of solutions for employers and health plans, including:Judi Rx, a public benefit corporation delivering full-service pharmacy benefit management (PBM) solutions to self-insured employers,Judi Health™, which offers full-service health benefit management solutions to employers, TPAs, and health plans, andJudi®, the industry's leading proprietary Enterprise Health Platform (EHP), which consolidates all claim administration-related workflows in one scalable, secure platform.Together with our clients, we're rebuilding trust in healthcare in the U.S. and deploying the infrastructure we need for the care we deserve. To learn more, visit SummaryAs an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production.
This role bridges the gap between model development and real-world usage by translating ambiguous product goals into measurable quality targets.We're looking for someone to lead evaluation end-to-end — from unit and integration testing to offline, online, and statistical evaluations of probabilistic systems. What we need is someone who can design and operate robust evaluation frameworks, partner with scientists and engineers, and ensure we can confidently answer questions like: "Did this change improve or degrade quality, safety, or user outcomes?"What You'll BuildEvaluation & Quality PipelinesBuild data evaluation pipelines that collect production conversations and agent interactionsReconstruct full sessions from traces, logs, recordings, and transcriptsApply labeling and scoring using human feedback signals (surveys, sentiment, outcomes) and automated evaluators (e.g., LLM-as-judge)Continuous Quality & Safety BenchmarkingOwn weekly and on-demand automated evaluation runs against staging and productionDefine benchmarks that track accuracy, reliability, and safety-related signalsProduce trend dashboards that clearly answer: "Did this deploy change quality or risk?"Unified Evaluation FrameworkDesign and extend a standardized evaluation framework that supports multiple agent types and workflowsTranslate high-level product expectations into concrete success criteria and metricsEnsure new agents and features can be evaluated consistently with minimal frictionSelf Service Evaluation ToolingBuild APIs and internal tools so data scientists and engineers can go from "interesting scenario" to "included in the eval suite" quicklyEnable scenario curation, dataset management, and eval execution without deep infrastructure knowledgeExperiment Tracking & VisibilityProvide shared visibility into prompt, model, and agent experimentsEnable reproducibility and comparison across runs so teams can build on each other's work instead of operating in silosPosition Responsibilities:Data EngineeringBuild and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback)Implement complex data stitching and session reconstruction logicManage dataset versioning, provenance, and lifecyclePlatform & ObservabilityDevelop dashboards and monitoring tools for AI quality metricsIntegrate evaluations into CI/CD pipelines for scheduled and gated runsImplement alerting on quality and safety signals, not just infrastructure healthAI / ML Evaluation ToolingApply and extend LLM-as-judge evaluation patternsDesign metrics and scoring approaches suitable for stochastic, non-deterministic systemsUse tools like LangSmith to track runs, traces, experiments, and evaluation resultsCollaborationPartner closely with data science, engineering, and product teamsTranslate between research goals, product intent, and engineering constraintsHelp define what "good" looks like for AI behavior in productionAdvocate for strong developer experience and usability in the tools you buildResponsible for adherence to the Judi Health Code of Conduct including the reporting of non-compliance.Required Qualifications4+ years of experience in data engineering, ML engineering, or software engineeringBachelor's or Master's degree in Computer Science, Machine Learning, or a related quantitative field • Strong proficiency in Python • Experience building and maintaining production data pipelines • Strong SQL skills • Experience working with at least one cloud platform (AWS preferred)Nice to HavesPrior work on LLM or agent evaluation infrastructureFamiliarity with designing metrics for safety, reliability, or quality in AI systemsExperience with voice or call-center data (audio, transcripts, sentiment)Experience with browser automation tools (e.g., Playwright) for end-to-end evalsDeep SQL expertiseNew York, NY Salary Range: $161,600 - $200,000 USDDenver, CO Salary Range: $148,400 - $185,000 USDCharlotte, NC Salary Range: $134,800 - $168,500 USDAll employees are responsible for adherence to the Judi Health Code of Conduct including the reporting of non-compliance. This position description is designed to be flexible, allowing management the opportunity to assign or reassign duties and responsibilities as needed to best meet organizational goals.We provide equal employment opportunities to all employees and applicants for employment and prohibit discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, medical condition, genetic information, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.By submitting an application, you agree to the retention of your personal data for consideration for a future position at Judi Health.
More details about Judi Health's privacy practices can be found at