Unknown Company

AI Evaluation Engineer

new york, ny • Posted 3 days ago
Remote Full Time General

Job TitleAI Evaluation EngineerLocationHybrid / RemoteEmployment TypeFull-timeJob SummaryWe are seeking an AI Evaluation Engineer to design, implement, and maintain evaluation frameworks for AI and machine learning systems, with a focus on Large Language Models (LLMs) and generative AI applications. The ideal candidate will develop robust evaluation methodologies, assess model performance across quality and safety dimensions, and collaborate with AI engineers, data scientists, and product teams to improve AI system reliability and user experience.Key ResponsibilitiesDesign and implement evaluation frameworks for AI, machine learning, and generative AI systems.Develop automated and manual evaluation pipelines to assess model quality and performance.Define evaluation metrics for accuracy, relevance, factuality, consistency, completeness, latency, and user satisfaction.Create benchmark datasets, test suites, and evaluation scenarios for AI models.Evaluate LLMs and AI applications for hallucinations, bias, toxicity, fairness, robustness, and safety.Measure Retrieval-Augmented Generation (RAG) performance, including retrieval quality and response grounding.Conduct A/B testing and comparative evaluations of models, prompts, and AI workflows.Analyze evaluation results and provide actionable recommendations for model improvement.Collaborate with AI engineers, data scientists, product managers, and QA teams throughout the AI development lifecycle.Monitor production AI systems and identify performance degradation, data drift, and model drift.Document evaluation methodologies, findings, and best practices.Ensure compliance with organizational AI governance, privacy, security, and responsible AI policies.Required QualificationsBachelor's degree in Computer Science, Artificial Intelligence, Data Science, Statistics, Mathematics, Software Engineering, or a related field.3–5+ years of experience in AI, machine learning, software engineering, data science, or model evaluation.Strong understanding of machine learning concepts and evaluation methodologies.Experience evaluating AI or generative AI applications.Proficiency in Python and SQL.Experience with REST APIs and cloud-based applications.Preferred QualificationsMaster's degree in AI, Machine Learning, Data Science, Computer Science, or a related field.Experience evaluating Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems.Knowledge of Responsible AI, AI safety, and AI governance principles.Experience with MLOps and model lifecycle management.AI or cloud certifications (AWS, Azure, Google Cloud).Technical SkillsPythonSQLMachine learning fundamentalsLLM evaluation methodologiesPrompt engineeringRAG evaluationAI benchmarking techniquesStatistical analysisData visualization (Power BI, Tableau, Matplotlib)Git and CI/CDREST APIsJSONAI evaluation frameworks (DeepEval, Ragas, LangSmith, Promptfoo)MLflowDocker and KubernetesCloud platforms (AWS, Azure, Google Cloud)Soft SkillsAnalytical thinkingCritical thinkingProblem-solvingStrong communication skillsTechnical documentationCollaborationAttention to detailTime managementContinuous learningPreferred ExperienceGenerative AI applicationsConversational AI and chatbotsEnterprise AI platformsRetrieval-Augmented Generation (RAG)Machine learning model validationAI-powered SaaS applicationsRegulated industries such as healthcare, finance, or insuranceSuccess MetricsAI evaluation coverageBenchmark quality and completenessImprovement in model accuracy and reliabilityReduction in hallucinations and unsafe responsesEvaluation pipeline automationDetection of model regressions before releaseProduction AI performance and stabilityStakeholder satisfactionCompliance with Responsible AI and governance standardsNice-to-Have SkillsExplainable AI (XAI)Model monitoring and observability toolsSynthetic data generationVector databasesLangChain or similar AI orchestration frameworksData annotation toolsExperiment tracking platformsA/B testing methodologiesAI red teamingFamiliarity with standards such as NIST AI Risk Management Framework (AI RMF) or ISO/IEC 42001

Back to Job Search