Research Engineer – Benchmarking, Evals & Failure AnalysisLocation: San Francisco Company Stage: Late-Stage / Series C (AI / Applied ML) Office Type: Onsite (5 Days a Week) Salary: $130,000 – $400,000 + EquityThis fast-growing AI company is operating at the forefront of applied machine learning and labor transformation. By partnering with leading AI labs and enterprises, they are building systems that combine human expertise with cutting-edge AI to improve model performance and unlock new categories of work. With strong revenue, scale, and backing from top-tier investors, they are shaping how frontier models are trained, evaluated, and deployed in real-world environments.What You Will DoDesign and implement benchmarking systems to evaluate model capabilities such as tool use, reasoning, and agent behaviorBuild and operate end-to-end evaluation pipelines, including scoring systems, dashboards, and reporting infrastructureConduct systematic failure analysis on model outputs, identifying key failure modes and translating them into actionable improvementsDevelop rubrics, evaluators, and scoring frameworks that balance rigor with scalability (human + automated evaluation)Partner with research and applied AI teams to align evaluation systems with training and product goalsAnalyze data quality and performance trends to inform model training, data generation, and post-training strategiesOwn evaluation and benchmarking systems in a fast-paced, high-iteration environmentIdeal BackgroundStrong applied AI or ML engineering experience, particularly in model evaluation, benchmarking, or failure analysisHands-on experience building or running LLM evaluation systems, benchmarks, or experimentation pipelinesStrong coding ability (Python or similar) with experience building production-quality systemsSolid understanding of data structures, algorithms, and backend systemsExperience working with APIs, databases (SQL/NoSQL), and cloud infrastructureAbility to reason deeply about model behavior, experimental results, and system performanceComfortable operating in ambiguous, high-ownership environments with rapid iteration cyclesPreferredExperience working on post-training, RL, or evaluation teams at AI labs or AI-first companiesFamiliarity with LLM evaluation techniques, benchmarking frameworks, or agent evaluation systemsExperience with synthetic data generation, rubric design, or reward modeling workflowsPublications or research experience in ML, especially in evaluation or benchmarkingExposure to large-scale experimentation systems or model performance tracking infrastructureCompensation and BenefitsCompetitive base salary ($130K – $400K) + meaningful equityRelocation and housing support availableMonthly meal stipend and premium wellness perks (e.g., fitness membership)Comprehensive health insuranceOpportunity to work directly with frontier AI labs and influence model development at the cutting edgeThis is a high-impact role at the intersection of engineering and applied AI research, ideal for candidates excited about defining how next-generation models are evaluated, improved, and deployed at scale.