Unknown Company

AI Evaluations Engineer, US Decision Intelligence

cupertino, ca • Posted 1 weeks ago
Onsite Contract Software Architecture & Engineering

AI Evaluations Engineer, US Decision Intelligence

Cupertino, California, United States Machine Learning and AI

Imagine what you could do here. At Apple, new ideas have a way of becoming outstanding products, services, and customer experiences very quickly. Bring passion and dedication to your job, and there's no telling what you could accomplish.Apple’s Sales organization generates the revenue needed to fuel our ongoing development of products and services. This, in turn, enriches the lives of hundreds of millions of people around the world. We are, in many ways, the face of Apple to our largest customers.Apple's US Decision Intelligence (DI) team is looking for a talented individual who is passionate about crafting, implementing, and operating AI solutions that have a direct and measurable impact on Apple Sales and its customers.

Description

We’re seeking a visionary AI Evaluations Engineer to own the end-to-end evaluation pipeline for our AI products and agentic workflows. This role will focus on implementing and maintaining evaluation frameworks, instrumentation, and workflows that help us understand how well our AI systems perform, where they fail, and how they improve over time. You own the evaluation gate and the standards.This role will operate in both capacities, to augment existing AI roadmap, as well as innovate and trailblaze new frontier-technology projects, crafting AI experiences that reduce time to insight and catalyze decision making.

Responsibilities

  • Architect a comprehensive Evals framework to trace at every layer, from agent responses breakdown, to skill level, and tool calling with the main objective to optimize for accuracy and performance improving latency and running AB tests to find the best recommendation across the system.
  • Build and operate AI evaluation workflows that measure the quality of LLM outputs across chat, summarization, recommendations, and agentic actions.
  • Implement rubric-based evals to score outputs for correctness, relevance, grounding, and consistency.
  • Move beyond LLM-as-a-judge to agent-as-a-judge, including harness-as-a-judge patterns applied against real traces in a sandbox.
  • Instrument LLM and agent workflows to capture traces, prompts and responses, metadata, and user feedback.
  • Own the platform-wide eval gate: define the pass/fail contract every track ships against, and hold a release when it isn't met.
  • Help define agent-specific evaluations (task completion, tool correctness, error recovery).
  • Partner with AI engineers and AI platform teams to translate product requirements into evaluation criteria.
  • Arbitrate eval disputes with track leads and own the standard that geo eval engineers implement agains.
  • Define rerun policy and variance thresholds.
  • Collaborate with the Evals team in India to maximize global impact.
  • Contribute to system design for observability, retries, and logging.

Minimum Qualifications

  • 5+ years of experience in data and AI-related fields such as AI engineering, software development, ML engineering, data science, or QA roles.
  • Eagerness and ability to learn new skills and solve dynamic problems in an encouraging and expansive environment.
  • Hands-on experience with AI evaluation techniques, such as Golden datasets, LLM-as-a-Judge, or rubric-based scoring.
  • Experience with different LLM ecosystems (OpenAI, Anthropic, Gemini, etc.), RAG pipelines, vector databases (e.g., Pinecone, FAISS, Milvus, PostgreSQL).
  • Proficiency in SQL and experience with at least one major data analytics platform, such as Hadoop, Spark, or Snowflake.
  • Experience with CI/CD or release validation workflows.
  • Experience working with data science teams on insights generation leveraging LLMs.
  • Strong time management skills with the ability to collaborate across multiple teams.
  • Able to balance competing priorities, long-term projects, and ad hoc requirements.
  • Ability to work in a fast-paced, dynamic, constantly evolving business environment.
  • Hands-on experience with Langfuse or similar tools for LLM observability.
  • Comfortable working with product/domain experts to translate fuzzy correctness criteria into measurable rubrics or metrics.
  • B.S. degree in Computer Science/Engineering, or equivalent work experience

Preferred Qualifications

  • Sound communication skills - expert at messaging domain and technical content, at a level appropriate for the audience. Strong ability to gain trust with stakeholders and senior leadership.
  • Familiarity with embeddings, retrieval algorithms, agents, and data modeling for vector and graph databases.
  • Other complementary technologies for distributed systems architecture and asynchronous messaging, agent communication, and caching like RabbitMQ, Redis, and Valkey are preferred.
  • Experience working across global teams to ensure alignment of product development.
  • Applied knowledge of GenAI and RAG strategies, microservices, recommendation systems, and context engineering.
  • Working knowledge of agent evaluation concepts like trajectory vs. end-to-end vs. component-level evaluation, tool-call correctness.
  • Advanced degree (MS or Ph.D.) in Economics, Electrical Engineering, Statistics, Data Science, or a similar quantitative field is preferred.

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $277,600, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

#J-18808-Ljbffr

AI Evaluations Engineer, US Decision Intelligence in cupertino at Unknown Company

This position is listed as contract and onsite.

Back to Job Search