\"Did it actually get better?\" If you want to be the one who answers that with numbers the whole team trusts, read on.
A San Francisco AI lab is building consumer-grade agents that run privately on phones, laptops and other devices. They've held the top spot on the industry's leading mobile-agent benchmark since late 2025, work with a leading mobile chipmaker and a global device maker, and are backed by top-tier VCs.
Without a serious evals function, releases are based on gut feeling. With one, the team knows whether each training run or prompt tweak helped, and trusts the number. You build that function: evals covering what the models can do, how the agents behave, how they perform on device, and what users actually experience. And you decide what \"shipped\" means, then hold that line when deadlines push.
What you'll own
- The eval suites every model and agent release has to pass: capabilities, behavior, regressions, plus rubrics scored by people for everything automated checks can't see
- Dashboards and tools that speed up research iterations and make leadership decisions simple
- The definition of \"ready to ship\", and the evidence behind it
- The bridge to research (measuring what they really care about), product engineers (tracking how real users behave on real hardware) and the partnerships team (translating \"it improved\" into commitments a hardware partner can verify)
Your first 90 days
- Day 30: you've reviewed every existing eval, written up what's useful, what's noise and what's missing, and closed the biggest gap
- Day 60: you've launched a new area of evaluation
- Day 90: releases go out against your standard, and you've stopped at least one regression from shipping
What you bring
- Eval harnesses for systems that never give the same answer twice: agents, tool calling, multi-step and long-horizon tasks, behavior across languages
- Strong engineering and tooling skills: dashboards, instrumentation, quick experiment cycles
- You design human-scored rubrics, and when a metric gets gamed you call it out without wrecking team morale
- Real urgency. You hold the standard under deadline pressure and say \"not ready\" when it isn't.
- Evals built in a real setting: a data or eval company, or a frontier lab's eval team
- Hypergrowth, founder or early-hire experience is a plus; relevant work beats a famous name
Bonus points
- Agentic, tool-use or on-device evals
- Consumer AI or QA at device-maker scale
What's in it for you
- $200k–$250k base + equity
- You define what \"better\" means at a frontier on-device lab
- Relocation and immigration support
Good to know
- Full-time, in person in San Francisco, 9-9-6
- Process: intro → quick talk with the founder → technical conversation (no live coding, no puzzles) → half-day onsite → work trial → offer
Research Engineer – Evals in san francisco at Unknown Company
This position is listed as full time and onsite.