Unknown Company

Research Scientist – RL Post-Training for Agents

san francisco, ca • Posted Today
Onsite Full Time Science

Want your RL research to land in agents that run for days in the real world, not in a paper appendix?

A brand-new AI lab in San Francisco is building autonomous agents that pursue complex goals over very long horizons. The founding team comes from leading frontier AI labs, autonomous-driving and robotics AI, big-tech research and a top quant firm. This is a research seat that builds real systems: you own ambitious bets from the first hypothesis and dataset all the way to a deployed capability.

They’re hiring 4 research scientists.

What you’ll own

  • Using RL, and whatever else works, to post-train LLM-based and multimodal agents
  • Long-horizon capabilities: planning, memory, error recovery, proactivity, self-improvement, and knowing when to escape to a human
  • Building the environments where agents use computers and tools, plus the data pipelines, benchmarks and evals around them
  • Ways for agents to grasp what a user wants and stay true to it over long runs
  • Careful experiment design, and taking what works all the way into production
  • A real say in the research agenda from day one

What you bring

  • Exceptional ML research and engineering skills
  • Depth in at least one of these: RL, LLM post-training, reasoning, agents, computer use, long-horizon systems, memory and context, evals, or how humans and agents work together
  • Good judgment on what to test and when to stop. You build complete systems and move easily between ideas, large-scale experiments and production.
  • You’ve led something significant: a model, an agent system, a benchmark, a paper, an open-source project or a big research bet
  • Roughly 3–6 years in; frontier-lab experience valued, exceptional outliers and senior leads welcome
  • High agency and comfort with uncertain directions

Bonus points

  • Hands-on RL post-training or computer-use work
  • A strong publication, open-source or benchmark record
  • Experience building environments and eval harnesses
  • A spike: olympiad (IOI/IMO), quant, or world-class competitive achievement

What’s in it for you

  • $250k–$500k base + 1–5% equity
  • Your own research bets, end to end, at a lab where results ship
  • Visa sponsorship available; if you’re outside the US, expect to go via an O-1

Good to know

  • Full-time, in person in San Francisco, 9-9-6
  • Process: informal talk with a founder → technical deep-dive on your research → paid 2–3 day work trial in person → offer
#J-18808-Ljbffr

Research Scientist – RL Post-Training for Agents in san francisco at Unknown Company

This position is listed as full time and onsite.

Back to Job Search