What You’ll Do
- Set up end-to-end evals to measure & improve agent performance.
- Experiment with new agentic techniques (e.g. multi-agent systems, reasoning-from-feedback, RFT, etc)
- Build lightweight tools, servers, and orchestration layers (e.g. MCP servers) that enable agents to operate reliably in production
- Stay on top of emerging research and blogs on LLM/AI agents and bring ideas into production experiments.
What We’re Looking For
- Amazing ability to speak with LLMs - Occam's razor in prompting
- Strong experience with Python
- 6+ years building in ML/AI
- Clear communicator - both in person and in writing
- Bonus: background in B2B SaaS and 0-1 experience
- Above all: drive, grit, and ownership . If you excel here, other requirements are secondary
Note that this is not a model-training role - you’ll be building orchestration and reasoning systems on top of existing LLMs (think Claude-Code over Claude-Model).
You can look forward to the following benefits:
- Fully covered, best-in-class health, dental, and vision benefits
- Competitive Compensation, meaningful stock options, company 401(k)
- Unlimited PTO
- Outstanding in-office culture in the heart of San Francisco
- Lunch and dinner onsite
- Team events, such as happy hours and off-sites
- Pre-tax Commuter benefits