- Lead the development of credible, repeatable evaluations for live AI agents.
- Own the technology stack supporting the evaluation platform, from APIs to orchestration.
Key Responsibilities:
- Design reusable evaluation primitives applicable across software categories.
- Partner with data science teams to develop proprietary benchmarks.
- Promote the effective use of evaluations across agent-focused products.
Requirements:
- 10+ years of backend or full-stack development experience.
- 2+ years of engineering management experience.
- Hands-on experience evaluating AI agents and using frontier models.
- Total earnings of approximately $260,000-$320,000.
- Unlimited PTO, equity, and comprehensive parental leave.
Partner Company
This company focuses on building credible, scalable evaluations of AI agents from software vendors. It operates as a fully remote, inclusive team with a flexible culture and a focus on professional growth.
#J-18808-LjbffrSoftware Engineering Director, Agentic Evaluations (Fully Remote) in workfromhome at Unknown Company
This position is listed as full time and able to be worked remotely.