You will own AI features end to end, from the first evaluation harness to the system running in a client’s cloud. We build copilots, agents, and RAG systems for regulated industries, so correctness, observability, and graceful failure matter as much as the happy path.
This is a hands-on engineering role on a small, senior team. You will write the hard parts yourself, set the patterns others follow, and stay close to the decisions that determine whether a build ships or stalls.
What you'll do
- Architect and build LLM applications: agents, tool use, retrieval, and evaluation pipelines.
- Stand up evals and observability so quality is measured continuously, not guessed at.
- Harden systems for production: latency, cost, prompt-injection, data leakage, and fallbacks.
- Integrate with messy client data and existing systems without breaking their controls.
- Set engineering patterns and review work across the team.
What we're looking for
- 5+ years building production software, with recent hands‑on LLM/ML application work.
- Strong Python and TypeScript; comfortable across the full stack of an AI system.
- Real experience with evaluation, retrieval, and the failure modes of LLM systems.
- A bias for shipping something narrow that works over something broad that demos.
- Clear written communication; you can explain trade‑offs to engineers and executives.
Nice to have
- Experience deploying into client VPCs / on‑prem with SSO and data‑residency constraints.
- Background in a regulated sector (finance, healthcare, legal, real estate).
- Open‑source contributions or published technical writing.