- Own the evaluation layer for Hilbert’s production agents, including harnesses, metrics, golden datasets, regression gates, and human-in-the-loop review
- Architect and implement agent workflows using LangChain, LangGraph, or equivalent frameworks
- Build state memory, routing, tool registries, and recovery paths
- Own systems from experimentation through production
- Operate production systems through tracing, monitoring, latency management, cost-per-task budgeting, and on-call
- Diagnose production failures and turn them into durable fixes and tests
- Expand agent capabilities across retrieval, orchestration, and execution
- Set technical standards, review designs, and define reusable patterns
- Collaborate with the founding team and cross-functional partners
- Communicate technical decisions, tradeoffs, and progress clearly
- Make pragmatic engineering decisions, ship, learn, and iterate
- Build intelligent retrieval combining RAG, graph-based retrieval, and other approaches
- Develop robust agentic workflows that handle edge cases, missing data, escalation, and human-in-the-loop checkpoints
- Integrate agents with external platforms and execute real-world workflows
Requirements
- 6+ years of production software engineering experience
- 2+ years building LLM or agent systems used by real users in production
- Experience with APIs, services, data infrastructure, tests, CI/CD, and on-call
- Hands-on experience with LangChain, LangGraph, or equivalent agent/orchestration frameworks
- Experience with agent architectures, state memory, routing, tool registries, recovery paths, and multi-step inference
- Ability to design evaluation harnesses, metrics, golden datasets, regression gates, and human-in-the-loop review
- Ability to diagnose hallucination, tool misuse, retrieval misses, and silent degradation
- Strong knowledge of retrieval-augmented generation, hybrid and graph retrieval, chunking, embeddings, ranking, and grounding
- Experience with LLM observability tools such as Langfuse or OpenTelemetry
- Knowledge of cost and latency optimization
- Knowledge of MCP, tool-calling frameworks, structured output, and constrained decoding
- Clear technical communication and ability to explain architecture tradeoffs
- Ability to take ownership and work effectively in ambiguity and at startup speed
- Willingness to commute to the San Francisco office for a hybrid work schedule
- Authorization to work in the United States without visa sponsorship
Core Competencies
Demonstrates expertise in architecting and implementing agent workflows using LangChain or LangGraph, with a strong focus on production software engineering and LLM systems. Capable of optimizing performance through effective monitoring, diagnostics, and technical communication.
Highest-signal resume keywords
- LangChain Framework
- LangGraph Framework
- Production Software Engineering
- LLM Systems Development
- Retrieval-Augmented Generation
ATS Optimization Keywords
Hard Skills
- APIs
- Data Infrastructure
- CI/CD
- Agent Architectures
- State Memory
- Routing
- Tool Registries
- Recovery Paths
- Evaluation Harnesses
- Metrics
Soft Skills
- Clear Technical Communication
- Ownership
- Ability to Work in Ambiguity
Industry Keywords
- Production Systems
- Human-in-the-Loop Review
- Cost Optimization
- Latency Management
- Hybrid Retrieval
Tools & Technologies
- Langfuse
- OpenTelemetry
Senior AI Engineer – Core in san francisco at Unknown Company
This position is listed as full time and hybrid.