Sage Care AI Agent Platform EngineerSage Care is a fast-growing, early-stage healthcare startup founded by exceptional leaders from Apple, Uber, Carbon Health and backed by top-tier venture capital (General Catalyst, Chelsea Clinton). With a strong customer pipeline, Sage Care is transforming healthcare by simplifying care navigation.Our platform makes it easier for patients to find the right doctor and helps providers focus on those who need them most through harnessing the latest AI innovations.Building on our successful collaborations with health systems across the U.S., we have expanded internationally to the MENA region. We are now partnering with health systems there to deploy our AI-powered care navigation platform.About the RoleEvery day, our AI agents handle real patient conversations for hospital systems.
Those conversations contain everything needed to make the agents better. Today, much of that learning is still manual: humans review calls, identify issues, investigate failures, and work with engineers to improve agent behavior.Your mission is to build the systems that close that loop.You will own the intelligence layer of our agent platform: the evaluation pipelines, feedback systems, and ML infrastructure that turn production conversations into measurable, continuous improvement.This role sits at the intersection of AI evaluation, ML pipelines, and production quality. You will work closely with the engineers who own the agent runtime, with the operations teams who review calls, and with product leaders who decide what good looks like for each hospital partner.What You'll DoBuild the evaluation and feedback platformDesign systems that continuously analyze production conversations, identify and cluster quality issuesBuild evaluation pipelines that measure agent performance across the dimensions that matter for patient careDevelop workflows that transform human feedback into actionable improvementsDesign mechanisms for routing issues to the appropriate AI, engineering, or operational ownersMake failures diagnosableInvestigate production failures and identify root causes across transcription, reasoning, retrieval, and orchestrationBuild tooling that helps engineers quickly understand why an agent behaved a certain wayEstablish quality metrics and reliability standards for production agentsAutomate the learning loopBuild ML pipelines that reduce the manual effort required to improve agentsWhat We're Looking ForRequired5+ years of software engineering experienceExperience building production systemsExperience working with LLMs, AI agents, or conversational AI applicationsExperience in one or more of the following:building AI evaluation platforms or frameworksdeveloping feedback systems for ML and LLM applicationscreating observability, reliability, or quality tooling for AI productsbuilding ML pipelines that improve model or agent performanceStrong backend engineering skills and systems thinkingExperience working with ambiguous problems and defining solutions from first principlesNice to HaveExperience with evaluation frameworks and model quality measurement at AI-native companiesExperience with voice AI systemsExperience designing human-in-the-loop workflowsExperience turning emerging research or new techniques into tested prototypesExperience building internal platforms used by engineering teamsExperience in healthcare or other regulated, safety-critical domainsWhat Success Looks LikeFailures are detected by systems, not discovered by peopleEvery production issue has a diagnosable root cause and a clear ownerHuman feedback measurably changes agent behavior, with the lag between the two shrinking over time