Own the production environment at a clinical-AI company where reliability decides whether a patient gets their medication that day.
Latent's AI automates prior authorization, the insurer paperwork that decides whether and when a patient gets treatment. It's live across 45+ US health systems and runs for millions of patients a year. The work is real and the stakes are real, so reliability here is the mission, not an afterthought.
This is a deliberately application-leaning SRE role. You'll own production end-to-end and write Python and TypeScript to make it reliable, fixing performance in the code itself, not just maintaining the infrastructure around it. We care about what you've built, not the logos on your CV: if you cut your teeth owning production at a seed or Series A startup and want that again, this is for you.
What you'll do
- Own the production environment end-to-end - design, build and run it (Kubernetes + Helm, Terraform IaC).
- Optimise the TypeScript and Python/ML deployment pipelines for high-velocity, high-reliability releases.
- Define and evolve SLIs/SLOs and tie reliability targets to patient impact.
- Standardise observability with OpenTelemetry + Honeycomb - metrics, traces, logs, dashboards, alerts.
- Hunt performance bottlenecks in real application code and fix them: query optimisation, indexing, caching, connection pooling, safe retries.
- Lead incident response and turn post-incident learnings into durable, shipped fixes.
You'll thrive here if you
- Grew up in high-growth startups and want ownership, not process or layers.
- Spent real time as a backend software engineer and still write strong Python and TypeScript.
- Know Kubernetes, Helm, Terraform and Terragrunt hands-on, on AWS, with PostgreSQL, Redis and Kafka.
- Have around 7 years of hands-on experience and are comfortable with ambiguity, scope and moving fast.
- Bonus: regulated / health-tech (HIPAA) experience, open-source contributions.