Unknown Company

Product Infrastructure Engineer, Data & Agent Systems

san francisco, ca • Posted 1 weeks ago
Onsite Full Time General

Product Infrastructure EngineerTruewind is building AI agents that help accounting teams close books faster and more accurately. Our agents read documents, prepare workpapers, reconcile transactions, draft structured outputs, and operate across ERP and financial systems.To make this reliable in production, we need a strong product infrastructure engineer who can build the data foundation and execution systems underneath the product.This is a backend-leaning infrastructure role for someone who can work across production data models, correctness-sensitive workflows, and agent execution systems. It is data-first in the near term: the primary focus is migrating Truewind from legacy data models into cleaner, more durable domain models while keeping live customer workflows working. As that foundation gets stronger, the role also expands into the execution infrastructure that lets AI agents safely complete real work.This is not a prompt engineering role, a pure analytics data role, or a pure DevOps/SRE role.

It is a product infrastructure role for someone who has lived through messy production systems and can move between backend services, data correctness, workflow reliability, and product-facing infrastructure.Why This Role MattersTruewind is in the middle of a major platform transition.We are migrating from legacy schemas into cleaner domain models. These systems need to run side by side while we move product modules, preserve customer behavior, validate correctness, and avoid breaking production workflows.Financial data has very little margin for silent error. A missing transaction, duplicated record, stale sync, or incorrect mapping can cascade into a wrong close. You will work with data from ERPs, banks, spreadsheets, PDFs, file uploads, and customer-provided documents that arrives in inconsistent formats and needs to be normalized, validated, audited, and made useful before humans or agents act on it.At the same time, our agents are becoming more capable.

They need reliable execution infrastructure: long-running jobs, retries, workspaces, artifacts, review flows, logs, traces, and failure recovery.In a larger company, this might be split across data platform, product infrastructure, and agent runtime teams. At our stage, we need someone who can work across these layers without losing sight of correctness or product impact.Team and StackYou will work directly with the engineering and product team on infrastructure that is already in production with real customers. The work sits between backend engineering and data infrastructure, with a stack that includes TypeScript, PostgreSQL/Supabase, Drizzle, queue and workflow systems, cloud infrastructure, and Python or similar tools where they are the right fit for data and automation work.What You'll Work On1. Data Infrastructure and Model MigrationYou will help move Truewind from legacy data models to cleaner, more durable domain models while the product stays live.This includes:Building and maintaining data pipelines that ingest, normalize, transform, and serve correctness-sensitive financial dataMigrating customer-facing product modules from legacy schemas to new domain modelsMaintaining compatibility while legacy and new systems run side by sideDesigning schemas, repositories, services, APIs, and workflows around complex data modelsWriting migrations, backfills, validation checks, and test coverageBuilding data quality checks to catch missing, duplicate, stale, inconsistent, or incorrectly mapped recordsImproving observability around syncs, transformations, model transitions, and downstream product behaviorPreserving tenant isolation, auditability, and correctness across data flowsCreating internal tools that help engineers debug data pipeline and migration failures faster2.

Agent Execution SystemsYou will also help make our AI agents reliable enough for real production workflows.This includes:Building orchestration for long-running agent workflows, including queues, retries, cancellations, checkpoints, resumability, and failure recoveryDesigning workspace and artifact handling for documents, workbooks, logs, generated outputs, and intermediate filesBuilding tool-calling infrastructure for agents to interact with files, APIs, documents, browsers, CLIs, and internal systemsImplementing human review flows where users can inspect, approve, reject, or modify agent outputsAdding traces, logs, workflow state, and root-cause debugging tools so agent work is auditable and debuggableIntroducing safer execution environments when agent tasks need to manipulate files, call tools, or run isolated codeYou May Be a Fit If You Have4+ years of experience in product infrastructure, backend engineering, data infrastructure, or distributed systemsStrong experience with relational databases, schema design, migrations, and data integrityExperience building data pipelines, ingestion systems, transformation layers, or backend services around complex data modelsExperience with async jobs, queues, workflow engines, retries, idempotency, and failure recoveryStrong coding ability in TypeScript, Python, Go, Rust, or similarStrong debugging instincts across data, backend, infrastructure, and workflow layersGood judgment around system boundaries, reliability, observability, and operational simplicityComfort working in a startup where you may need to move between product features, infrastructure, data pipelines, and internal toolingInterest in building systems where AI agents do real work, not just generate textStrong SignalsYou have migrated a production system from one data model to another while keeping the product runningYou have built or maintained production data pipelinesYou have worked on systems where data correctness really mattersYou have designed validation gates, audit logs, approval flows, or data quality checksYou have built workflow engines, internal platforms, automation infrastructure, or developer toolsYou have experience with multi-tenant SaaS systemsYou have worked with Postgres, Drizzle, Supabase, Temporal, Dagster, Airflow, Celery, BullMQ, Sidekiq, or similar systemsYou have worked with LLM agents, tool-calling systems, or human review workflowsYou enjoy turning messy real-world workflows into reliable, observable systemsWhat Makes This Role DifferentThis role sits at the intersection of data infrastructure, backend systems, product workflows, and agent execution. The center of gravity is not model behavior or demos. It is the production substrate that makes AI workflows reliable.You will help build:data model migration pathsbackend services around durable domain modelsvalidation and observability for correctness-sensitive dataworkflow reliability for long-running jobsartifact, trace, and review systems for agent outputssafer execution environments when agent workflows need themThe best person for this role is likely a strong product infrastructure engineer: someone backend-capable, data-correctness-minded, and comfortable moving systems forward without breaking production.Not a Fit IfYou mainly want to write promptsYou only want to work on model behaviorYou prefer demos over production reliabilityYou are looking for a pure DevOps/SRE role disconnected from product and data modelingYou are looking for a pure analytics or warehouse data engineering roleYou are uncomfortable working with complex data modelsYou do not enjoy migrations, backfills, validation, and system cleanupYou do not care about logs, traces, retries, idempotency, and observabilityYou want a narrowly scoped role with only one type of problem

Back to Job Search