- The role The primary responsibility for this position is to own the pipelines which touch protected health information (PHI). These are written in Python, using Ibis / DuckDB / Pandas, and run on Dagster as well as Kubernetes.
- You will be the design owner for our PHI pipeline portfolio and for our data ingestion system, working within an architecture the CTO sets (lake layout, PHI segregation, orchestration patterns). You are being hired to make good engineering decisions inside a well-de?ned boundary, explain them clearly, get them reviewed, and ship them with strong and creative use of the Codex AI coding assistant.
- You will also be the engineer of last resort for these pipelines. When a PHI pipeline fails at a moment that matters to a client, you are the person who answers.
- What you’ll own PHI pipeline portfolio. Design, build, and operate the pipelines that handle identi?ed patient data in Dagster on our DuckLake/Ibis/DuckDB lakehouse. This includes the patient-identity and record- linkage logic that determines whether a client’s data lines up correctly.
- Ingestion. Own the ingestion lifecycle end to end: client ?le submission through preprocessing, validation, enrichment, masking, and audit. Ingestion is event-driven on Azure Event Grid and Service Bus, and your job is to make it more stable, more observable, and more automated than it is today.
- Data-ops automation. Our ingestion recovery and data movement are still partly manual. Turn them into fully idempotent, re-runnable, record-level-retryable processes with enterprise-grade observability and governance.
- How we work AI-assisted development is how we build. Every engineer here uses Codex, and the interview includes a live pairing session with your own tools. What we’re looking for is not speed of generation but quality of steering: you use AI to move faster on work you could do yourself, and you can review, question, and defend everything it produces. You’ll work closely with the CTO and our other US engineers to keep re?ning how we do this.
- You can defend your résumé. For every system you list, we will ask why it was designed that way, what the alternatives were, and what you’d change now. We would rather see three systems you owned deeply than ten you were near.
- Design within an architecture, through code review. You’ll be working inside decisions you didn’t make and having your designs and code reviewed by people who did. If that’s uncomfortable, this isn’t the right role. If you know how to disagree, commit, and then improve things from the inside, it is.
- Overlap with Remote Workers. Many of our data-pipeline and data-operations engineers are in UTC+5. We ask for reliable availability by 8:00am Eastern for a few hours of daily overlap.
- Requirements Healthcare data literacy (required). You have worked with patient-level data under HIPAA, and you know what an MRN, a registry, a quality measure, a benchmark, and a BAA are without being told.
- Production SQL - you write it, tune it, and read query plans.
- Deep experience with one modern orchestrator (Dagster, Air?ow, Prefect, or similar) and the judgment to explain why it was the right choice and where it wasn’t. Dagster is a big plus.
- Ingestion lifecycle experience - ?le intake, validation, event-driven or queue-driven processing, failure handling and recovery, audit.
- Kubernetes - you have deployed and operated workloads on it. You don’t need to run the cluster but you know the di?erence between a node, an image, and a pod.
- Observability instincts - you build things so that when they fail, you know, and you know why.
- AI-assisted development practice with the discipline to own the output.
- Sound engineering reasoning across the board - you can explain the tradeo?s behind every design decision you’ve made.
- US-based, authorized to work in the US without sponsorship. Background check required.
- Eastern or Central time zone preferred. Will take PST candidates as well.
- Strongly preferred (one or more) Clinical registry or quality-reporting experience: NCDR, STS, GWTG harvest data, risk models,
- physician or service-line performance reporting.
- EMR/EHR integration: Epic, Cerner, Allscripts, HL7 v2/FHIR feeds at multi-hospital scale.
- Patient identity resolution / record linkage - if you’ve built it, we want to hear about it.
- On our stack: Azure, Dagster, DuckLake/DuckDB/Ibis, Event Grid and Service Bus, MySQL, AKS.
- We don’t expect you to know this exact combination. We do expect you to be the kind of engineer who has learned a new orchestrator and a new cloud before, and can tell us how it went.
Senior Data Engineer in Location not specified at Unknown Company
This position is listed as full time and able to be worked remotely.