Unknown Company

Site Reliability Engineer

chicago, il • Posted Today
Onsite Full Time Quality Engineering

Prolaio believes that continuous learning and collaboration can make a significant difference in how heart care is administered. We are creating smarter ways to address heart disease and heart risks by uniting patients, care teams, and researchers on a secure, technology-enabled platform that drives clinical innovation and offers a path towards better patient outcomes.

This is precision cardiology, and we know it's within reach.

What Will You Do?

The Overview

The Site Reliability Engineer will ensure that Prolaio's cardiovascular data platform reliably captures, transports, and delivers continuous biosensor data from participants to the clinical researchers and trial sponsors who depend on it. In this role, you will build and operate the monitoring, service level objectives, incident response, and automation needed to keep critical data flows healthy, identify failures before they impact studies, and ensure that device downtime never goes unnoticed simply because a shipment record says a device was delivered.

This role is ideal for a self-starter who thrives at the intersection of software reliability, connected devices, and clinical data. You will investigate incidents across the full participant-to-platform stack, combining service telemetry with device data and participant reports to determine whether an issue originated with the hardware, phone, firmware, mobile application, connectivity, or the platform itself. You will turn those investigations into durable improvements, automating manual checks, strengthening observability, reducing operational toil, and building the reliability practices that keep Prolaio's clinical data complete and trustworthy at scale.

The Specifics

  • Build monitoring that treats missing data as a failure rather than a quiet day - freshness, volume and quality checks that verify delivery from outside the system instead of trusting a job's exit code.
  • Define and maintain SLIs, SLOs and error budgets for customer-facing data services, and report attainment honestly, including when the news is bad.
  • Join the on-call rotation, run incident response, and write blameless postmortems that produce owned follow-up actions.
  • Investigate device failures end to end: correlate discharge curves, Bluetooth disconnects and wear detection against the participant's complaint, and establish or exonerate each candidate cause on evidence.
  • Treat the fleet as a population - failure rates per device-month, clustering by build lot, firmware version or site - rather than as a queue of individual tickets.
  • Close the loop. When an analysis cannot reach a verdict, name the missing metric or threshold and route it back into the
#J-18808-Ljbffr

Site Reliability Engineer in chicago at Unknown Company

This position is listed as full time and onsite.

Back to Job Search