Cognativ in United States is seeking an experienced Site Reliability Engineer to own uptime across a large mixed estate, including GPU AI pipelines, Java services, and edge devices. This reliability-first role focuses on measurable reliability, not busywork, and seeks someone who defines SLOs, leads incident response, and drives automation.
You will plan capacity, manage disaster recovery, and advance observability with dashboards, alerts, and instrumentation.
#J-18808-Ljbffr