NVIDIA Corporation in Santa Clara, CA is hiring a DevOps Engineer to operate our AI Data Center telemetry platform. You’ll own reliability, incident response, and postmortems for telemetry ingestion, processing, storage, and APIs/dashboards used by operators.
Expect to lead Kubernetes deployments end-to-end, build runbooks, and partner with Software and Systems Engineering to translate platform signals into actionable, trustworthy alerts and automation.
#J-18808-LjbffrSenior SRE, AIOps Platform for GPU Data Centers in santa clara at Unknown Company
This position is listed as full time and onsite.