CoreWeave is seeking an experienced Production Engineer/SRE to ensure the reliability and scalability of one of the industry’s largest Kubernetes environments. You will build, operate, and scale infrastructure, develop automation in Go and other languages, and shape the systems that keep CoreWeave’s cloud running smoothly.
You will implement monitoring, alerting, and observability with Grafana-based tooling, drive incident response and post‑mortem improvements, and collaborate with multiple
#J-18808-LjbffrKubernetes Reliability Engineer - Scale AI Cloud in washington at Unknown Company
This position is listed as full time and onsite.