You’ll define reliability standards, own incident management, and drive improvements for agents, proxies, and RAG pipelines. Strong Python, AWS, Kubernetes, and CI/CD experience are required.
#J-18808-LjbffrAI Platform SRE: Reliability, Observability & Scale in new york at Unknown Company
This position is listed as full time and onsite.