Cambridge Mobile Telematics (CMT) is seeking a Principal Site Reliability Engineer I, Machine Learning to help us change the world by building reliable ML infrastructure in the cloud. You will own SLOs, incident response, and the health of Ray clusters on AWS and Databricks workloads.
You will implement observability with CloudWatch and Datadog, craft automation with Terraform and CI/CD, and guide cost optimization across EC2/EKS, while collaborating across teams to deliver scalable, resilient
#J-18808-LjbffrPrincipal ML SRE: Scalable AI Infra on AWS in ma at Unknown Company
This position is listed as full time and onsite.