Are you passionate about Site Reliability Engineering, automation, observability, and AI/ML-driven operations ? We’re looking for an experienced engineer who can help transform how mission-critical enterprise applications are deployed, monitored, and supported.
What You’ll Do
- Champion the SRE mindset and drive systematization of operational processes
- Build innovative tools and automation that reduce manual toil and improve operational efficiency
- Develop production-grade scripts, frameworks, and infrastructure automation
- Design AI/ML-driven automation, anomaly detection, predictive alerting, and proactive response solutions
- Expand automation across deployment, monitoring, alerting, and self-healing workflows
- Enhance observability and leverage AIOps/ML-assisted monitoring at scale
- Develop CI/CD orchestration and promote GitOps practices
- Use data-driven trend analysis and ML-informed forecasting for capacity planning
- Troubleshoot and resolve critical production issues across mission-critical platforms
- Partner with Engineering, Scrum, and Operations teams to improve availability, reliability, and rollout success
- Participate in an on-call rotation and support real-time production operations
We’re looking for someone who combines strong SRE/production engineering experience with a passion for automation, cloud platforms, observability, CI/CD, AIOps, and AI/ML.
#J-18808-LjbffrSite Reliability Engineer in austin at Unknown Company
This position is listed as full time and onsite.