Your CareerPalo Alto Networks is at the forefront of cloud-native infrastructure, where reliability, scale, and intelligent automation define the future of operations. As a Senior Site Reliability Engineer, you will design and operate the platforms that power our applications across GCP, AWS, and global data centers — and you'll push the boundary of what's possible by leveraging AI and machine learning to transform how we approach SRE.This isn't just about keeping the lights on. You'll build intelligent systems that predict incidents before they happen, automate root cause analysis, and continuously optimize our infrastructure.
You'll be a critical bridge between engineering and our Infrastructure Platform, combining deep SRE expertise with AI-driven automation to deliver unprecedented levels of reliability and operational efficiency.If you're excited about applying AI to real-world infrastructure challenges — and you thrive in an environment where automation isn't just a nice-to-have but a core philosophy — this is your next career.Your ImpactDesign, build, and operate cloud infrastructure that enables reliable, rapid deployment of microservices with resilient operations and effective monitoringLeverage AI/ML to automate incident detection, root cause analysis, and remediation — reducing toil and accelerating mean time to resolutionBuild and integrate AI-powered tools (e.g., LLM-based agents, AIOps platforms) into SRE workflows for intelligent alerting, log analysis, and capacity planningWrite automation code for provisioning and operating infrastructure at massive scaleDevelop self-healing systems that can automatically detect anomalies, diagnose issues, and take corrective action with minimal human interventionWork with development teams to ensure applications are production-ready, scalable, and reliable from the ground upIdentify and drive opportunities to improve automation for code deployment, management, and observability of application servicesEstablish end-to-end monitoring and alerting on all critical components, incorporating AI-driven anomaly detection and predictive analyticsParticipate in the on-call rotation supporting the platform and production applicationsLead root cause analysis of critical business and production issues, building runbooks and automation to prevent recurrenceMentor other SREs on best practices in infrastructure orchestration, production troubleshooting, and AI-augmented operationsRepresent SRE in design reviews and work cross-functionally with engineering teams on operational readinessQualificationsYour Experience7+ years of experience in DevOps, Site Reliability, or infrastructure engineeringExpertise in multi-cloud environments — strong hands-on experience with GCP, AWS, and familiarity with OCI (Oracle Cloud Infrastructure)Experience designing and operating infrastructure across multiple cloud providers, including networking, identity management, and cross-cloud connectivityExpertise in Infrastructure as Code with tools such as Terraform, AnsibleStrong proficiency in Python and shell scripting for automationStrong experience with Linux and distributed systems handling high-volume transactionsFamiliarity with CI/CD pipelines, GitLab, and ArtifactoryStrong fundamentals in HTTP, web servers, and networkingBS or MS in Computer Science, a related field, or equivalent professional experienceExcellent problem solving, critical thinking, communication, and teamwork skillsSelf-disciplined, self-managed, self-motivated with a strong sense of ownership, urgency, and driveExperience applying AI/ML to operational workflows (e.g., AIOps, intelligent alerting, automated remediation, or LLM-powered tooling) is a strong plusExperience with cloud compliance frameworks (FedRAMP, IL5) and operating in regulated environments is a plusExperience building and managing large database systems — relational (MySQL, PostgreSQL) and non-relational (Redis, BigQuery, etc.) — is a plus
Principal Cloud Infrastructure Engineer (Advanced Threat Protection) in santa clara at Unknown Company
This position is listed as full time and onsite.