SRE & DevOps EngineerAt CrowdStrike, our engineering organization depends on shared infrastructure platforms that power critical product capabilities at global scale. These platforms require dedicated engineering ownership to operate reliably, scale safely, harden for security, and mature into self-service capabilities that teams across the organization can depend on.As an SRE & DevOps Engineer, you will own production infrastructure spanning multiple cloud providers and regions, serving engineering teams across CrowdStrike. The work is equal parts reliability engineering and DevOps engineering - building automation, hardening security, establishing governance, and enabling consuming teams to adopt these platforms effectively.You will work across a rich technology landscape including Kubernetes, Kafka, Cassandra, PostgreSQL, Apache Pinot, OpenSearch, and Apache Spark — operating and scaling microservices-based distributed systems that process millions of security events per second with zero tolerance for data loss or downtime.What You'll Do:Run production infrastructure - Deploy, upgrade, and maintain platform services across multiple clouds and regions on Kubernetes, including microservices-based distributed systemsOwn delivery pipelines - Build and maintain scalable CI/CD pipelines using Jenkins, GitLab CI, and Bitbucket Pipelines with GitOps workflows via ArgoCD or FluxOwn capacity planning - Track usage, forecast growth, right-size clusters, and optimize infrastructure costs across multi-cloud environmentsBuild observability - Implement metrics, dashboards, alerts (Prometheus/Grafana), distributed tracing (Jaeger/OpenTelemetry), and actionable runbooksOwn on-call and incidents - Participate in on-call rotation, lead incident resolution, write blameless postmortems, and automate repeat problemsDrive reliability - Apply SRE principles and AI-driven automation to move from reactive firefighting to proactive operationsHarden security - Implement auth, encryption, secret rotation, and network policies.Own disaster recovery - Build and test backup strategies and failover mechanisms ensuring zero data lossOperate data infrastructure - Maintain reliability of Kafka, Cassandra, PostgreSQL, Apache Pinot, and OpenSearch in productionEnable and collaborate - Support engineering teams with templates and patterns; partner with Infrastructure, SRE, and Data Services on shared operational problemsExperience and Background:8+ years in SRE, Devops engineering, or infrastructure engineeringHands-on experience running stateful distributed systems and microservices architectures on Kubernetes in productionBachelor's degree in Computer Science or related field, or equivalent work experienceReliability Engineering:Deep understanding of SRE principles - SLOs, SLIs, error budgets applied to large-scale distributed systemsStrong incident management background - on-call ownership, blameless postmortems, and turning operational pain into automationExperience with chaos engineering and resilience validation for production systemsProven ability to build and maintain systems with zero tolerance for data loss or downtimeAdvanced observability experience including Prometheus, Grafana, distributed tracing (Jaeger/OpenTelemetry), and large-scale log aggregation (ELK/Splunk) with a focus on building custom SLO dashboards and reliability scorecardsProgramming & Automation:Proficiency in Python and/or Golang for automation, tooling, and platform servicesStrong scripting and automation skills - if you do it by hand more than once, you automate itPlatform and Delivery Engineering:CI/CD pipeline experience - Building and owning scalable delivery pipelines using Jenkins, GitLab CI, Bitbucket Pipelines, Tekton, or equivalentStrong proficiency in Infrastructure as Code (IaC) — Terraform, Ansible, Pulumi, or equivalentCloud and Big Data Exposure:Proficiency in at least one cloud environment (AWS, Azure, GCP) with emphasis on multi-region architecture, cloud-native reliability patterns, and security-first cloud designStrong experience with Kubernetes at scale - managing large cluster fleets, workload orchestration, and container lifecycle managementFamiliarity with distributed data systems including relational databases (PostgreSQL), NoSQL (Cassandra), OLAP (Pinot), Indexing(OpenSearch) and real-time streaming platforms (Kafka, Flink)Exposure to Big Data and analytics technologies like Spark, StormBenefits of Working at CrowdStrike:Market leader in compensation and equity awardsComprehensive physical and mental wellness programsCompetitive vacation and holidays for rechargePaid parental and adoption leavesProfessional development opportunities for all employees regardless of level or roleEmployee Networks, geographic neighborhood groups, and volunteer opportunities to build connectionsVibrant office culture with world class amenitiesGreat Place to Work Certified™ across the globeCrowdStrike is proud to be an equal opportunity employer.
We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program.CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law.
We base all employment decisions--including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay-offs, return from lay-off, terminations and social/recreational programs--on valid job requirements.If you need assistance accessing or reviewing the information on this website or need help submitting an application for employment or requesting an accommodation, please contact us at for further assistance.Find out more about your rights as an applicant.CrowdStrike participates in the E-Verify program.Notice of E-Verify ParticipationRight to WorkCrowdStrike, Inc. is committed to fair and equitable compensation practices. Placement within the pay range is dependent on a variety of factors including, but not limited to, relevant work experience, skills, certifications, job level, supervisory status, and location.
The base salary range for this position for all U.S. candidates is $120,000 - $180,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off.For detailed information about the U.S. benefits package, please click here.