Unknown Company

Site Reliability Engineer (DevOps Engineer)

saratoga, ca • Posted 3 days ago
Onsite Full Time IT & Technology


  • Design, deploy, and maintain highly-scalable, highly-available software systems in AWS

  • Architect and manage containerized applications on Amazon EKS with focus on reliability and performance

  • Build and maintain Infrastructure as Code using Terraform for AWS cloud resources

  • Develop and optimize CI/CD pipelines for automated testing, deployment, and rollback capabilities

  • Implement comprehensive monitoring, alerting, and observability solutions using CloudWatch, Prometheus, and Grafana

  • Ensure system reliability through SLI/SLO definition, error budgets, and incident response procedures

  • Collaborate directly with engineering teams to optimize application deployment and operations

  • Manage deployments and scaling strategies to support mission-critical operations

  • Automate and enforce cloud security, governance, and compliance controls

  • Participate in on-call rotation and lead incident response for production level systems



  • Experience with database operations and scaling (RDS, Aurora, or similar)

  • Proficiency in Python and Bash scripting for automation

  • 5+ years of experience in SRE, DevOps, or Platform Engineering roles

  • Deep experience with Kubernetes, EKS, Helm, and container orchestration

  • Strong CI/CD pipeline development and management experience (Bitbucket preferred)

  • Experience with monitoring and observability tools (Prometheus, Grafana, ELK Stack)

  • Proven experience designing and operating mission-critical, highly-available systems within AWS

  • Knowledge of capacity planning and performance optimization

  • Advanced proficiency in Infrastructure as Code using Terraform (OpenTofu)

  • AWS Solutions Architect Professional, Certified Kubernetes Administrator (CKA), or equivalent expertise

  • Experience with incident management and post-mortem processes

  • Experience with GitOps workflows and tools (ArgoCD, Flux)

  • Experience with chaos engineering and disaster recovery planning

  • Knowledge of service mesh technologies (Istio, Linkerd)

  • Experience with Zero Trust Networking (ZTNA) or VPN solutions

  • Background in aerospace, defense, or other mission-critical industries

  • Strong intellectual curiosity and commitment to continuous learning

  • Exceptional attention to detail and an ownership mentality

#J-18808-Ljbffr
Back to Job Search