Unknown Company

Principal Engineer - Resiliency

austin, tx • Posted 4 days ago
Onsite Full Time IT & Technology

Seeking a Principal Engineer to own the resiliency strategy and reference architecture for a large-scale AWS environment. This individual will establish availability standards, lead business continuity and disaster recovery initiatives, and drive high-availability architecture across infrastructure, containers, data platforms, and engineering teams.

Required Skills & Experience

  • 12+ years of engineering experience, including 7+ years owning large-scale resiliency or high-availability architecture.
  • Deep hands-on AWS experience across multi-account, multi-region environments, networking, IAM, Route 53, load balancing, and Global Accelerator.
  • Proven ownership of multi-AZ/multi-region architecture, BC/DR strategy, automated failover, RTO/RPO, and recovery testing.
  • Expertise in the AWS Well-Architected Framework and production container platforms, including EKS and ECS/Fargate.
  • Strong data resiliency experience with Aurora/RDS and at least one of DynamoDB, ElastiCache, or S3 replication.
  • Advanced Terraform/IaC and infrastructure CI/CD experience, including state management, guardrails, and drift detection.
  • Hands-on experience with chaos engineering, observability, SLOs, error budgets, and incident response.
  • Strong cross-functional leadership with the ability to balance availability, cost, risk, and operational complexity.

Key Responsibilities

  • - Own AWS resiliency strategy, availability standards, and multi-AZ/multi-region architecture.
  • - Define appropriate active-active, active-passive, warm-standby, and pilot-light patterns.
  • - Lead BC/DR planning, including RTO/RPO targets, automated failover, runbooks, DR testing, and game days.
  • Conduct AWS Well-Architected Reviews and drive remediation efforts.
  • - Build resiliency across ECS/Fargate, EKS, Aurora, RDS, - DynamoDB, ElastiCache, and S3.
  • - Automate recovery using IaC, self-healing, drift detection, and fault injection.
  • - Establish SLOs, error budgets, health checks, dependency mapping, and observability standards.
  • - Lead availability incident response and convert recurring failures into architectural improvements.
  • - Partner with engineering, product, finance, and business leaders to balance reliability, cost, and operational complexity.

#J-18808-Ljbffr

Principal Engineer - Resiliency in austin at Unknown Company

This position is listed as full time and onsite.

Back to Job Search