Seeking a Principal Engineer to own the resiliency strategy and reference architecture for a large-scale AWS environment. This individual will establish availability standards, lead business continuity and disaster recovery initiatives, and drive high-availability architecture across infrastructure, containers, data platforms, and engineering teams.
Required Skills & Experience
- 12+ years of engineering experience, including 7+ years owning large-scale resiliency or high-availability architecture.
- Deep hands-on AWS experience across multi-account, multi-region environments, networking, IAM, Route 53, load balancing, and Global Accelerator.
- Proven ownership of multi-AZ/multi-region architecture, BC/DR strategy, automated failover, RTO/RPO, and recovery testing.
- Expertise in the AWS Well-Architected Framework and production container platforms, including EKS and ECS/Fargate.
- Strong data resiliency experience with Aurora/RDS and at least one of DynamoDB, ElastiCache, or S3 replication.
- Advanced Terraform/IaC and infrastructure CI/CD experience, including state management, guardrails, and drift detection.
- Hands-on experience with chaos engineering, observability, SLOs, error budgets, and incident response.
- Strong cross-functional leadership with the ability to balance availability, cost, risk, and operational complexity.
Key Responsibilities
- - Own AWS resiliency strategy, availability standards, and multi-AZ/multi-region architecture.
- - Define appropriate active-active, active-passive, warm-standby, and pilot-light patterns.
- - Lead BC/DR planning, including RTO/RPO targets, automated failover, runbooks, DR testing, and game days.
- Conduct AWS Well-Architected Reviews and drive remediation efforts.
- - Build resiliency across ECS/Fargate, EKS, Aurora, RDS, - DynamoDB, ElastiCache, and S3.
- - Automate recovery using IaC, self-healing, drift detection, and fault injection.
- - Establish SLOs, error budgets, health checks, dependency mapping, and observability standards.
- - Lead availability incident response and convert recurring failures into architectural improvements.
- - Partner with engineering, product, finance, and business leaders to balance reliability, cost, and operational complexity.
Principal Engineer - Resiliency in austin at Unknown Company
This position is listed as full time and onsite.