About the Opportunity
Our client is seeking an innovative, hands-on senior reliability engineer who thrives at the intersection of cloud reliability, technical problem solving, and application team partnership. In this role, you will work closely with cross-functional teams to strengthen the availability, recoverability, and operational maturity of our client’s cloud platform in a regulated financial services environment.
Key Duties/Responsibilities:
- Debug and resolve complex issues across AWS services, Terraform configurations, and GitLab CI/CD pipelines in direct partnership with application teams
- Diagnose and remediate infrastructure failures across Lambda, ECS, RDS/Aurora, DynamoDB, SQS, S3, EventBridge, and API Gateway
- Design and implement multi-region AWS architectures aligned to application-specific RTO/RPO targets
- Guide teams through Terraform errors, state issues, provider misconfigurations, and module integration problems
- Troubleshoot GitLab pipeline failures — job dependencies, OIDC auth issues, runner configuration, artifact handling, and deployment stage failures
- Lead application teams through reliability assessments, production readiness reviews, and resilience improvement plans
- Establish reusable reliability patterns, standards, and scorecards that improve consistency across application teams
- Execute fault injection scenarios using AWS Fault Injection Service (FIS) across compute, database, network, and dependency layers
- Monitor application and infrastructure behavior during FIS experiments — capturing metrics, surfacing gaps, and driving remediation with owning teams
What You Bring:
- Bachelor’s degree in Computer Science, Software Engineering, or equivalent practical experience
- 10+ years of experience in the technology industry, including 5+ years of hands-on cloud engineering expertise
- 5+ years of hands-on experience in cloud engineering or site reliability engineering on AWS
- Expert knowledge of AWS architecture and reliability-related services including Lambda, ECS, RDS/Aurora, DynamoDB, S3, SQS, EventBridge, CloudWatch, and Route 53
- Strong diagnostic skills across AWS infrastructure, Terraform-based infrastructure as code, and CI/CD delivery workflows
- Hands-on experience with AWS Fault Injection Service (FIS) or equivalent chaos engineering tooling to validate system behavior under failure conditions
- Familiarity with AWS Resilience Hub for resiliency scoring, recommendations, and policy enforcement
- Experience with production readiness reviews, reliability scorecards, and resilience improvement planning
- Proficiency in Terraform to design, build, and maintain infrastructure as code with a focus on reliability patterns
- Expert-level knowledge of AWS IAM policy development, least-privilege role design, and Service Control Policies (SCPs)
- Strong CI/CD process experience with tools such as GitLab
- Proficiency writing, testing, and debugging Python applications and automation scripts
- Experience developing containerized and serverless microservices with a reliability-first approach
- Ability to influence reliability standards, mentor engineers, and drive enterprise-wide adoption of resilient cloud patterns
- Strong grasp of software development lifecycle (SDLC) and agile delivery concepts
- Comfortable delivering independently while also serving as a technical leader for cross-functional engineering teams
Preferred Certifications:
- AWS Certified Solutions Architect – Associate
- AWS Certified DevOps Engineer – Associate
- Other relevant certifications
Applications will be accepted through October 30th, 2026.
About Us:
Brooksource, Medasource, and Calculated Hire are part of the Eight Eleven Group family of companies and operate under Eight Eleven Group, LLC. All employees receive the same benefits, policies, and terms of employment.
EEO Statement:
Eight Eleven Group is an equal opportunity employer that does not discriminate on the basis of actual or perceived race, color, creed, religion, national origin, ancestry, citizenship status, age, sex or gender (including pregnancy, childbirth, lactation, and related medical conditions), gender identity or gender expression, sexual orientation, marital status, military service and veteran status, physical or mental disability, protected medical condition as defined by applicable state or local law, genetic information, or any other characteristic protected by applicable federal, state, or local laws and ordinances.
Benefits & Perks:
Brooksource offers competitive medical, dental, vision, Health Savings Account, Dependent Care FSA, and supplemental coverage with plans that can fit each employee’s needs. We offer a 401k plan that includes a company match and is fully vested after you become eligible, paid time off, sick time, and paid company holidays. We also offer an Employee Assistance Program (EAP) that provides services like virtual counseling, financial services, legal services, life coaching, etc.
Pay Disclaimer:
The pay range for this job level is a general guideline only and not a guarantee of compensation or salary. Additional factors considered in extending an offer include (but are not limited to) responsibilities of the job, education, experience, knowledge, skills, and abilities, as well as internal equity, alignment with market data, applicable bargaining agreement (if any), or other law.
Senior Cloud Reliability Engineer in Charlotte at Brooksource
This position is listed as full time and onsite. It was posted 3 days ago.