Unknown Company

Sr. Site Reliability Engineer

az • Posted 6 days ago
Onsite Contract IT & Technology

Senior Site Reliability Engineer — Cloud Communications / High-Compliance Enterprise SaaS

Our client is a large-scale, high-compliance cloud communications and healthcare data company that does business with the federal government and requires all employees to be U.S. citizens. For over 25 years, they've been a leader in their space and they're continuing to invest heavily in infrastructure modernization, automation, and platform reliability at global scale.

They are adding a 4th member to a tight-knit SRE team supporting mission-critical infrastructure that processes and delivers data across a global, high-volume, always-on platform. This is a highly visible role — you'll partner directly with Engineering and Information Security leadership on infrastructure strategy, not just execute tickets.

This role is built for a systems-first engineer — ideally a ~15–20 year Linux/UNIX systems engineer who has evolved into an SRE, rather than a developer who transitioned into infrastructure. This is not a role for someone purely cloud-native with no deep OS/networking background. The team needs someone who has genuinely run enterprise-scale platforms and lived through the operational fires that come with it — not someone who has only worked greenfield, cloud-native environments.

What you'll own:

  • Design, write, and maintain Terraform and Terragrunt modules (not just consumption) managing AWS infrastructure at scale, including EC2, S3, RDS, VPC, IAM, Lambda, and both ECS and EKS environments
  • Own and evolve CI/CD pipelines across GitHub Actions and AWS CodePipeline (GitLab experience a plus) — both for infrastructure and application deployments
  • Troubleshoot deep networking and mail delivery issues — DNS, IP/sender reputation, deliverability at global scale, and the class of problems that come with high-volume mail transport
  • Support and maintain Apache and Postfix-based mail infrastructure in production
  • Provide SRE support into Tomcat/Java-based application environments — diagnosing the class of problems Java/Tomcat teams run into (memory/GC issues, thread pool exhaustion, connection pooling, API contract issues) even without being a full-time Java developer
  • Build and maintain observability and alerting across the stack using tools like Prometheus, Grafana, OpenTelemetry, and the ELK stack; define and manage SLOs/SLIs for critical services
  • Support and extend configuration automation using tools such as AWS Config, SSM, Ansible, Puppet, or Chef
  • Work hands-on with containerization (Docker) and container orchestration, with particular depth in Amazon ECS
  • Use APM tooling (New Relic, Jaeger, Zipkin, or OpenTelemetry) to identify and resolve performance bottlenecks across distributed systems
  • Operate in a heavy on-call rotation supporting large-scale, high-availability systems (once fully staffed: 1 week on, 3 weeks off); lead and contribute to blameless postmortems following incidents
  • Draft and contribute to RFCs and internal standards for IaC, automation, and infrastructure design patterns
  • Mentor other engineers — including junior/lower-level ops staff — on Linux systems, IaC, CI/CD, and operational best practices
  • Ramp expectation: contribute to the codebase within 30 days, understand the environment deeply and take on smaller scoped projects by 60 days, own projects independently by 90 days

What you bring:

  • 15+ years of Linux/UNIX systems engineering, with meaningful time in SRE/DevOps roles at enterprise scale
  • Expert-level AWS (EC2, S3, RDS, VPC, IAM, Lambda, ECS/EKS, CloudWatch) and infrastructure design best practices
  • Expert-level Terraform and Terragrunt (module authorship required — this is not a "consume existing modules" role)
  • Strong production experience in Python and Go
  • Deep enterprise networking fundamentals — DNS, TCP/IP, mail routing and delivery troubleshooting, sender/IP reputation management
  • Hands-on Apache and Postfix administration experience
  • Prior exposure to Tomcat/Java-based production environments and the operational issues specific to that stack
  • Experience with configuration automation tooling (Ansible, Puppet, Chef, or AWS Config/SSM)
  • Solid containerization background (Docker), with strong familiarity in the ECS ecosystem specifically
  • Experience with observability/monitoring frameworks at scale (Prometheus, Grafana, OpenTelemetry, ELK, Thanos)
  • Experience operating in large-scale, high-compliance enterprise environments (healthcare, fintech, or similarly regulated industries — PCI/HITRUST/FedRamp not required, but that class of environment strongly preferred)
  • Customer-facing operational support background (UNIX/Windows) is a strong plus
  • Comfortable working within Agile/Scrum and Waterfall environments, using Jira/Confluence
  • Comfortable digesting large amounts of information quickly and moving from context to independent contribution fast
  • A self-starter mentality — able to work independently with minimal oversight while staying aligned with team goals

You’ll stand out if you also have:

  • Experience mentoring engineers into stronger DevOps/SRE practices
  • Background operating in PCI, HITRUST, or FedRamp/GovCloud-adjacent environments
  • Experience leading a team or platform through an SDLC/DevOps maturity transition
  • A GitHub profile or code samples that showcase personal or professional projects

#J-18808-Ljbffr

Sr. Site Reliability Engineer in az at Unknown Company

This position is listed as contract and onsite.

Back to Job Search