Unknown Company

Site Reliability Engineer

westlake, tx • Posted 5 days ago
Onsite Full Time Architecture and Engineering Occupations

We are currently sourcing for a Site Reliability Engineer to work in a leading Enterprise Infrastructure Group in Westlake, TX or Merrimack, NH.
The Role
Our Site Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the Development Experience to deliver services at high scale, high availability, and resilience through automation and Infrastructure as Code. We build reliability into our ecosystem by applying standard methodologies in Resiliency Engineering, Automation, Observability, and Chaos Testing.
The team comes from diverse technical backgrounds, and the responsibilities provide opportunities for a variety of challenges. Ideal candidates will have a background in either software engineering or systems engineering with a desire to learn the other, or previous experience as an SRE. We are looking for a Systems Thinking, SRE Engineer who has helped teams scale through production insights, operational automation, developer guidance, real-time metrics, and automation.
Team
The Production Services organization is a centralized support services team within the Enterprise Infrastructure & Operations group. The team supports 3000+ applications across multiple business units and provides roster-based on-call rotation support (follow-the-sun model). Services include Platform, Application, and Batch Support through Incident Management, Change Management, Environment Management, Cloud, UI, Middle Tier & Database Services, Mainframe Operations, Release Services, and Performance Engineering Services.
The Expertise You Have

  • Bachelor's degree or equivalent experience or higher in a technology-related field (e.g., Engineering, Computer Science, etc.) required; Master's degree is a plus.
  • 5+ years of hands-on experience deploying and/or supporting highly distributed multi-tiered systems at scale.
  • 1-2 years of experience in Cloud development (AWS) and migration skills; experience building and operating highly resilient platforms in AWS cloud environments.
  • 2-4 years of experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation.
  • Hands-on experience with container orchestration, preferably Kubernetes.
  • Experience operating and implementing distributed and highly concurrent service-based architectures.
The Skills You Bring
  • Ability to automate using scripting languages such as Python and Shell scripting.
  • Experience managing systems using Infrastructure as Code tools (IAM, ARM, Terraform, Chef, etc.).
  • Solid understanding of Cloud Computing and DevOps concepts, including CI/CD pipelines.
  • Hands-on Kubernetes experience and knowledge.
  • Experience with one or more observability tools (Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, Datadog, etc.).
  • Experience building, operating, monitoring, logging, and alerting distributed systems at scale.
  • Proven experience maintaining scalability and resiliency in complex environments.
  • Proven experience implementing advanced observability practices and techniques at scale.
  • Demonstrated experience with modern monitoring tools (Datadog, Prometheus, Splunk, etc.).
  • Ability to triage issues, perform root cause analysis, and make decisions under pressure.
  • Experience managing and interpreting large datasets using query languages and visualization tools.
  • Strong communication skills with the ability to engage both technical and non-technical audiences.
  • Ability to learn new software, methods, and practices and introduce them to development teams.
  • Ability to work collaboratively with diverse individuals and groups, both in person and virtually, while building and maintaining effective relationships.

Site Reliability Engineer in westlake at Unknown Company

This position is listed as full time and onsite.

Back to Job Search