Unknown Company

Site Reliability Engineer – III

irvine, ca • Posted 4 days ago
Onsite Full Time Engineering

Responsibilities

  • Create, update, or automate internal business processes or tools to reduce toil and improve team productivity.
  • Understand and monitor the Taco Bell Digital ecosystem for performance, availability, and accuracy of transactional data.
  • Communicate and collaborate with both technical and non‑technical stakeholders on issues, upcoming changes, and updates to system health.
  • Perform final validation tests on various mobile and web‑based applications, reporting on, and offering feedback on areas for improvement.
  • Build expertise in serverless infrastructure and initiatives while also learning aspects of modern SRE practices and terms, such as SLIs, SLOs, Observability, toil, and incident response with blameless post‑mortems.
  • Work with and adopt Agile practices while participating in a 24/7 on‑call rotation.
  • Collaborate within the team and with cross‑functional partners on high‑impact business issues that affect revenue and brand reputation.

Requirements

  • Bachelor’s degree in computer science, engineering, OR a related field, OR equivalent work experience.
  • At least 2+ years of experience in the SRE space, with a focus on observability and automation.
  • Familiarity with SRE core principles (e.g. SLO, SLA, SLI, Error Budget, etc.).
  • Hands‑on experience creating monitors, dashboards, SLOs, and other observability capabilities.
  • Experience with logging solutions or platforms such as DataDog, CloudWatch log insights, etc.
  • Understanding of incident management practices, including leading bridge calls, conducting RCAs, and facilitating postmortems.
  • Familiarity with modern observability practices and tools, such as distributed tracing, APM, OpenTelemetry, and RUM.
  • Excellent communication and collaboration skills, with the ability to work effectively in a fast‑paced environment as a member of a team.
  • A fundamentally complete understanding of Observability principles (not just monitoring) + experience using tools like DataDog, Lumigo, CloudWatch, or similar.
  • General level understanding of Agile methods such as Kanban, Scrum, etc.
  • Advanced troubleshooting skills.
  • A curious mindset and the desire to always keep learning.
  • Proactive self‑starter capable of operating autonomously.
  • Ability to participate in an on‑call rotation.
  • Proficiency in SQL.
  • Fundamental understanding of JavaScript, Python, Go, or TypeScript.
  • Working knowledge of AWS services commonly used in serverless and cloud‑native environments, including Lambda, API Gateway, Fargate, S3, DynamoDB, and EventBridge.

#J-18808-Ljbffr
Back to Job Search