Unknown Company

Site Reliability Engineer

san ramon, ca • Posted 6 days ago
Onsite Full Time Architecture and Engineering Occupations
Site Reliability Engineer

We are looking for a Site Reliability Engineer based in Latin America to work on a long-term project for one of our clients, a software company based in San Ramon, California. Our client provides cloud solutions trusted by governments worldwide to accelerate their digital transformation, deliver vital services, and build stronger communities.

Responsibilities
  • Lead deep technical analysis during critical incidents, coordinating cross-functional teams to identify root causes, restore services quickly, and drive long-term reliability improvements.
  • Automation and Toil Reduction: Writing code and scripts (Python, Go, or Bash) to eliminate repetitive manual work and handle provisioning.
  • Incident Management: Responding to system alerts, troubleshooting production issues, joining on-call rotations, and restoring services quickly.
  • SLO and Error Budget Management: Establishing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to balance feature delivery speed with system stability.
  • Monitoring and Observability: Building dashboards and setting up tools (such as Prometheus, Grafana, or Datadog) to track system health and latency.
  • Post-Incident Reviews: Conducting blameless post-mortems to find root causes and prevent repeat failures.
  • Capacity Planning: Analyzing resource usage trends to forecast future infrastructure and scaling needs.
Requirements
  • Advanced level of English.
  • 5+ years of experience working as a Site Reliability Engineer.
  • 1+ years of experience within Microsoft Azure.
  • Strong experience with Python, Bash or Go for automation purposes.
  • Experience working with Monitoring and Observability tools such as Prometheus, Grafana, or Datadog.
  • Ability to keep a focused mind during high-severity production outages to lead teams effectively.
  • Proven experience explaining complex infrastructure failures to non-technical business stakeholders in clear, simple terms.
  • Strong prioritization instincts, with the ability to quickly determine whether to resolve a live issue manually or engineer a permanent automated solution.
  • Experience working with AI apps or agents.
Bonus Points
  • Bachelor's Degree in Computer Science, Systems Engineering or related fields.
  • SaaS experience
What We Offer
  • Long term positions
  • Compensation in USD
  • Paid time off
  • Cool clients and products
  • Work with great engineers

Site Reliability Engineer in san ramon at Unknown Company

This position is listed as full time and onsite.

Back to Job Search