Site Reliability Engineer
We are looking for a Site Reliability Engineer based in Latin America to work on a long-term project for one of our clients, a software company based in San Ramon, California. Our client provides cloud solutions trusted by governments worldwide to accelerate their digital transformation, deliver vital services, and build stronger communities.
Responsibilities
- Lead deep technical analysis during critical incidents, coordinating cross-functional teams to identify root causes, restore services quickly, and drive long-term reliability improvements.
- Automation and Toil Reduction: Writing code and scripts (Python, Go, or Bash) to eliminate repetitive manual work and handle provisioning.
- Incident Management: Responding to system alerts, troubleshooting production issues, joining on-call rotations, and restoring services quickly.
- SLO and Error Budget Management: Establishing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to balance feature delivery speed with system stability.
- Monitoring and Observability: Building dashboards and setting up tools (such as Prometheus, Grafana, or Datadog) to track system health and latency.
- Post-Incident Reviews: Conducting blameless post-mortems to find root causes and prevent repeat failures.
- Capacity Planning: Analyzing resource usage trends to forecast future infrastructure and scaling needs.
Requirements
- Advanced level of English.
- 5+ years of experience working as a Site Reliability Engineer.
- 1+ years of experience within Microsoft Azure.
- Strong experience with Python, Bash or Go for automation purposes.
- Experience working with Monitoring and Observability tools such as Prometheus, Grafana, or Datadog.
- Ability to keep a focused mind during high-severity production outages to lead teams effectively.
- Proven experience explaining complex infrastructure failures to non-technical business stakeholders in clear, simple terms.
- Strong prioritization instincts, with the ability to quickly determine whether to resolve a live issue manually or engineer a permanent automated solution.
- Experience working with AI apps or agents.
Bonus Points
- Bachelor's Degree in Computer Science, Systems Engineering or related fields.
- SaaS experience
What We Offer
- Long term positions
- Compensation in USD
- Paid time off
- Cool clients and products
- Work with great engineers
Site Reliability Engineer in san ramon at Unknown Company
This position is listed as full time and onsite.