Unknown Company

Site Reliability Engineer (SRE)

phoenix, az • Posted 4 days ago
Onsite Full Time Architecture and Engineering Occupations
Site Reliability Engineer (SRE)

We are seeking an experienced Site Reliability Engineer (SRE) to ensure the reliability, availability, and performance of enterprise applications and infrastructure. The ideal candidate will have strong expertise in production support, system monitoring, incident management, automation, and operational excellence. This role requires close collaboration with application, infrastructure, database, and network teams to maintain production stability, drive continuous improvements, and minimize operational risk.

Key Responsibilities
  • Monitor, maintain, and support enterprise applications and infrastructure to ensure high availability, reliability, and performance.
  • Monitor system health, application performance, and production environments, responding promptly to alerts, incidents, and escalations.
  • Serve as a senior technical resource during production incidents, coordinating troubleshooting, resolution, and stakeholder communication.
  • Perform production maintenance activities, including system patching, upgrades, configuration changes, and environment validation.
  • Troubleshoot complex issues across application, infrastructure, operating systems, network, middleware, and database layers.
  • Partner with development, infrastructure, database, and operations teams to implement production changes while adhering to change management processes.
  • Conduct Root Cause Analysis (RCA) for production incidents and implement preventive measures to improve system reliability.
  • Identify recurring operational issues and develop automation solutions to eliminate manual tasks and improve operational efficiency.
  • Create, maintain, and update operational documentation, runbooks, SOPs, and knowledge base articles.
  • Support application governance, security, audit, compliance, and regulatory requirements.
  • Participate in disaster recovery, business continuity, resilience planning, and system recovery testing.
  • Ensure adherence to ITIL best practices for incident, problem, change, and release management.
  • Continuously improve system monitoring, observability, automation, and operational processes.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
  • Strong experience in Site Reliability Engineering (SRE), Production Support, or DevOps.
  • Hands-on experience with production monitoring, incident management, and operational support.
  • Strong knowledge of Linux/Unix operating systems and system administration.
  • Experience with scripting and automation using Python, Shell, Bash, or PowerShell.
  • Knowledge of monitoring and observability tools such as Splunk, Dynatrace, Grafana, Prometheus, AppDynamics, or similar.
  • Experience troubleshooting application, infrastructure, database, middleware, and network-related issues.
  • Understanding of CI/CD pipelines, deployment processes, and automation practices.
  • Experience with change management, release management, and ITIL processes.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent communication and stakeholder management abilities.
Preferred Qualifications
  • Experience with Kubernetes, Docker, OpenShift, or cloud platforms (AWS, Azure, or GCP).
  • Experience with Infrastructure as Code (Terraform, Ansible, or similar tools).
  • Knowledge of databases such as Oracle, SQL Server, PostgreSQL, or MongoDB.
  • SRE, DevOps, Cloud, or ITIL certifications are a plus.
Back to Job Search