We are seeking an experienced Site Reliability Engineer (SRE) with strong data center and bare-metal infrastructure experience. This role focuses on maintaining and improving the reliability of data center infrastructure through monitoring, automation, troubleshooting, and incident response.
You will work across servers, hardware, power/cooling infrastructure, monitoring, and automation , partnering with infrastructure, hardware, and data center operations teams to improve reliability and operational efficiency.
Key Responsibilities
- Monitor and improve data center infrastructure using Prometheus, Grafana, and Splunk , including server health and power/cooling telemetry.
- Develop Python and Shell scripts to automate troubleshooting, incident response, alert management, and other operational processes.
- Maintain and improve NetBox or similar data center inventory systems, including device, rack, and infrastructure information.
- Build and maintain Grafana dashboards to monitor server health, infrastructure performance, capacity, and power/cooling metrics.
- Use SQL and Splunk queries to troubleshoot infrastructure issues, analyze system data, and identify potential bottlenecks.
- Troubleshoot bare-metal servers, hardware, and data center infrastructure , including issues involving IPMI, PDUs, and power feeds.
- Participate in incident response and on-call support , including troubleshooting, root cause analysis, mitigation, and resolution of infrastructure issues.
- Develop and maintain runbooks and operational documentation for common server, hardware, power, cooling, and facility-related issues.
- Work with software, hardware, and data center operations teams to support reliable infrastructure deployments and ongoing operations.
Qualifications
- 8+ years of experience in SRE, Production Operations, Data Center Infrastructure, or a similar role.
- Strong hands-on experience with bare-metal servers, hardware troubleshooting, provisioning, and data center infrastructure .
- Experience with NetBox or similar DCIM/inventory tools .
- Strong SQL experience and experience working with REST APIs .
- Hands-on experience with Prometheus, Grafana, and Splunk .
- Strong understanding of IPMI, out-of-band server management, PDUs, and power distribution .
- Practical understanding of data center power and cooling systems , including HVAC, liquid cooling, hot/cold aisle containment, and air handling.
- Strong Python and Shell scripting skills with experience automating on-premises infrastructure.
- Bachelor's degree in Computer Science, Computer Engineering, or a related technical field preferred .
Site Reliability Engineer in austin at Unknown Company
This position is listed as full time and onsite.