Unknown Company

Site Reliability Engineer

austin, tx • Posted Yesterday
Onsite Full Time IT & Technology

We are seeking an experienced Site Reliability Engineer (SRE) with strong data center and bare-metal infrastructure experience. This role focuses on maintaining and improving the reliability of data center infrastructure through monitoring, automation, troubleshooting, and incident response.

You will work across servers, hardware, power/cooling infrastructure, monitoring, and automation , partnering with infrastructure, hardware, and data center operations teams to improve reliability and operational efficiency.

Key Responsibilities

  • Monitor and improve data center infrastructure using Prometheus, Grafana, and Splunk , including server health and power/cooling telemetry.
  • Develop Python and Shell scripts to automate troubleshooting, incident response, alert management, and other operational processes.
  • Maintain and improve NetBox or similar data center inventory systems, including device, rack, and infrastructure information.
  • Build and maintain Grafana dashboards to monitor server health, infrastructure performance, capacity, and power/cooling metrics.
  • Use SQL and Splunk queries to troubleshoot infrastructure issues, analyze system data, and identify potential bottlenecks.
  • Troubleshoot bare-metal servers, hardware, and data center infrastructure , including issues involving IPMI, PDUs, and power feeds.
  • Participate in incident response and on-call support , including troubleshooting, root cause analysis, mitigation, and resolution of infrastructure issues.
  • Develop and maintain runbooks and operational documentation for common server, hardware, power, cooling, and facility-related issues.
  • Work with software, hardware, and data center operations teams to support reliable infrastructure deployments and ongoing operations.

Qualifications

  • 8+ years of experience in SRE, Production Operations, Data Center Infrastructure, or a similar role.
  • Strong hands-on experience with bare-metal servers, hardware troubleshooting, provisioning, and data center infrastructure .
  • Experience with NetBox or similar DCIM/inventory tools .
  • Strong SQL experience and experience working with REST APIs .
  • Hands-on experience with Prometheus, Grafana, and Splunk .
  • Strong understanding of IPMI, out-of-band server management, PDUs, and power distribution .
  • Practical understanding of data center power and cooling systems , including HVAC, liquid cooling, hot/cold aisle containment, and air handling.
  • Strong Python and Shell scripting skills with experience automating on-premises infrastructure.
  • Bachelor's degree in Computer Science, Computer Engineering, or a related technical field preferred .
#J-18808-Ljbffr

Site Reliability Engineer in austin at Unknown Company

This position is listed as full time and onsite.

Back to Job Search