Unknown Company

Site Reliability Engineer

austin, tx • Posted 1 weeks ago
Onsite Contract IT & Technology

Role: Site Reliability Engineer

Location: Southlake / Austin, TX - Onsite 4 days weekly

Duration: 12 Months

Job Summary

We are seeking a motivated Site Reliability Engineer (Contractor) with 3 to 5 years of experience in automation, cloud infrastructure, and production operations. The ideal candidate has a strong automation-first mindset and enjoys solving complex operational challenges. This role focuses on improving reliability, observability, operational efficiency, and automation across on-premises and cloud platforms.

Key Responsibilities

Automation & Platform Engineering

  • Develop Python-based automation solutions to reduce manual operational effort.
  • Automate infrastructure management across Linux, Windows, Kubernetes, GCP, and cloud-native environments.
  • Integrate tools and platforms through APIs and client libraries.
  • Assist in implementing infrastructure automation using Ansible, Terraform, or similar technologies.
  • Support CI/CD automation and deployment reliability initiatives.

Reliability & Operations

  • Monitor and maintain production systems to meet reliability and availability objectives.
  • Participate in incident response, troubleshooting, and root cause analysis activities.
  • Develop automation and operational improvements to prevent recurring issues.
  • Support disaster recovery, failover testing, and operational readiness activities.
  • Perform performance analysis and system health reviews.

Observability & Monitoring

  • Build and maintain dashboards, alerts, and monitoring solutions using Splunk, Grafana, Prometheus, GCP Operations Suite, or similar tools.
  • Improve visibility into application and infrastructure health through metrics, logs, and traces.
  • Investigate alerts and identify opportunities to reduce noise and improve detection.

AIOps & Intelligence

  • Explore AI/ML-driven operational improvements such as anomaly detection, intelligent alerting, and log analytics.
  • Assist in developing automation solutions that leverage AI to improve operational efficiency.
  • Participate in evaluating emerging AIOps capabilities and observability technologies.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience.
  • 3 to 5 years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or Platform Engineering.
  • Strong programming skills in Python for automation and tooling development.
  • Experience supporting Kubernetes and cloud platforms (GCP, AWS, or Azure).
  • Familiarity with infrastructure automation and configuration management tools.
  • Experience with monitoring and observability platforms such as Splunk, Grafana, Prometheus, Datadog, or similar.
  • Understanding of Linux systems, networking, and distributed applications.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Ability to work effectively in fast-paced, mission-critical environments.

Preferred Qualifications

  • Experience with Terraform, Ansible, or Infrastructure as Code solutions.
  • Exposure to OpenTelemetry and modern observability practices.
  • Experience with CI/CD pipelines and deployment automation.
  • Knowledge of AI/ML, AIOps, or intelligent operational tooling.
  • Experience supporting highly available production systems in regulated or enterprise environments.

Must Have:

  • 3-5 years of hands‑on Site Reliability Engineering or Production Engineering experience
  • Must have supported production systems at scale.
  • Operations ownership, incident response, and reliability engineering experience required.
  • Pure DevOps, build/release, or CI/CD-only backgrounds are not a fit.
  • Strong Python and operation automation Development Skills
  • Demonstrated experience building automation tools, scripts, frameworks, or operational solutions.
  • Candidate should be able to provide examples of automation they personally developed.
  • Python must be a primary skill, not just basic scripting.
  • Production Operations & Incident Management Experience
  • Experience troubleshooting critical production incidents.
  • Root Cause Analysis (RCA) participation and problem remediation.
  • Experience reducing operational toil through automation.

#J-18808-Ljbffr

Site Reliability Engineer in austin at Unknown Company

This position is listed as contract and onsite.

Back to Job Search