Unknown Company

Site Reliability Engineer

birmingham, al • Posted 3 days ago
Onsite Full Time IT & Technology

We are looking for a Site Reliability Engineer (SRE) to help build, maintain, and improve the infrastructure and systems that power our connected products and applications. This is a hands-on role focused on reliability, automation, monitoring, performance, and keeping production systems running smoothly at scale.

Responsibilities

  • Design, build, and maintain reliable cloud infrastructure and production systems in Microsoft Azure
  • Support and troubleshoot production applications and infrastructure across a high-scale, distributed environment
  • Work with Azure App Service, App Service Plans, Azure Functions, Event Grid, Event Hubs, Service Bus, and Azure SQL
  • Build and improve monitoring, alerting, logging, and observability to identify and resolve issues before they impact customers
  • Participate in production deployments, incident response, troubleshooting, and root cause analysis
  • Partner with software engineers to improve application reliability, performance, scalability, and resiliency
  • Develop automation that reduces manual operational work and improves consistency
  • Contribute to CI/CD pipelines and deployment processes
  • Perform capacity planning, performance testing, load testing, and reliability testing
  • Help identify system bottlenecks and implement improvements across infrastructure and applications
  • Participate in on-call and production support as needed
  • Establish and improve reliability standards, operational processes, and best practices across the engineering organization

Qualifications

  • 3+ years of experience working in Site Reliability Engineering, DevOps, Cloud Engineering, Infrastructure Engineering, or a similar role

Required Skills

  • Strong hands-on experience with Microsoft Azure
  • Experience supporting and troubleshooting production cloud environments
  • Experience with distributed systems, microservices, and cloud-based applications
  • Strong understanding of monitoring, logging, alerting, and application observability
  • Experience with CI/CD and automated deployments
  • Strong troubleshooting and problem-solving skills
  • Experience with scripting or automation using PowerShell, Python, Bash, or similar
  • Understanding of networking, databases, application performance, and cloud infrastructure
  • Ability to work closely with software engineers, DevOps engineers, and other technical teams
  • Comfortable taking ownership of production issues and driving problems through resolution

Preferred Skills

  • Experience with infrastructure as code is a plus
  • Experience with Azure App Services, Functions, Event Grid, Event Hubs, Service Bus, or Azure SQL
  • Experience with performance, load, capacity, or chaos testing
  • Experience with Kubernetes or containerized workloads
  • Experience with Terraform, Bicep, or other Infrastructure as Code tools
  • Experience with observability platforms such as Application Insights, Azure Monitor, Grafana, Prometheus, Splunk, or similar
  • Experience designing systems for high availability, scalability, and resiliency

#J-18808-Ljbffr

Site Reliability Engineer in birmingham at Unknown Company

This position is listed as full time and onsite.

Back to Job Search