Job Description
nWe are seeking a Site Reliability Engineer (SRE) / DevOps Engineer to join a growing team responsible for supporting and scaling a Microsoft Azure environment. This role is ideal for an engineer who enjoys blending infrastructure, automation, deployment support, and reliability engineering to ensure highly available, scalable systems.
nAs the organization continues to grow its customer base, this individual will play a key role in improving observability, monitoring system health, diagnosing issues, and proactively preventing outages before they occur. While the position currently leans more heavily toward DevOps and infrastructure support, it will evolve into a balanced SRE/DevOps role with increased ownership of reliability initiatives, capacity planning, root cause analysis, and system performance optimization.
nResponsibilities
nSupport and maintain cloud infrastructure within Microsoft Azure
nBuild and manage CI/CD pipelines and deployment automation
nMonitor application and system health using Azure observability tools
nAnalyze logs, diagnose incidents, and troubleshoot production issues
nImprove monitoring, alerting, and overall system observability
nPartner with engineering teams to improve application reliability and performance
nImplement Infrastructure-as-Code (IaC) solutions
nParticipate in preventative maintenance, system optimization, and capacity planning efforts
nContribute to a culture of proactive reliability engineering and operational excellence
nWe are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy:
nSkills and Requirements
nExperience
n2-5 years of experience in a DevOps Engineer, Site Reliability Engineer (SRE), Cloud Engineer, or related role
nExperience working within Microsoft Azure environments
nStrong troubleshooting, problem-solving, and systems-thinking abilities
nExperience supporting applications in a .NET ecosystem
nAzure & Cloud Technologies
nAzure DevOps
nAzure Monitor
nApplication Insights
nLog Analytics
nAzure CLI
nAzure Container Apps
nInfrastructure & Automation
nYAML-based pipeline deployments
nBicep
nGit repositories and source control best practices
nInfrastructure-as-Code (IaC) experience
nReliability & Operations
nLog analysis and troubleshooting
nIncident diagnosis and resolution
nSystem health monitoring
nStrong observability mindset Kusto Query Language (KQL)
nARM Templates
nAdditional Azure platform expertise
nExperience implementing or managing observability solutions
nCapacity planning experience
nRoot cause analysis and post-incident review experience
nPrevious ownership of SRE initiatives or reliability programs
nBachelor's degree in Computer Science, Engineering, Information Systems, or a related field
SRE - PERM REMOTE in Remote at Unknown Company
This position is listed as full time and onsite.