Job Description
nWe are seeking a Senior Site Reliability Engineer (SRE) to play a leadership role in driving reliability, scalability, and observability across a modern, cloud-based application platform. This role will be instrumental in transitioning teams from reactive support to proactive engineering practices, establishing SRE standards, and mentoring junior team members.
nThe ideal candidate will bring deep expertise in cloud environments, application monitoring, and automation, combined with strong collaboration skills and the ability to influence engineering teams toward operational excellence.
nKey Responsibilities
nDesign, implement, and continuously improve monitoring, alerting, and observability frameworks across production environments
nLead troubleshooting of complex production issues and drive thorough root cause analysis to prevent recurrence
nOwn and enhance cloud-based application and platform reliability at scale
nPartner with and influence engineering teams to improve system performance, scalability, and resiliency through architectural guidance and best practices
nArchitect and implement automation of operational tasks and workflows to significantly reduce manual intervention
nDrive improvements in incident response processes, playbooks, and on-call procedures to minimize downtime
nLead monitoring and administration of containerized environments (Kubernetes/AKS)
nEstablish and champion SRE best practices and standards across teams, mentoring mid-level and junior engineers in SRE principles
nHelp guide teams through the cultural and technical transition from reactive support to proactive SRE practices
nWe are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy:
nSkills and Requirements
n~6–8+ years of experience in a Site Reliability Engineer or similar role (10+ years total IT experience)
nDeep expertise working in Azure cloud environments at scale
nExtensive hands-on experience with monitoring and observability tools (e.g., Elastic, Prometheus, Grafana, or similar), including designing and architecting monitoring strategies
nProven experience supporting production applications in complex, high-availability environments (application-focused SRE vs. infrastructure-only)
nStrong knowledge of Kubernetes (AKS) for monitoring, alerting, administration, and troubleshooting
nAbility to troubleshoot and debug applications at a deep level, including reading, understanding, and reviewing code
nSolid experience with .NET/C# application environments
nExperience with databases (SQL and/or NoSQL such as Cosmos DB, PostgreSQL, etc.)
nDemonstrated ability to mentor engineers and drive SRE adoption across teams Experience in multi-cloud or hybrid-cloud environments (Azure + AWS/GCP)
nExposure to EKS or non-Azure Kubernetes environments
nExperience supporting single-page applications (e.g., Angular)
nStrong background in automation and scripting (e.g., Python, Bash, PowerShell, Terraform)
nPrior experience leading teams through the transition to SRE best practices
nExperience working in global or distributed teams
nTrack record of reducing MTTR, improving SLOs/SLIs, and driving operational maturity