Job Description
nA large healthcare client of ours is seeking a Site Reliability Engineer to join a fast-paced operations team responsible for maintaining platform stability, managing critical incidents, and supporting enterprise applications. This individual will play a key role in incident response, service reliability, cross-functional coordination, and operational excellence.
nThe ideal candidate brings a strong operations mindset, excels under pressure, and can effectively lead technical incident management efforts while communicating with both technical and business stakeholders.
nResponsibilities
nLead and coordinate response efforts for P1/P2 production incidents.
nServe as Incident Commander during major outages and service disruptions.
nDrive incident triage and engage appropriate engineering, infrastructure, and business teams.
nMonitor application and platform health to ensure system availability and reliability
nFacilitate root cause analysis activities and support post-incident reviews
nManage escalation processes and ensure timely resolution of critical production issues.
nDraft executive-level and customer-facing communications during outages and service-impacting events.
nPartner with engineering, infrastructure, cloud, application support, and business teams to improve operational processes.
nParticipate in an on-call rotation supporting enterprise operations.
nIdentify opportunities for automation and operational efficiency improvements.
nSupport reliability initiatives focused on reducing incident frequency and improving service performance
npayrate - $50-60
nWe are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy:
nSkills and Requirements
n3-7 years of experience in Reliability Engineering, Production Support, Operations Engineering, or Incident Management
nHands-on experience managing high-severity production incidents (P1/P2).
nStrong experience leading incident bridges and coordinating cross-functional response efforts
nExcellent verbal and written communication skills with the ability to communicate effectively with technical teams and executive leadership.
nAbility to thrive in a fast-paced operational environment.
nStrong troubleshooting and problem-solving skills
nExperience working within an on-call support model Experience supporting cloud-based environments (AWS, Azure, or GCP)
nKnowledge of observability and monitoring tools such as Dynatrace, Splunk, Datadog, AppDynamics, Prometheus, or Grafana
nFamiliarity with ITIL processes and Major Incident Management frameworks
nExperience supporting large-scale enterprise applications within healthcare, retail, or highly regulated environments
nExposure to automation and scripting using Python, Bash, PowerShell, or similar technologies
Site Reliability Engineer (SRE) in woonsocket at Unknown Company
This position is listed as full time and onsite.