Unknown Company

Site Reliability Engineer (SRE)

woonsocket, ri • Posted 5 days ago
Onsite Full Time Architecture and Engineering Occupations

Job Description

n

A large healthcare client of ours is seeking a Site Reliability Engineer to join a fast-paced operations team responsible for maintaining platform stability, managing critical incidents, and supporting enterprise applications. This individual will play a key role in incident response, service reliability, cross-functional coordination, and operational excellence.

n

The ideal candidate brings a strong operations mindset, excels under pressure, and can effectively lead technical incident management efforts while communicating with both technical and business stakeholders.

n

Responsibilities

n

Lead and coordinate response efforts for P1/P2 production incidents.

n

Serve as Incident Commander during major outages and service disruptions.

n

Drive incident triage and engage appropriate engineering, infrastructure, and business teams.

n

Monitor application and platform health to ensure system availability and reliability

n

Facilitate root cause analysis activities and support post-incident reviews

n

Manage escalation processes and ensure timely resolution of critical production issues.

n

Draft executive-level and customer-facing communications during outages and service-impacting events.

n

Partner with engineering, infrastructure, cloud, application support, and business teams to improve operational processes.

n

Participate in an on-call rotation supporting enterprise operations.

n

Identify opportunities for automation and operational efficiency improvements.

n

Support reliability initiatives focused on reducing incident frequency and improving service performance

n

payrate - $50-60

n

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy:

n

Skills and Requirements

n

3-7 years of experience in Reliability Engineering, Production Support, Operations Engineering, or Incident Management

n

Hands-on experience managing high-severity production incidents (P1/P2).

n

Strong experience leading incident bridges and coordinating cross-functional response efforts

n

Excellent verbal and written communication skills with the ability to communicate effectively with technical teams and executive leadership.

n

Ability to thrive in a fast-paced operational environment.

n

Strong troubleshooting and problem-solving skills

n

Experience working within an on-call support model Experience supporting cloud-based environments (AWS, Azure, or GCP)

n

Knowledge of observability and monitoring tools such as Dynatrace, Splunk, Datadog, AppDynamics, Prometheus, or Grafana

n

Familiarity with ITIL processes and Major Incident Management frameworks

n

Experience supporting large-scale enterprise applications within healthcare, retail, or highly regulated environments

n

Exposure to automation and scripting using Python, Bash, PowerShell, or similar technologies

Site Reliability Engineer (SRE) in woonsocket at Unknown Company

This position is listed as full time and onsite.

Back to Job Search