Unknown Company

Principal Engineer - Major Incident Response & ITIL Platform

marlborough, ma • Posted 1 weeks ago
Remote Full Time IT & Technology

The Principal Engineer, Major Incident Response & ITIL Platform Lead is a senior individual contributor and program leader who combines deep technical expertise with operational discipline. This role owns the design, configuration, and continuous evolution of the ITIL practice — including Major Incident Response (MIR), Post-Incident Review (PIR), and Problem Management — while serving as a hands-on engineer within ServiceNow and adjacent tooling platforms.

Unlike a traditional SDM role, this position is explicitly technical: you will architect workflows, build automation, instrument observability, and drive platform maturity across stores, distribution centers, and digital environments. You will also lead a high-performing offshore team and act as the primary program authority during high-severity events — bridging the gap between engineering execution and executive communication.

Key Responsibilities:

Major Incident Response (MIR) Program Leadership

  • Own and operate the MIR program end-to-end — from playbook authorship to real-time bridge command — for incidents impacting stores, DCs, POS, fuel, e-commerce, and membership systems.
  • Serve as Incident Commander during P1/P2 events, driving technical triage, stakeholder communication, and escalation decisions under pressure.
  • Design and maintain a universal MIR playbook with consistent execution standards 24x7, including on-call rotations for nights, weekends, and holidays.
  • Establish leadership notification templates, technical bridge protocols, and business-facing communication cadences during major incidents.
  • Instrument incident severity classification logic, auto-routing, and escalation thresholds directly within ServiceNow.

Post-Incident Review & Postmortem Excellence

  • Own the end-to-end PIR lifecycle — blameless, data-driven reviews completed within SLA — and enforce action-item closure rigor.
  • Build and maintain an enterprise-wide RCA library, problem signatures, and trend intelligence within ServiceNow's CMDB and Problem Management modules.
  • Partner with SRE and Software Engineering to translate RCA findings into reliability-driven design improvements and automated runbooks.
  • Configure and manage PIR workflows, SLA timers, and notification rules natively in ServiceNow — no manual handoffs.

ITIL Platform Engineering & ServiceNow Ownership

  • Act as a hands-on technical owner of ServiceNow ITSM modules: Incident, Problem, Change, and Event Management.
  • Design and build ServiceNow workflows, business rules, UI policies, Flow Designer automations, and integration spokes connecting monitoring platforms (Dynatrace, Splunk, PagerDuty/AlertOps, etc.).
  • Develop and maintain custom dashboards, real-time KPI reporting, and SLA/SLO tracking within ServiceNow Performance Analytics.
  • Own the Problem Management lifecycle: identification, logging, root cause investigation, routing, and verified resolution.
  • Surface recurring incident patterns from trend analysis and feed intelligence back into MIR and Service Excellence programs.
  • Ensure complete, accurate, and timely documentation of all Problems in ServiceNow with appropriate categorization and linkage to incidents and changes.
  • Lead and develop a high-performing offshore operations team, setting clear goals aligned to MIR and ITIL program objectives.
  • Drive a culture of automation-first thinking: identify manual toil and eliminate it through ServiceNow scripting, Flow Designer, and third-party integrations.
  • Conduct regular retrospectives, process audits, and tooling reviews; translate findings into prioritized improvement backlog items.
  • Present program health, metrics, and roadmap updates to senior IT and business leadership.

Key Outcomes:

  • Faster stabilization of high-severity events through structured, technically informed incident command.
  • Measurable reduction in repeat incidents via high-quality, action-tracked RCAs.
  • A mature, automated ServiceNow platform that minimizes manual effort and accelerates response and reporting.
  • Predictable, trust-building communication to business stakeholders during and after major incidents.
  • Continuous improvement embedded into operational DNA — not a periodic exercise.

KPIs & Success Metrics:

Implement and manage service level agreements (SLAs and SLOs) to meet organizational goals and user expectations.

  • Mean Time to Acknowledge (MTTA) and Mean Time to Resolve (MTTR) for P1/P2 incidents.
  • Postmortem SLA compliance (e.g., 100% PIR completion within 5 business days).
  • Action item closure rate from PIRs within agreed timelines.
  • Reduction in repeat incidents (measured quarterly).
  • Problem Management throughput (number of problems logged, analyzed, and resolved).
  • Leadership communication SLA adherence during major incidents.
  • Continuous improvement initiatives delivered (e.g., automation, process optimization).
  • Stakeholder satisfaction scores from incident and problem management processes.

Requirements:

Education & Certifications

  • Bachelor's degree in Computer Science, Information Systems, or equivalent experience.
  • ITIL 4 Strategic Leader or Managing Professional certification strongly preferred; ITIL Expert acceptable.
  • ServiceNow certifications (CSA, CIS-ITSM, or CIS-Event Management) highly desirable.

Experience

  • 8+ years in IT Service Management with a strong technical bias — hands-on platform work, not just process governance.
  • 5+ years of direct experience administering or engineering ServiceNow (workflow design, scripting, integrations, Performance Analytics).
  • Proven track record leading Major Incident Response in large-scale retail, e-commerce, or distributed digital environments.
  • Experience owning Problem Management and PIR programs with measurable outcomes (repeat incident reduction, MTTR improvement).
  • Demonstrated ability to manage and develop onshore/offshore teams in a follow-the-sun operations model.
  • Hybrid working model: 3 days onsite (Tue, Wed, Thu), 2 days remote (Mon, Fri).
  • Occasional travel to company locations or industry events.
  • Flexibility in working hours to accommodate global operations and time zone differences.
  • Participation in Major Incident on-call rotation.

Technical Skills

  • Deep ServiceNow platform expertise understanding platform mechanics that drive configuration and administration.
  • Working knowledge of tools such as Dynatrace, Splunk, AlertOps/PagerDuty, or equivalent; ability to build event-to-incident automation bridges. Observability & Monitoring:
  • MS Teams, Jira — including integration design with ServiceNow. Collaboration & Comms:

Leadership & Soft Skills

  • Able to command a major incident bridge — calm, decisive, technically credible under pressure.
  • Executive-level communication: concise, audience-aware, and trustworthy during crises.
  • Program management discipline: roadmaps, metrics, stakeholder alignment, and backlog ownership.
  • Growth mindset with a bias toward automation and measurable improvement.

This is a hybrid role. Tuesday through Thursday are in-office days at BJ's Club Support Center in Marlborough, MA and Monday and Friday are remote days.

In accordance with the Pay Transparency requirements, the following represents a good faith estimate of the compensation range for this position. At BJ’s Wholesale Club, we carefully consider a wide range of non-discriminatory factors when determining salary. Actual salaries will vary depending on factors including but not limited to location, education, experience, and qualifications. The pay range for this position is $137,500.00 - $175,500.00

#J-18808-Ljbffr

Principal Engineer - Major Incident Response & ITIL Platform in marlborough at Unknown Company

This position is listed as full time and able to be worked remotely.

Back to Job Search