The Principal Engineer, Major Incident Response & ITIL Platform Lead is a senior individual contributor and program leader who combines deep technical expertise with operational discipline. This role owns the design, configuration, and continuous evolution of the ITIL practice — including Major Incident Response (MIR), Post-Incident Review (PIR), and Problem Management — while serving as a hands-on engineer within ServiceNow and adjacent tooling platforms.
Unlike a traditional SDM role, this position is explicitly technical: you will architect workflows, build automation, instrument observability, and drive platform maturity across stores, distribution centers, and digital environments. You will also lead a high-performing offshore team and act as the primary program authority during high-severity events — bridging the gap between engineering execution and executive communication.
Key Responsibilities:
Major Incident Response (MIR) Program Leadership
- Own and operate the MIR program end-to-end — from playbook authorship to real-time bridge command — for incidents impacting stores, DCs, POS, fuel, e-commerce, and membership systems.
- Serve as Incident Commander during P1/P2 events, driving technical triage, stakeholder communication, and escalation decisions under pressure.
- Design and maintain a universal MIR playbook with consistent execution standards 24x7, including on-call rotations for nights, weekends, and holidays.
- Establish leadership notification templates, technical bridge protocols, and business-facing communication cadences during major incidents.
- Instrument incident severity classification logic, auto-routing, and escalation thresholds directly within ServiceNow.
Post-Incident Review & Postmortem Excellence
- Own the end-to-end PIR lifecycle — blameless, data-driven reviews completed within SLA — and enforce action-item closure rigor.
- Build and maintain an enterprise-wide RCA library, problem signatures, and trend intelligence within ServiceNow's CMDB and Problem Management modules.
- Partner with SRE and Software Engineering to translate RCA findings into reliability-driven design improvements and automated runbooks.
- Configure and manage PIR workflows, SLA timers, and notification rules natively in ServiceNow — no manual handoffs.
ITIL Platform Engineering & ServiceNow Ownership
- Act as a hands-on technical owner of ServiceNow ITSM modules: Incident, Problem, Change, and Event Management.
- Design and build ServiceNow workflows, business rules, UI policies, Flow Designer automations, and integration spokes connecting monitoring platforms (Dynatrace, Splunk, PagerDuty/AlertOps, etc.).
- Develop and maintain custom dashboards, real-time KPI reporting, and SLA/SLO tracking within ServiceNow Performance Analytics.
- Own the Problem Management lifecycle: identification, logging, root cause investigation, routing, and verified resolution.
- Surface recurring incident patterns from trend analysis and feed intelligence back into MIR and Service Excellence programs.
- Ensure complete, accurate, and timely documentation of all Problems in ServiceNow with appropriate categorization and linkage to incidents and changes.
- Lead and develop a high-performing offshore operations team, setting clear goals aligned to MIR and ITIL program objectives.
- Drive a culture of automation-first thinking: identify manual toil and eliminate it through ServiceNow scripting, Flow Designer, and third-party integrations.
- Conduct regular retrospectives, process audits, and tooling reviews; translate findings into prioritized improvement backlog items.
- Present program health, metrics, and roadmap updates to senior IT and business leadership.
Key Outcomes:
- Faster stabilization of high-severity events through structured, technically informed incident command.
- Measurable reduction in repeat incidents via high-quality, action-tracked RCAs.
- A mature, automated ServiceNow platform that minimizes manual effort and accelerates response and reporting.
- Predictable, trust-building communication to business stakeholders during and after major incidents.
- Continuous improvement embedded into operational DNA — not a periodic exercise.
KPIs & Success Metrics:
Implement and manage service level agreements (SLAs and SLOs) to meet organizational goals and user expectations.
- Mean Time to Acknowledge (MTTA) and Mean Time to Resolve (MTTR) for P1/P2 incidents.
- Postmortem SLA compliance (e.g., 100% PIR completion within 5 business days).
- Action item closure rate from PIRs within agreed timelines.
- Reduction in repeat incidents (measured quarterly).
- Problem Management throughput (number of problems logged, analyzed, and resolved).
- Leadership communication SLA adherence during major incidents.
- Continuous improvement initiatives delivered (e.g., automation, process optimization).
- Stakeholder satisfaction scores from incident and problem management processes.
Requirements:
Education & Certifications
- Bachelor's degree in Computer Science, Information Systems, or equivalent experience.
- ITIL 4 Strategic Leader or Managing Professional certification strongly preferred; ITIL Expert acceptable.
- ServiceNow certifications (CSA, CIS-ITSM, or CIS-Event Management) highly desirable.
Experience
- 8+ years in IT Service Management with a strong technical bias — hands-on platform work, not just process governance.
- 5+ years of direct experience administering or engineering ServiceNow (workflow design, scripting, integrations, Performance Analytics).
- Proven track record leading Major Incident Response in large-scale retail, e-commerce, or distributed digital environments.
- Experience owning Problem Management and PIR programs with measurable outcomes (repeat incident reduction, MTTR improvement).
- Demonstrated ability to manage and develop onshore/offshore teams in a follow-the-sun operations model.
- Hybrid working model: 3 days onsite (Tue, Wed, Thu), 2 days remote (Mon, Fri).
- Occasional travel to company locations or industry events.
- Flexibility in working hours to accommodate global operations and time zone differences.
- Participation in Major Incident on-call rotation.
Technical Skills
- Deep ServiceNow platform expertise understanding platform mechanics that drive configuration and administration.
- Working knowledge of tools such as Dynatrace, Splunk, AlertOps/PagerDuty, or equivalent; ability to build event-to-incident automation bridges. Observability & Monitoring:
- MS Teams, Jira — including integration design with ServiceNow. Collaboration & Comms:
Leadership & Soft Skills
- Able to command a major incident bridge — calm, decisive, technically credible under pressure.
- Executive-level communication: concise, audience-aware, and trustworthy during crises.
- Program management discipline: roadmaps, metrics, stakeholder alignment, and backlog ownership.
- Growth mindset with a bias toward automation and measurable improvement.
This is a hybrid role. Tuesday through Thursday are in-office days at BJ's Club Support Center in Marlborough, MA and Monday and Friday are remote days.
In accordance with the Pay Transparency requirements, the following represents a good faith estimate of the compensation range for this position. At BJ’s Wholesale Club, we carefully consider a wide range of non-discriminatory factors when determining salary. Actual salaries will vary depending on factors including but not limited to location, education, experience, and qualifications. The pay range for this position is $137,500.00 - $175,500.00
#J-18808-LjbffrPrincipal Engineer - Major Incident Response & ITIL Platform in marlborough at Unknown Company
This position is listed as full time and able to be worked remotely.