Unknown Company

Senior Manager, Incident Management

workfromhome • Posted 6 days ago
Remote Full Time Management Occupations

Leading a team of Site Reliability Engineers, the full-time Senior Manager, Incident Management will oversee the "Order-to-Cash," "Procure-to-Pay," and "Record-to-Report" lifecycles, ensuring the resilience and scalability of the global SaaS ecosystem while transitioning from reactive support to proactive engineering in a remote environment. Key responsibilities Lead, mentor, and grow a team of SREs while fostering a culture of accountability and psychological safety Architect observability across complex business processes and partner with business owners to define Service Level Objectives (SLOs) Own the Major Incident Response process and lead Root Cause Analysis (RCA) efforts to promote continuous learning and reduce manual toil Required qualifications 8+ years in SRE, DevOps, or Production Engineering, including 2+ years in people management Deep understanding of Order-to-Cash or Procure-to-Pay cycles and their impact on business operations Experience managing enterprise ecosystems such as NetSuite, SAP, Workday, or Salesforce Solid knowledge of Networking (SD-WAN, VPNs), Identity (IAM), and Endpoint Management Proficiency with monitoring tools like Datadog, Splunk, New Relic, or Prometheus

Back to Job Search