Unknown Company

Site Reliability Engineer

charlotte, nc • Posted 1 weeks ago
Hybrid Full Time Insurance

The position is described below. If you want to apply, click the Apply button at the top or bottom of this page. You'll be required to create an account or sign in to an existing one.If you have a disability and need assistance with the application, you can request a reasonable accommodation.

Send an email to Accessibility (accommodation requests only; other inquiries won't receive a response).Regular or Temporary:RegularLanguage Fluency:  English (Required)Work Shift:1st Shift (United States of America)Please review the following job description:Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This is a full-time leadership role responsible for owning monitoring design, driving platform decisions, and guiding engineering teams toward modern SRE practices.This individual will act as the technical authority for monitoring and alerting, shaping how signals from Dynatrace flow into ServiceNow and enterprise messaging/paging platforms, and enabling a shift toward automated, intelligent, and self-healing operations.Key ResponsibilitiesStrategic Leadership & Decision-MakingDefine and own the enterprise monitoring and SRE observability strategyServe as the subject matter expert for Dynatrace, ServiceNow integration, and alerting architectureEvaluate and recommend tooling, integration patterns, and platform directionDrive decisions on alerting philosophy, noise reduction, and signal quality improvementPlatform Ownership & ArchitectureArchitect and standardize end-to-end monitoring and SRE pipelines:Dynatrace ServiceNow incident lifecycleAlert correlation, deduplication, and prioritizationIntegration with paging systems (PagerDuty, SMS, voice, Teams)Establish best practices for:Event ingestion and enrichmentIncident routing and automated assignmentIntegration with CMDB and service mappingSite Reliability Engineering (SRE) LeadershipLead adoption of SRE principles, including:SLIs, SLOs, and error budgetsReliability engineering practices across servicesProactive monitoring and resilience designChampion a shift from reactive operations to proactive reliability engineeringInfluence application and platform teams to build observable, resilient systems by designAutomation & Self-Healing EnablementDrive development of automated remediation and self-healing capabilitiesLeverage Dynatrace workflows, Azure services, and automation frameworks to:Reduce manual incident handlingEliminate repeatable operational tasksMinimize unnecessary pagingServiceNow & Observability Integration LeadershipOwn integration between Dynatrace and ServiceNow ITSM/ITOM, including:Incident, Event Management, and CMDB alignmentService mapping and dependency visibilityGovernance for application/service taggingDefine standards for:Automated incident creation and resolutionPriority assignment and routing logicMonitoring-to-ITSM data synchronizationTeam Leadership & Cross-Functional InfluenceProvide technical leadership and mentorship across SRE, platform, and application teamsAct as a central point of coordination between engineering, cloud, and ITSM teamsLead workshops and working sessions to:Drive monitoring standardizationAlign teams on reliability practicesInfluence upstream architectural decisionsOperational ExcellenceEstablish KPIs and drive improvement in:Incident response and resolution timesAlert quality and paging effectivenessMonitoring coverage across critical servicesProvide leadership with clear visibility into service health and reliability trendsRequired Qualifications7+ years in Site Reliability Engineering, monitoring, or production engineeringProven experience in a technical leadership or lead engineer roleDeep hands-on experience with:Dynatrace (or equivalent observability platforms)Microsoft Azure (IaaS, PaaS, networking, identity)ServiceNow ITSM / ITOM (incident, event management, CMDB)Demonstrated ability to:Design and lead enterprise monitoring/SRE architecturesDrive platform and tooling decisionsIntegrate observability, ITSM, and paging solutionsPreferred QualificationsExperience leading SRE or observability transformation initiativesStrong expertise with Dynatrace–ServiceNow integrationsExperience modernizing or consolidating paging/on-call toolingFamiliarity with:Azure-based SRE tooling or AI-assisted operationsAutomation frameworks (GitHub Actions, Runbooks, etc.)Infrastructure as Code (Terraform, ARM, Bicep)Success MetricsReduction in alert noise and unnecessary pagingImproved incident routing accuracy and MTTRIncreased adoption of self-healing and automated workflowsStrong alignment between monitoring, CMDB, and service ownershipEnterprise-wide adoption of SRE and monitoring standardsGeneral Description of Available Benefits for Eligible Employees of CRC Group: At CRC Group, we're committed to supporting every aspect of teammates' well-being – physical, emotional, financial, social, and professional. Our best-in-class benefits program is designed to care for the whole you, offering a wide range of coverage and support.

Eligible full-time teammates enjoy access to medical, dental, vision, life, disability, and AD&D insurance; tax-advantaged savings accounts; and a 401(k) plan with company match. CRC Group also offers generous paid time off programs, including company holidays, vacation and sick days, new parent leave, and more. Eligible positions may also qualify for restricted stock units and/or a deferred compensation plan.CRC Group supports a diverse workforce and is an Equal Opportunity Employer that does not discriminate against individuals on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status or other classification protected by law.

CRC Group is a Drug Free Workplace.EEO is the LawPay Transparency Nondiscrimination ProvisionE-VerifySummaryLocation: Charlotte NC - 600 S Tryon St.Type: Full time

Site Reliability Engineer in charlotte at Unknown Company

This position is listed as full time and hybrid.

Back to Job Search