Unknown Company

Site Reliability Engineer

nj • Posted 2 days ago
Onsite Full Time IT & Technology

  • Participate in design reviews, sprint zero, and delivery planning to define and validate reliability requirements
  • Collaborate with Major Release Management to ensure releases meet SRE standards and support-readiness requirements
  • Define and improve monitoring, observability, dashboards, telemetry coverage, and alert strategy
  • Assist in major incident response and root cause analysis
  • Drive automation, intelligent tooling, and AI-assisted remediation
  • Serve as the operational readiness authority before production releases
  • Lead capacity, performance, workload trend, and resiliency analysis
  • Establish and track reliability metrics including availability, incident volume, MTTx, alert quality, automation coverage, and change failure rate
  • Participate in application reliability governance and service reviews
  • Prepare executive reporting on reliability posture, release readiness, observability maturity, incident trends, risks, and improvement outcomes
  • Promote SRE practices through mentoring, standards adoption, best-practice sharing, and approved AI tools

Requirements

  • Minimum of 10+ years of related technical experience across application support engineering, software engineering, site reliability engineering, production support, or application operations
  • Bachelor’s degree preferred or equivalent practical experience
  • Experience supporting business-critical applications in production environments
  • SRE, observability, automation, or ITIL certifications are a plus
  • Proven experience in one or more in-scope roles including Application Support Engineer, SDET, Software Engineer, or SRE
  • Strong understanding of monitoring and observability platforms, including dashboard design, alert tuning, telemetry coverage, log analysis, metrics, traces, and event correlation
  • Programming or scripting proficiency in one or more languages such as Python, Java, Go, PowerShell, or similar
  • Familiarity with distributed applications, middleware, messaging, batch processing, real-time processing, and production application behavior in high-availability environments
  • Experience in financial services, capital markets, regulated environments, or other high-availability operational settings
  • Demonstrated participation in disaster recovery, performance testing, resiliency testing, release readiness, incident response, and root cause analysis
  • Knowledge of AI concepts, data platforms, anomaly detection, incident correlation, and intelligent automation use cases
  • Strong collaboration skills across application support, application development, release management, risk, security, business, and vendor stakeholders
  • Ability to translate production support insights into actionable engineering improvements

Core Competencies

Demonstrates expertise in Site Reliability Engineering (SRE) practices, including monitoring, observability, and automation, while effectively collaborating with cross-functional teams to enhance production readiness and incident response. Proven ability to analyze performance metrics and drive improvements in high-availability environments.

Highest-signal resume keywords

  • Site Reliability Engineering (SRE)
  • Monitoring And Observability
  • Automation And Intelligent Tooling
  • Programming Proficiency In Python, Java, Go, Or PowerShell
  • Experience In Financial Services Or Regulated Environments

ATS Optimization Keywords

Hard Skills

  • Monitoring And Observability Platforms
  • Dashboard Design
  • Alert Tuning
  • Telemetry Coverage
  • Log Analysis
  • Metrics And Traces
  • Event Correlation
  • Disaster Recovery
  • Performance Testing
  • Resiliency Testing

Soft Skills

  • Strong Collaboration Skills
  • Mentoring And Best-Practice Sharing

Certifications & Qualifications

  • SRE Certification
  • Observability Certification
  • Automation Certification
  • ITIL Certification

Industry Keywords

  • Application Support Engineering
  • Software Engineering
  • Production Support
  • Application Operations
  • High-Availability Environments
  • Capital Markets
  • Regulated Environments

Tools & Technologies

  • AI Concepts
  • Data Platforms
  • Anomaly Detection
  • Incident Correlation
  • Intelligent Automation

#J-18808-Ljbffr

Site Reliability Engineer in nj at Unknown Company

This position is listed as full time and onsite.

Back to Job Search