Ready to Lead Observability & Reliability Architecture?
At-a-Glance Snapshot
- Role: Senior Site Reliability Engineer / Observability Specialist
- Location / Work Model: Hybrid / Remote (US)
- Compensation / Perks: Competitive Salary + Equity + Comprehensive Benefits
Why Join Our Client?
- Massive Financial Scale: Build and optimize high-throughput liquidity and corporate financial platforms handling real-time, global payment flows.
- Greenfield Ownership: Establish core incident management and reliability frameworks from scratch—you'll set the playbook, not just follow one.
- Engineering-First Consulting: Enjoy the balance of hands-on technical execution (80%) and high-impact cross-functional team coaching (20%).
What You'll Do
- Engineered Observability: Build and optimize New Relic instrumentation (NRQL, APM, Logs) across Azure and AWS to streamline metrics collection using RED/USE frameworks.
- Infrastructure as Code (IaC): Author and maintain Terraform modules to manage monitoring configurations, alert pipelines, and cloud resources across multi-cloud environments.
- Incident Architecture: Establish early-stage incident response foundations, deploying Incident.IO, Slack/OpsGenie integrations, and automated escalation workflows to drive down MTTD/MTTR.
- Reliability Strategy: Define and govern actionable SLOs/SLIs and error budgets, optimizing log ingestion costs while eliminating alert fatigue for stream-aligned engineering teams.
- Technical Enablement: Partner directly with platform and product teams through hands-on coaching, post-incident reviews (PIRs), and Azure DevOps CI/CD pipeline automation.
What You Bring
- 7+ Years in SRE/DevOps: Deep experience scaling enterprise cloud infrastructure and production reliability in high-availability environments.
- Observability Mastery: Expert-level hands-on skills with New Relic (NRQL, Synthetics, APM) and defining actionable SLOs/SLIs. *MUST HAVE*
- Strong IaC & Automation: Advanced proficiency in Terraform and PowerShell scripting within enterprise Windows (80%) and Linux (20%) environments. *MUST HAVE*
- Cloud & CI/CD Expertise: Proven track record in Azure (App Services, Virtual Machines, Azure SQL) combined with Azure DevOps or Octopus Deploy pipelines.
- Incident Management Focus: Experience designing on-call rotations, runbooks, and incident response operations via Incident.IO, PagerDuty, or OpsGenie.
Ready to reshape global financial infrastructure and take complete ownership of production reliability?
#J-18808-Ljbffr