Unknown Company

Senior Software Engineer, Robinhood Command Center

new york, ny • Posted 1 weeks ago
Onsite Full Time Software Architecture & Engineering

Requirements

  • 5+ years of software engineering experience, including significant experience operating production systems
  • 2+ years focused on reliability engineering, infrastructure, distributed systems, or production operations
  • Hands-on experience serving in incident leadership roles (e.g., IMOC, incident commander, primary oncall)
  • Strong communication and cross-functional collaboration skills, especially during high-severity incidents
  • Deep knowledge of systems reliability, observability frameworks, and fault-tolerant architecture design
  • Experience with multi-region or multi-cluster architectures, capacity planning, and failover strategies
  • Familiarity with modern observability stacks (e.g., OpenTelemetry, Prometheus, Grafana)
  • Demonstrated ability to drive measurable improvements in MTTD, MTTR, availability, or customer impact

What the job involves

  • We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards
  • The Robinhood Command Center (RCC) is a newly formed reliability team that serves as the front line for detecting, coordinating, and mitigating production incidents across Robinhood
  • As part of Robinhood’s broader reliability initiative, RCC works closely with product engineering, reliability, observability, infrastructure, and business teams to reduce customer impact and shorten incident duration
  • As a Senior Engineer, you will be part of the founding RCC team, helping define how Robinhood responds to and learns from incidents at scale
  • This is a highly visible role focused on incident leadership, operational excellence, and reliability tooling
  • You will not own product services or core infrastructure, but you will own the processes and tools that enable fast, high-quality incident response
  • Serve as a senior technical leader driving the long-term reliability and observability strategy across Robinhood’s infrastructure
  • Partner closely across many different types of engineers to raise the bar for operational excellence and incident response
  • Lead incident mitigation efforts by coordinating service owners, facilitating time-sensitive decisions like rollbacks, traffic shifts, and maintaining a clear source of truth during active incidents
  • Develop and maintain incident management processes and procedures to ensure timely resolution and minimize customer impact
  • Own incident discovery at the company level by defining and maintaining global dashboards and alerts tied to critical user journeys (CUJs), availability, and business-impact metrics
  • Own and evolve incident response tooling and processes, including education, adoption, and measurement of MTTD/MTTR improvements
  • Drive post-incident governance and learning, defining standards for postmortems, SEV reviews, and follow‑up tracking to ensure durable reliability improvements
  • Design and implement next‑generation failure mitigation strategies that avoid full‑region or full‑datacenter failovers
  • Define and build frameworks to improve monitoring, alerting, and observability across hundreds of services and systems
  • Define and own the roadmap of bringing observability to critical user journeys for Robinhood’s products
  • Deliver key insights and executive‑level reporting to enable better business decisions around service quality and reliability
  • Act as a force multiplier through mentoring, technical influence, and contributions to hiring and engineering culture

#J-18808-Ljbffr
Back to Job Search