Unknown Company

Site Reliability Engineer

new york, ny • Posted 1 weeks ago
Onsite Full Time Engineering

Harrison Clarke partners exclusively with venture-backed technology companies building category-defining products. We are currently conducting a confidential retained search on behalf of one of our flagship portfolio companies, a well-capitalised, mission-driven technology firm that has been operating at scale since 2014, with millions of global users and a reputation for rigorous engineering.

The Role

As Senior Site Reliability Engineer , you will own the infrastructure foundation that the entire engineering organization depends on. This isn't a support function, it's a strategic one. You'll work at the intersection of reliability, scalability, and developer experience, ensuring that a high-availability distributed system stays fast, resilient, and ready to grow.

What You'll Be Doing

  • Maintaining, improving, and securing cloud infrastructure across AWS and GCP , alongside Linux systems at scale
  • Partnering directly with engineering teams to streamline deployment, packaging, and troubleshooting of complex distributed applications
  • Driving CI/CD maturity using Jenkins and adjacent tooling
  • Building, operating, and continuously improving Kubernetes clusters in production; serving as the internal authority on container orchestration
  • Leading application migration efforts onto Kubernetes in close collaboration with development squads
  • Owning internal platform services including Prometheus and ELK
  • Monitoring high-availability environments and responding to incidents with urgency and rigour; conducting thorough, blameless post-mortems
  • Participating in architecture and code reviews, setting the bar for infrastructure best practice
  • Evaluating emerging technologies and making pragmatic decisions on adoption
  • Identifying and eliminating toil through intelligent automation

What You Bring

  • 5+ years in cloud-based systems operations as an SRE or DevOps engineer
  • Hands-on experience with infrastructure as code and configuration management
  • Strong command of SRE methodologies : SLOs/SLIs/error budgets, capacity planning, disaster recovery testing
  • Deep understanding of networking fundamentals
  • Proven track record managing production workloads with sophisticated monitoring and alerting
  • Comfort with on-call responsibilities and a systematic approach to incident management
  • Proficiency in at least one scripting or programming language
  • A collaborative, low-ego working style — you raise the team up, especially under pressure

Bonus Points

  • Ability to read and reason about Go , Rust , C++ , or TypeScript
  • Experience applying AI-driven approaches to operational automation

#J-18808-Ljbffr

Site Reliability Engineer in new york at Unknown Company

This position is listed as full time and onsite.

Back to Job Search