A global leading quantitative trading firm is looking for an experienced Site Reliability Engineer (SRE) to help build, operate, and continuously improve the infrastructure that powers the electronic trading platform. You'll work closely with software engineers, quantitative researchers, and traders to ensure the systems remain reliable, scalable, secure, and performant in a fast-paced, low-latency environment.
What You'll Do
- Design, build, and maintain highly available production infrastructure.
- Improve the reliability, scalability, and performance of trading and supporting systems.
- Develop automation to eliminate manual operational tasks.
- Build and maintain monitoring, alerting, and observability platforms.
- Lead incident response, troubleshooting, and post-incident reviews.
- Partner with development teams to improve application reliability throughout the software lifecycle.
- Drive infrastructure-as-code and configuration management best practices.
- Implement capacity planning, performance testing, and disaster recovery processes.
- Champion operational excellence through documentation, automation, and continuous improvement.
What We're Looking For
- Strong experience as a Site Reliability Engineer, DevOps Engineer, or Infrastructure Engineer.
- Excellent Linux systems administration skills.
- Strong scripting or programming experience in Python or GO.
- Experience with Kubernetes.
- Experience with on-premise infrastructure.
- Strong understanding of networking fundamentals (TCP/IP, DNS, routing, load balancing).
- Experience with monitoring and observability.
- Strong troubleshooting skills with a methodical approach to solving complex production issues.
- Excellent communication and collaboration skills.
Site Reliability Engineer in New York City Metropolitan Area at Radley James
This position is listed as full time and onsite. It was posted 5 days ago.