Senior Site Reliability Engineer - Platform
ChicagoNew York
We arelookingfora Site Reliability Engineer, to join our growing Platform Engineering team,who can cultivate our SRE philosophy, processes, and technologies from the ground up.This roleentailsdriving standards and fostering adoption across our technology teams, whilstcloselypartnering with our DevOps and Cloud teams.
With a hands-on approach,you'llwork across both cloud and on-premises hosting platforms, ensuring the reliability and scalability of ourtradingsystemsand production environments. This is a chance to play a pivotal role in transforming our operational capabilities and enhancing performance across a wide array of environments and platforms.
Key Responsibilities:
- Develop and promote our SRE philosophy,establishingbest practices and processes that will be instrumental in scaling our infrastructure.
- Implementand scaleend-to-end observability and monitoring solutions using Prometheus, Grafana, Loki, andTempo, ensuring high visibility into application performance and infrastructure health.
- Participate in on-call rotation with approximately 1 week per month of on-call time shared equally across members of the team
- Review and define standards forapplication reliability requirements within ourKubernetesenvironment, ensuringapplication configuration isoptimizedfor performance,costand reliability.
- Develop automation and tooling to improve efficiency and reliability of deployment pipelines, system health checks, and recovery procedures.
- Collaborate with development teams to enhance service stability, scalability, and fault tolerance through SRE best practices like blameless post-mortems and service levelobjectives(SLOs).
To be considered a good fit, you musthave:
- 8+ years of experience in SRE or similar roles within complex, distributed systems environments.
- SMEwith key SRE technologies such as Prometheus, Grafana, Loki,Tempo,andOpenTelemetry.
- Extensive knowledge of container orchestration using Kubernetes and containerization with Docker.
- Hands-on experience with both cloud (AWS preferred) and on-premiseshosting platforms.
- Proven ability to script in languages like Python, Bash, or Go, to automate routine tasks and deployment pipelines.
- Strong understanding of CI/CD principles, agile methodologies, and DevOps culture.
- High levelof initiative, passion for reliability engineering, detail orientation, and follow-through capabilities.
- Exceptional interpersonal and communication skills, with the ability to explain complex technical concepts to a diverse audience.
With respect to NY, CA, and IL based applicants, the starting base pay range for this role is between USD and USD annually. The actual base pay is dependent upon several factors, including, but not limited to, relevant experience, business needs and market demands. This role may also be eligible for bonus compensation and employee benefits.
#J-18808-Ljbffr