Unknown Company

Sr Software Engineer - Reliability Engineering

village of north hills, ny • Posted 2 weeks ago
Onsite Full Time IT & Technology

We're hiring a Sr. Software Engineer - Reliability Engineer who can code across the stack and cares deeply about reliability. You'll design and build infrastructure, observability tooling, and operational systems—treating resilience and debuggability as first-class concerns. You'll own projects end-to-end: from architecture → code → deployment → production. You'll split time between infrastructure-as-code, incident response, system improvements, and mentoring. You'll work on a team that ships quality systems while maintaining operational excellence across a platform serving millions of dealership transactions daily.

What You'll Do:

SRE Best Practices & System Health Management

  • Design and implement resilience initiatives: redundancy, failover, disaster recovery, data protection.
  • Write infrastructure-as-code (Terraform); manage 50+ AWS accounts with infrastructure patterns.
  • Own system health: proactively maintain application performance, minimize downtime, ensure consistent user experience.
  • Evolve team's SRE standards and practices.

Application Monitoring & Observability

  • Build observability into systems: logging, metrics, distributed tracing, alert design.
  • Improve monitoring frameworks; enable faster incident detection and resolution.
  • Design dashboards and alerts that help teams understand system behavior.
  • Partner with application teams on service instrumentation.

AWS Cost Optimization

  • Drive significant reductions in cloud spend through architectural improvements and resource utilization.
  • Review infrastructure for efficiency; identify and eliminate waste.
  • Balance cost, performance, and reliability in design decisions.

Software Development & Architecture

  • Build production systems, APIs, internal tools, and automation with clean, well-tested code.
  • Design for maintainability, operational simplicity, and reliability.
  • Participate in code review and technical design discussions.
  • Mentor junior engineers on code quality and architectural thinking.

Operations & Incident Response

  • Participate in on-call rotations; debug and resolve production incidents.
  • Conduct postmortem analysis; drive systemic improvements.
  • Develop operational procedures and runbooks.

Qualifications:

  • 5+ years software engineering, platform engineering, or infrastructure engineering experience.
  • Strong coding in Python, Go, Java, or equivalent; writes clean, testable code.
  • AWS hands-on: EC2, RDS, DynamoDB, S3, Aurora, Lambda, VPCs, Athena.
  • Terraform or equivalent infrastructure-as-code experience.
  • Docker and container orchestration (Kubernetes or similar).
  • Debugging on Linux and Windows platforms; able to troubleshoot complex systems using logs, metrics, and architectural knowledge.
  • System design thinking: can architect scalable systems and reason about trade-offs.
  • Availability for rotational on-call duties outside of standard business hours may be required.

  • Bachelor’s degree in a related discipline and 4 years’ experience in a related field. The right candidate could also have a different combination, such as a master’s degree and 2 years’ experience; a Ph.D. and up to 1 year of experience; or 16 years’ experience in a related field.

Highly Valued:

  • Experience with observability tools (New Relic, Splunk, Prometheus).
  • Incident response experience; familiar with postmortem practices.
  • Interest in or hands-on experience with SRE concepts (SLOs, resilience, failure modes).
  • Cost optimization mindset; has identified and eliminated cloud waste.
  • Windows and Linux system troubleshooting and performance analysis.
  • Experience with CI/CD pipelines and deployment automation.

#J-18808-Ljbffr

Sr Software Engineer - Reliability Engineering in village of north hills at Unknown Company

This position is listed as full time and onsite.

Back to Job Search