We're hiring a Sr. Software Engineer - Reliability Engineer who can code across the stack and cares deeply about reliability. You'll design and build infrastructure, observability tooling, and operational systems—treating resilience and debuggability as first-class concerns. You'll own projects end-to-end: from architecture → code → deployment → production. You'll split time between infrastructure-as-code, incident response, system improvements, and mentoring. You'll work on a team that ships quality systems while maintaining operational excellence across a platform serving millions of dealership transactions daily.
What You'll Do:
SRE Best Practices & System Health Management
- Design and implement resilience initiatives: redundancy, failover, disaster recovery, data protection.
- Write infrastructure-as-code (Terraform); manage 50+ AWS accounts with infrastructure patterns.
- Own system health: proactively maintain application performance, minimize downtime, ensure consistent user experience.
- Evolve team's SRE standards and practices.
Application Monitoring & Observability
- Build observability into systems: logging, metrics, distributed tracing, alert design.
- Improve monitoring frameworks; enable faster incident detection and resolution.
- Design dashboards and alerts that help teams understand system behavior.
- Partner with application teams on service instrumentation.
AWS Cost Optimization
- Drive significant reductions in cloud spend through architectural improvements and resource utilization.
- Review infrastructure for efficiency; identify and eliminate waste.
- Balance cost, performance, and reliability in design decisions.
Software Development & Architecture
- Build production systems, APIs, internal tools, and automation with clean, well-tested code.
- Design for maintainability, operational simplicity, and reliability.
- Participate in code review and technical design discussions.
- Mentor junior engineers on code quality and architectural thinking.
Operations & Incident Response
- Participate in on-call rotations; debug and resolve production incidents.
- Conduct postmortem analysis; drive systemic improvements.
- Develop operational procedures and runbooks.
Qualifications:
- 5+ years software engineering, platform engineering, or infrastructure engineering experience.
- Strong coding in Python, Go, Java, or equivalent; writes clean, testable code.
- AWS hands-on: EC2, RDS, DynamoDB, S3, Aurora, Lambda, VPCs, Athena.
- Terraform or equivalent infrastructure-as-code experience.
- Docker and container orchestration (Kubernetes or similar).
- Debugging on Linux and Windows platforms; able to troubleshoot complex systems using logs, metrics, and architectural knowledge.
- System design thinking: can architect scalable systems and reason about trade-offs.
Availability for rotational on-call duties outside of standard business hours may be required.
Bachelor’s degree in a related discipline and 4 years’ experience in a related field. The right candidate could also have a different combination, such as a master’s degree and 2 years’ experience; a Ph.D. and up to 1 year of experience; or 16 years’ experience in a related field.
Highly Valued:
- Experience with observability tools (New Relic, Splunk, Prometheus).
- Incident response experience; familiar with postmortem practices.
- Interest in or hands-on experience with SRE concepts (SLOs, resilience, failure modes).
- Cost optimization mindset; has identified and eliminated cloud waste.
- Windows and Linux system troubleshooting and performance analysis.
- Experience with CI/CD pipelines and deployment automation.
Sr Software Engineer - Reliability Engineering in village of north hills at Unknown Company
This position is listed as full time and onsite.