- Build Canoe's Site Reliability Engineering practice from the ground up, including SLOs and error budgets on critical services, observability and alerting standards, production readiness reviews, and on-call design
- Own incident management across Canoe end to end, including process, tooling, escalation paths, retrospective quality, and follow-through on corrective actions
- Lead the Quality Engineering organization and build lasting in-house quality leadership and expertise
- Define a risk-based quality strategy, including quality gates in CI/CD pipelines
- Continue the shift from manual verification toward automation-first, AI-assisted testing, and specification-driven development
- Define and report executive-level KPIs for quality and reliability, such as escape rates, incident trends, and error-budget consumption
- Use quality and reliability KPIs to inform engineering priorities across teams
- Serve as Technical Governance Owner for testing and incident management topics
- Partner with architects to bring reliability, testability, and operability perspectives to technical decisions
- Report to the VP of Application Development as a peer of other engineering guild leads
Requirements
- Minimum of 10 years of software engineering experience
- At least 4 years leading quality, reliability, or platform engineering teams
- Experience building an SRE, reliability, or quality engineering practice from scratch
- Experience formally owning incident management for a company or major platform, from incident response through high-quality retrospectives and follow-through
- Deep quality engineering background, including ownership of quality strategy and experience leading distributed or offshore QA teams
- Demonstrated expertise applying LLMs and agentic tools such as Claude Code to real engineering workflows
- Fluency with quality and reliability metrics such as escape rate, flake rate, MTTD/MTTR, and SLO attainment, including experience reporting them at the executive level
- Excellent leadership, communication, and interpersonal skills with geographically distributed teams across multiple time zones
- Ability to work effectively with ambiguous or missing information and adapt quickly in a startup-like environment
- Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience)
- Preferred: hands-on experience with Python, TypeScript, PHP/Laravel, Kafka, AWS, Terraform, PostgreSQL, Redis, ElasticSearch, or Datadog
- Preferred: fintech or other environments where data integrity and client trust are the product
Core Competencies
Demonstrates extensive experience in building and leading Site Reliability Engineering and Quality Engineering practices, with a strong focus on incident management, quality strategy, and automation-first testing methodologies. Proficient in defining and reporting quality and reliability metrics to inform engineering priorities and drive organizational excellence.
Highest-signal resume keywords
- Site Reliability Engineering Leadership
- Incident Management Ownership
- Quality Engineering Strategy Development
- Automation-First Testing Methodologies
- Quality and Reliability Metrics Reporting
Hard Skills
- Software Engineering
- Quality Engineering
- Incident Management
- Automation Testing
- SLO Development
- CI/CD Pipeline Quality Gates
- LLMs and Agentic Tools
- KPI Reporting
- Risk-Based Quality Strategy
- Distributed QA Team Leadership
Soft Skills
- Leadership
- Communication
- Interpersonal Skills
- Adaptability
- Problem-Solving
Industry Keywords
- Fintech
- Data Integrity
- Client Trust
- Quality Metrics
- Incident Trends
Tools & Technologies
- Python
- TypeScript
- PHP/Laravel
- Kafka
- AWS
- Terraform
- PostgreSQL
- Redis
- ElasticSearch
- Datadog
Director of Engineering – Quality & Reliability in new york at Unknown Company
This position is listed as full time and onsite.