- Own the quality strategy for teams shipping largely agent-generated code
- Build and maintain automated test infrastructure across unit, integration, end-to-end, contract, API, performance, and regression testing
- Define and operate evaluation frameworks for AI-powered features, including datasets, scoring systems, model-as-judge approaches, human calibration, regression tracking, and release gates
- Design testing strategies for non-deterministic and agentic systems
- Use AI coding agents to generate and maintain meaningful, risk-based test suites
- Test agent behavior, including multi-step task execution, tool usage, failure recovery, safety enforcement, and human handoff workflows
- Instrument production quality signals and use real-world feedback to improve evaluation and testing frameworks
- Partner with engineering teams throughout the development lifecycle, from refinement through release
Requirements
- 8+ years of industry experience in test automation or quality engineering
- Strong programming skills in Python, TypeScript, Java, or C#
- Track record supporting complex software products
- Hands-on experience testing AI and model-based systems, including evaluation frameworks and model quality assessment
- Ability to build and scale automated testing platforms, including CI/CD integration, test data strategies, framework development, and flaky test management
- Strong understanding of evaluating agentic systems, including tool calling, multi-step workflows, safety mechanisms, and failure handling
- Regular use of AI coding agents with critical review and accountability for generated software and test assets
- Experience with API testing, performance testing, and at least one of desktop, web, or distributed backend platforms
- Deep understanding of AI failure modes, including hallucinations, prompt injection, context degradation, non-determinism, and silent regressions
- Practical experience with agentic development environments, MCP or comparable tooling, and repository-based management of prompts, instructions, and evaluation assets
- Strong communication skills and ability to collaborate across distributed teams
- Passion for helping shape a modern AI-native organization
Core Competencies
Demonstrates expertise in building and maintaining automated test infrastructures, particularly for AI-powered features and agentic systems. Proficient in evaluating model quality and implementing risk-based testing strategies while collaborating effectively with engineering teams.
Highest-signal resume keywords
- Test Automation
- Python Programming
- AI Testing Frameworks
- CI/CD Integration
- Agentic Systems Evaluation
ATS Optimization Keywords
Hard Skills
- Test Automation
- Python Programming
- TypeScript Programming
- Java Programming
- C# Programming
- API Testing
- Performance Testing
- Automated Testing Platforms
- Model Quality Assessment
- Flaky Test Management
Soft Skills
- Strong Communication Skills
- Collaboration Across Distributed Teams
Industry Keywords
- Quality Engineering
- Agent-Generated Code
- Non-Deterministic Systems
- Multi-Step Workflows
- Safety Mechanisms
Tools & Technologies
- AI Coding Agents
- MCP Tooling
- Evaluation Frameworks
- Test Data Strategies
- Agentic Development Environments