EPAM Systems is seeking an Agent Evaluation Engineer to design and maintain an evaluation framework for AI agents, including automated tests and CI/CD deployment gates. The role focuses on creating multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulating multi-turn conversations, and defining reliability metrics.
You will implement staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production
#J-18808-LjbffrRemote AI Agent Evaluation Engineer in workfromhome at Unknown Company
This position is listed as full time and onsite.