AuraOne seeks a remote contractor to evaluate multi-turn ranking reward model evaluation prompts and responses. You will compare outputs, assign severity tags, and provide structured feedback to retrain the model using AuraOne's rubric.
You will identify hallucinations, instruction-following failures, and unsafe content, while calibrating against gold-standard examples weekly and reporting recurring issues to improve future rubrics.
#J-18808-Ljbffr