Cincinnatus LLC is seeking specialists to red-team frontier AI models, crafting multi-step tasks and benchmarks to reveal vulnerabilities and edge cases. You will work in a collaborative, research-focused environment, designing, testing, and documenting high-quality evaluation tasks for leading AI systems.
The role supports remote, full-time engagement within the United States, with a flexible weekly cadence around 35 hours and close coordination with researchers to iteratively improve
#J-18808-Ljbffr