Unknown Company

Remote LLM Evaluation Scientist: Benchmarking Models

Remote • Posted 4 days ago
Remote Full Time Other

Anyone AI Labs is seeking a Research Scientist for LLM Evaluations and Benchmarking. This remote role spans LatAm/US and requires designing robust evaluation methodologies for frontier models, building benchmarks across reasoning, coding, agents, tool use, and multi-modal tasks.

The candidate will lead expert pools, validate ground truth, and publish results in venues like NeurIPS, ICLR, and ACL, with strong English proficiency and preference for Spanish speakers.

#J-18808-Ljbffr
Back to Job Search