Pinterest Labs is seeking researchers to advance vision-centric large language models. You will help build VLMs that perceive visual details and aesthetics, enabling communication through visual assets with multimodal search and text-to-image tools. The role supports collaboration across a core visual pod and production-ready research.
You will contribute to evaluation benchmarks, work with Pinterest Canvas data for RLHF and fine-tuning, and publish findings to the broader research community.
#J-18808-LjbffrVision AI Engineer for Multimodal LLMs (Equity Eligible) in san francisco at Unknown Company
This position is listed as full time and onsite.