Unknown Company

Research Scientist - Multimodal Generative AI

san francisco, ca • Posted 4 weeks ago
Onsite Full Time Bio & Pharmacology & Health

As a Research Scientist , you'll lead cutting-edge research that advances the state of generative AI for long-form storytelling. You'll work at the intersection of LLMs, multimodal AI, agentic systems, and scalable machine learning, translating novel research into production systems that delight millions of listeners worldwide.

Drive original research on large language models and multimodal foundation models for narrative generation, dialogue synthesis, controllable generation, planning, and creative reasoning. Explore novel model architectures, training paradigms, and alignment techniques that push beyond the current state of the art.

Advance Fine-Tuning & Model Adaptation

Design new approaches for adapting foundation models (e.g., Llama, Mistral, Qwen) using techniques such as PEFT, RLHF, DPO, preference optimization, continual learning, and domain adaptation. Improve model quality, efficiency, and factual consistency for long-form storytelling applications.

Develop novel methods that unify text, speech, audio, and other modalities into coherent generative systems capable of producing immersive storytelling experiences. Contribute to research in expressive generation, multimodal reasoning, and cross-modal representation learning.

Invent Agentic AI Systems

Conduct research on planning agents, tool-using language models, multi-agent collaboration, memory, reasoning, and autonomous creative workflows. Design new algorithms that improve reliability, controllability, and long-horizon generation.

Build Novel Evaluation Frameworks

Design robust benchmarks and evaluation methodologies for creative AI, combining automated metrics with human preference modeling. Establish rigorous research practices for measuring storytelling quality, coherence, creativity, factual consistency, and user engagement.

Publish and Influence the Research Community

Stay at the forefront of AI research by reading, reproducing, and extending work from leading conferences including NeurIPS, ICML, ICLR, ACL, EMNLP and CVPR. Contribute novel ideas through patents, publications, open-source contributions, and internal research initiatives wherever appropriate.

Translate Research into Production Impact

Partner closely with Research Engineers, Applied Scientists, and Product teams to transform research breakthroughs into scalable production systems that power Pocket FM's AI ecosystem.

The Ideal Candidate

We're looking for researchers who enjoy solving open-ended scientific problems and seeing their ideas impact products used by millions.

  • PhD in Computer Science, Artificial Intelligence, Machine Learning, NLP, Speech, Multimodal AI, or a closely related field.
  • Exceptional candidates with equivalent research experience and an outstanding publication record will also be considered.

Strong Research Track Record

  • Publications in leading conferences such as NeurIPS , ICML , ICLR , ACL , EMNLP , NAACL , CVPR , ICCV , ECCV , or equivalent research contributions.
  • Demonstrated ability to independently define research problems and deliver impactful solutions.

Expertise in Foundation Models

  • Deep understanding of modern language models, representation learning, transformer architectures, scaling laws, alignment techniques, and efficient adaptation methods such as LoRA, QLoRA, RLHF, DPO, or related approaches.
  • Research experience spanning multiple modalities including text, speech, audio, vision, or video, with a strong understanding of multimodal representation learning and generation.
  • Experience designing reproducible experiments, benchmarking systems, ablation studies, and statistically rigorous evaluations.
  • Production Mindset
  • Able to bridge the gap between research and deployment by working closely with engineering teams to productionize promising ideas.

Technical Expertise

Required

  • Strong programming skills in Python and deep expertise in PyTorch
  • Experience with the Hugging Face ecosystem (Transformers, Accelerate, PEFT)
  • Hands-on experience training or adapting large-scale foundation models
  • Experience with distributed training frameworks such as FSDP , DeepSpeed , or Megatron-LM
  • Strong understanding of optimization, large-scale training, and model evaluation

Preferred

  • Research experience in Agentic AI , reasoning systems, planning, or autonomous AI workflows
  • Experience with Retrieval-Augmented Generation (RAG)
  • Experience designing evaluation frameworks for generative AI
  • Familiarity with orchestration frameworks such as LangGraph or Ray
  • Experience deploying research systems to large-scale production environments

Why Join Pocket FM?

  • Build the Future of AI-Native Entertainment
  • Help define how millions of people consume stories by advancing the frontier of generative AI.
  • Research That Ships
  • Unlike traditional research labs, your work won't remain in papers—it will directly power products used globally.
  • Solve Open Research Problems
  • Work on challenges involving long-context reasoning, creative generation, multimodal intelligence, personalization, agentic systems, and large-scale AI infrastructure.
  • Collaborate with World-Class Researchers
  • Join a team of researchers and engineers working across LLMs, multimodal AI, reinforcement learning, speech, and creative AI.
  • Your research will shape the next generation of storytelling experiences for one of the world's largest AI-powered entertainment platforms.

#J-18808-Ljbffr

Research Scientist - Multimodal Generative AI in san francisco at Unknown Company

This position is listed as full time and onsite.

Back to Job Search