Unknown Company

Multimodal ML Engineer - Vision, Audio & Text AI

mountain view, ca • Posted 5 days ago
Onsite Full Time Electrical & Energy Engineering

Google DeepMind is seeking a Machine Learning Engineer focused on multimodal models in Mountain View, CA. You will build and optimize models that fuse vision, audio, text, and video data, targeting low latency and high throughput at scale.

You should have 3+ years of industry experience with Python and deep learning frameworks (TensorFlow or PyTorch), plus a strong grasp of CV, NLP, and multimodal fusion techniques. Excellent collaboration skills are required.

#J-18808-Ljbffr

Multimodal ML Engineer - Vision, Audio & Text AI in mountain view at Unknown Company

This position is listed as full time and onsite.

Back to Job Search