Google DeepMind is seeking a Machine Learning Engineer focused on multimodal models in Mountain View, CA. You will build and optimize models that fuse vision, audio, text, and video data, targeting low latency and high throughput at scale.
You should have 3+ years of industry experience with Python and deep learning frameworks (TensorFlow or PyTorch), plus a strong grasp of CV, NLP, and multimodal fusion techniques. Excellent collaboration skills are required.
#J-18808-LjbffrMultimodal ML Engineer - Vision, Audio & Text AI in mountain view at Unknown Company
This position is listed as full time and onsite.