All videosMultimodality1:02

Multimodality in Generative AI

What it means to build systems that can perceive, align, reason, generate and act across text, image, audio, video, documents and other signals from the world.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

Contents

02

1. The real problem: what counts as multimodality

What a modality is and why text, image, audio, video, documents and sensors behave differently. Why "turn everything into text" solves some tasks while throwing away part of the problem.

03

2. Alignment: from pairs to interactions

How a system learns that two different signals refer to the same object, event or context. What changes when alignment is not only image-text, but audio-text, video-audio, document-layout…