Multimodality in Generative AI
What it means to build systems that can perceive, align, reason, generate and act across text, image, audio, video, documents and other signals from the world.
Links containing ?t= open the video at a specific second.
Video summary
The ideas to retain
Contents
1. The real problem: what counts as multimodality
What a modality is and why text, image, audio, video, documents and sensors behave differently. Why "turn everything into text" solves some tasks while throwing away part of the problem.
2. Alignment: from pairs to interactions
How a system learns that two different signals refer to the same object, event or context. What changes when alignment is not only image-text, but audio-text, video-audio, document-layout…


