All videosMultimodality2:12

Connecting representations and models

The system returns a record; writing an answer requires a separate generative stage.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

1. Visual encoder + connector + language model

The most widespread approach in recent years uses three components in sequence: a visual encoder that processes the image and produces a high-dimensional representation, a connector that…

02

2. Fusion through cross-attention

Flamingo, published by DeepMind in 2022, introduced a different approach: instead of processing the image before the text and passing its representation as input, it inserted…

03

3. Native multimodal tokenization

The third approach removes the connector boundary entirely. Instead of connecting a visual encoder to a language model through a separate module, the system discretizes images or audio…

Key moments

Jump directly to a section

  1. Two encoders can retrieve without generating
  2. A connector learns to change representations
  3. A query can select visual information
  4. Sequence position links images and questions
  5. Compressing regions can hide a detail
  6. Text and audio outputs need coordination
Read the reviewed transcript

This video has no narration. This transcript reproduces its on-screen text; it does not invent a spoken track.

Two encoders can retrieve without generating

One encoder transforms the image and another transforms the text query.

Their comparable representations allow candidates to be ordered by similarity.

The system returns a record; writing an answer requires a separate generative stage.

A connector learns to change representations

The visual encoder emits a three-component vector in this example.

A learned matrix projects it into four components accepted by the receiving model.

Changing dimensions connects interfaces but does not guarantee that all visual evidence survives.

A query can select visual information

The text query asks about the object on the right.

Attention weights visual regions and combines their representations.

Selection depends on the query; an attention map alone does not establish a causal explanation.

Sequence position links images and questions

The input interleaves a first image with its caption and then a second image.

The question refers to the second image, not an unordered mixture of both.

Keeping modality boundaries and positions preserves which evidence accompanies each fragment.

Compressing regions can hide a detail

Four visual regions contain information, including a small label in the fourth.

A two-representation summary merges or discards some detail.

Fewer tokens can reduce work, but the small label must still be tested for recoverability.

Text and audio outputs need coordination

A system can produce text and audio chunks at different rates.

Identifiers and boundaries associate each chunk with the correct content.

Coordination belongs to the system: multiple modalities do not prove the entire path is full duplex.