All videosMultimodality2:13

Structure, resolution and time

If that position answered the question, the reduction already removed the needed evidence.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

1. A modality is not only an input type

We call a modality a distinct way of encoding information about the world. Text is a discrete sequence of symbols, an image has spatial structure, audio adds temporal continuity, tone,…

02

2. The problem is not adding modalities, but crossing them without destroying them

A system can accept an image and still remain deeply text-centric: it only has to turn the image into a caption and perform all subsequent reasoning over that caption.

03

3. The shared space matters, but it is not the only possible architecture

The most common version of an introduction to multimodality treats the shared representation space almost as the essence of the entire field.

Key moments

Jump directly to a section

  1. Lower resolution can erase a difference
  2. Reading numbers is not enough to read a table
  3. A transcript can omit how words were spoken
  4. An event can fall between two samples
  5. Synchronization requires knowing the offset
  6. A number needs units and a reference frame
Read the reviewed transcript

This video has no narration. This transcript reproduces its on-screen text; it does not invent a spoken track.

Lower resolution can erase a difference

Two four-pixel patterns differ in the position of one active pixel.

Averaging each pattern into one cell gives the same value: one quarter.

If that position answered the question, the reduction already removed the needed evidence.

Reading numbers is not enough to read a table

The columns are price and quantity; the row contains twelve and two.

A flat list retains the numbers but can lose which column each belongs to.

The correct structure means twelve euros per unit and two units, not the reversed interpretation.

A transcript can omit how words were spoken

Two voice traces contain the same words: “it is done”.

One ends with a rising pitch contour and the other with a falling contour.

The text matches but the prosody differs; no universal emotion is assigned to either contour.

An event can fall between two samples

Video is sampled at seconds zero, two and four.

The light is on only between one and one and a half: no sample captures that interval.

Absence from the received frames does not prove the event never happened.

Synchronization requires knowing the offset

A flash and its sound belong to one event, but their recordings use different clocks.

Audio appears three tenths of a second later on the video time scale.

Correcting the offset aligns the event; uncalibrated timestamps can invent an order.

A number needs units and a reference frame

The camera reports eighty centimeters relative to its own origin.

The map uses meters and its origin is twenty centimeters behind the camera origin.

The same position is one meter in map coordinates: the reference changes, not the object.