Structure, resolution and time
If that position answered the question, the reduction already removed the needed evidence.
Links containing ?t= open the video at a specific second.
The ideas to retain
1. A modality is not only an input type
We call a modality a distinct way of encoding information about the world. Text is a discrete sequence of symbols, an image has spatial structure, audio adds temporal continuity, tone,…
2. The problem is not adding modalities, but crossing them without destroying them
A system can accept an image and still remain deeply text-centric: it only has to turn the image into a caption and perform all subsequent reasoning over that caption.
3. The shared space matters, but it is not the only possible architecture
The most common version of an introduction to multimodality treats the shared representation space almost as the essence of the entire field.
Jump directly to a section
Read the reviewed transcript
This video has no narration. This transcript reproduces its on-screen text; it does not invent a spoken track.
Lower resolution can erase a difference
Two four-pixel patterns differ in the position of one active pixel.
Averaging each pattern into one cell gives the same value: one quarter.
If that position answered the question, the reduction already removed the needed evidence.
Reading numbers is not enough to read a table
The columns are price and quantity; the row contains twelve and two.
A flat list retains the numbers but can lose which column each belongs to.
The correct structure means twelve euros per unit and two units, not the reversed interpretation.
A transcript can omit how words were spoken
Two voice traces contain the same words: “it is done”.
One ends with a rising pitch contour and the other with a falling contour.
The text matches but the prosody differs; no universal emotion is assigned to either contour.
An event can fall between two samples
Video is sampled at seconds zero, two and four.
The light is on only between one and one and a half: no sample captures that interval.
Absence from the received frames does not prove the event never happened.
Synchronization requires knowing the offset
A flash and its sound belong to one event, but their recordings use different clocks.
Audio appears three tenths of a second later on the video time scale.
Correcting the offset aligns the event; uncalibrated timestamps can invent an order.
A number needs units and a reference frame
The camera reports eighty centimeters relative to its own origin.
The map uses meters and its origin is twenty centimeters behind the camera origin.
The same position is one meter in map coordinates: the reference changes, not the object.


