Preserving evidence across modalities
That caption lost a spatial relation needed to answer where each object is.
Links containing ?t= open the video at a specific second.
The ideas to retain
Contents
1. The real problem: what counts as multimodality
What a modality is and why text, image, audio, video, documents and sensors behave differently. Why "turn everything into text" solves some tasks while throwing away part of the problem.
2. Alignment: from pairs to interactions
How a system learns that two different signals refer to the same object, event or context. What changes when alignment is not only image-text, but audio-text, video-audio, document-layout…
Jump directly to a section
Read the reviewed transcript
This video has no narration. This transcript reproduces its on-screen text; it does not invent a spoken track.
Naming objects does not preserve their relation
The scene contains a circle and a square, with the circle on the left.
The caption “a circle and a square” still fits after their positions are swapped.
That caption lost a spatial relation needed to answer where each object is.
The same states can form a different sequence
We see a closed, open and closed door at three moments.
Keeping only the first and last frame hides the intervening opening.
Video requires evidence of events between frames, not only of visible objects.
Cross-modal search requires correspondences
A text query asks for a triangular sign. The archive contains images of three shapes.
An aligned system compares text and image representations.
The match returns a candidate; its geometry and context still need checking.
An embedding is not a generated image
An image representation can be used to find another archive entry.
A generator instead produces a new output conditioned on the request.
The first operation retrieves an existing record; the second constructs new content.
The answer must retain its evidence
A table places the twelve-euro amount in the row for order A.
The answer needs the value and its row relation, not just recognition of the number.
Keep region, row and document identity so the claim can be checked.
Seeing an instruction does not authorize an action
A screenshot contains the text of a write request.
Reading the text and proposing a tool are different from executing the operation.
Permission is checked outside the observed content, before the external effect.


