All videosMultimodality2:10

Preserving evidence across modalities

That caption lost a spatial relation needed to answer where each object is.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

Contents

02

1. The real problem: what counts as multimodality

What a modality is and why text, image, audio, video, documents and sensors behave differently. Why "turn everything into text" solves some tasks while throwing away part of the problem.

03

2. Alignment: from pairs to interactions

How a system learns that two different signals refer to the same object, event or context. What changes when alignment is not only image-text, but audio-text, video-audio, document-layout…

Key moments

Jump directly to a section

  1. Naming objects does not preserve their relation
  2. The same states can form a different sequence
  3. Cross-modal search requires correspondences
  4. An embedding is not a generated image
  5. The answer must retain its evidence
  6. Seeing an instruction does not authorize an action
Read the reviewed transcript

This video has no narration. This transcript reproduces its on-screen text; it does not invent a spoken track.

Naming objects does not preserve their relation

The scene contains a circle and a square, with the circle on the left.

The caption “a circle and a square” still fits after their positions are swapped.

That caption lost a spatial relation needed to answer where each object is.

The same states can form a different sequence

We see a closed, open and closed door at three moments.

Keeping only the first and last frame hides the intervening opening.

Video requires evidence of events between frames, not only of visible objects.

Cross-modal search requires correspondences

A text query asks for a triangular sign. The archive contains images of three shapes.

An aligned system compares text and image representations.

The match returns a candidate; its geometry and context still need checking.

An embedding is not a generated image

An image representation can be used to find another archive entry.

A generator instead produces a new output conditioned on the request.

The first operation retrieves an existing record; the second constructs new content.

The answer must retain its evidence

A table places the twelve-euro amount in the row for order A.

The answer needs the value and its row relation, not just recognition of the number.

Keep region, row and document identity so the claim can be checked.

Seeing an instruction does not authorize an action

A screenshot contains the text of a write request.

Reading the text and proposing a tool are different from executing the operation.

Permission is checked outside the observed content, before the external effect.