All videosMultimodality2:11

Learning alignment

The contrastive objective distinguishes those batch correspondences without guaranteeing every pair is correctly labelled.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

1. Image–text pairs: the foundation and its limits

The simplest starting point is a set of image–text pairs that belong together: a photograph with its caption, a product image with its description, or a diagram next to its legend. The…

02

2. Beyond the pair: aligning multiple modalities

A key recent finding is that multimodal alignment does not need direct pairs between every modality. ImageBind demonstrated this directly: using only pairs in which image is the common…

03

3. Visual instruction: the next level

Image–text pairs, and their extensions to other modalities, train representations but do not teach the model to follow instructions. For a system to answer "what anomalies are in this…

Key moments

Jump directly to a section

  1. The matching pair is a training signal
  2. Similarity can depend on direction
  3. Changing a relation creates a hard negative
  4. A pairing can teach a false association
  5. Aligning a clip does not locate every event
  6. A shared anchor does not replace evaluation
Read the reviewed transcript

This video has no narration. This transcript reproduces its on-screen text; it does not invent a spoken track.

The matching pair is a training signal

A batch contains three images and their three associated captions.

Each image is compared with every text: matched pairs lie on the diagonal.

The contrastive objective distinguishes those batch correspondences without guaranteeing every pair is correctly labelled.

Similarity can depend on direction

We normalize two vectors and compare their angle using cosine similarity.

Bringing their directions closer increases positive-pair similarity.

Training also contrasts other candidates; high cosine similarity is not a probability of truth.

Changing a relation creates a hard negative

Both captions name a circle and a square.

One places the circle on the left; the other places it on the right.

Distinguishing them requires the relation, not just recognition of the object list.

A pairing can teach a false association

An image of three dots arrives paired with the caption “two dots”.

Treating that pair as positive favors a wrong correspondence.

Review corrects or excludes the item; a larger batch alone does not eliminate this noise.

Aligning a clip does not locate every event

A recording contains silence, a knock and another pause.

A clip label says a knock is present but not when it starts or ends.

A temporal annotation binds the event to the relevant interval with explicit boundaries.

A shared anchor does not replace evaluation

Image–text and image–audio pairs are trained around a shared concept.

The image can connect representations that were not directly paired.

Test text–audio retrieval: a shared anchor does not guarantee equal quality for every relation.