Learning alignment
The contrastive objective distinguishes those batch correspondences without guaranteeing every pair is correctly labelled.
Links containing ?t= open the video at a specific second.
The ideas to retain
1. Image–text pairs: the foundation and its limits
The simplest starting point is a set of image–text pairs that belong together: a photograph with its caption, a product image with its description, or a diagram next to its legend. The…
2. Beyond the pair: aligning multiple modalities
A key recent finding is that multimodal alignment does not need direct pairs between every modality. ImageBind demonstrated this directly: using only pairs in which image is the common…
3. Visual instruction: the next level
Image–text pairs, and their extensions to other modalities, train representations but do not teach the model to follow instructions. For a system to answer "what anomalies are in this…
Jump directly to a section
Read the reviewed transcript
This video has no narration. This transcript reproduces its on-screen text; it does not invent a spoken track.
The matching pair is a training signal
A batch contains three images and their three associated captions.
Each image is compared with every text: matched pairs lie on the diagonal.
The contrastive objective distinguishes those batch correspondences without guaranteeing every pair is correctly labelled.
Similarity can depend on direction
We normalize two vectors and compare their angle using cosine similarity.
Bringing their directions closer increases positive-pair similarity.
Training also contrasts other candidates; high cosine similarity is not a probability of truth.
Changing a relation creates a hard negative
Both captions name a circle and a square.
One places the circle on the left; the other places it on the right.
Distinguishing them requires the relation, not just recognition of the object list.
A pairing can teach a false association
An image of three dots arrives paired with the caption “two dots”.
Treating that pair as positive favors a wrong correspondence.
Review corrects or excludes the item; a larger batch alone does not eliminate this noise.
Aligning a clip does not locate every event
A recording contains silence, a knock and another pause.
A clip label says a knock is present but not when it starts or ends.
A temporal annotation binds the event to the relevant interval with explicit boundaries.
A shared anchor does not replace evaluation
Image–text and image–audio pairs are trained around a shared concept.
The image can connect representations that were not directly paired.
Test text–audio retrieval: a shared anchor does not guarantee equal quality for every relation.


