Chapter 2 — Alignment: from pairs to interactions¶
This article explains how multimodal systems learn that two different signals refer to the same content: the alignment process. It covers the data needed to build shared representations, what changes when alignment extends beyond image–text pairs (as ImageBind does across six modalities), and why training-data quality matters more than architecture in determining where a system fails in production. It is useful whether you have a technical background in machine learning or simply want to understand why models handle some images well and fail systematically on others.
The previous chapter established that multimodality is not limited to text–image pairs or to a single shared representation space. This chapter explains how alignment between modalities is built: what data is needed, how the problem changes when there are more than two modalities, and what happens when the data is low quality or poorly structured.
This sequence reflects how multimodal systems are actually trained. Each stage builds on the previous one, so failures that appear late in the pipeline often originate in weaknesses introduced earlier.
1. Image–text pairs: the foundation and its limits¶
The simplest starting point is a set of image–text pairs that belong together: a photograph with its caption, a product image with its description, or a diagram next to its legend. The internet contains billions of these pairs, and foundational vision–language models such as CLIP or ALIGN were trained by extracting them at scale CLIP.
Training on these pairs uses contrastive learning: the model projects image and text into the same representation space so that matching pairs end up close together and non-matching pairs end up far apart. The loss penalizes the model when it associates an image with another image's description and rewards it when the representations of the correct pair are similar. This produces aligned visual and text encoders that can find the description most compatible with an image without having seen that exact pair during training, because the shared space captures semantic relationships rather than simply memorizing examples.
The main limitation is noisy data. A photograph's caption does not always describe precisely what appears in the image: it may refer to something that happened before or after, describe context rather than visible content, or simply be irrelevant.
When a model learns from millions of noisy pairs, the resulting representations can be robust in aggregate while remaining fragile on fine details. That weakness becomes visible when the system faces questions that require precision.
The field did not respond by abandoning contrastive learning, but by refining it. SigLIP replaced CLIP's classic loss with a sigmoid loss over independent pairs, improving stability with smaller batches and producing representations that transfer better to localization tasks SigLIP.
DINOv2 took a different direction: instead of depending on text supervision, it trains the visual encoder through self-supervision on a curated image collection, producing denser representations that capture fine spatial structure and generalize better to segmentation and visual retrieval DINOv2. Both approaches point to the same conclusion: representation quality depends not only on alignment with text, but also on the quality and richness of the visual encoder itself.
The deeper limitation is that image–text pairs leave entire modalities out. Audio, video, documents and continuous signals do not fit into that framework without additional extensions, which constrains the systems that can be built from this type of data alone.
2. Beyond the pair: aligning multiple modalities¶
A key recent finding is that multimodal alignment does not need direct pairs between every modality. ImageBind demonstrated this directly: using only pairs in which image is the common anchor (image–text, image–audio, image–depth, image–thermal, image–IMU), the system learns a shared embedding space across six modalities without ever requiring direct audio–text or audio–depth pairs ImageBind.
As a result, a text query can retrieve audio, an image can retrieve depth or thermal data, and all modalities become aligned transitively through the visual anchor.
The same alignment pattern now appears in production embedding systems. Gemini Embedding 2 treats multimodal embedding as a native primitive that unifies text, images, video, audio and documents in a single representation space usable for cross-modal search, classification and clustering Gemini Embedding 2. The shift is qualitative, not merely support for more modalities: the embedding is no longer a by-product of an understanding model, but the central object of the system.
Going beyond text–image pairs also changes the structure of the training data. For vision–language, the internet provided billions of natural pairs. For audio–text, video–text or document-layout, high-quality pairs are much scarcer, noisier and more dependent on human work or controlled synthesis. That asymmetry in data availability largely explains why system capabilities are asymmetric: models understand images better than audio, and audio better than documents with complex layout.
anchor
3. Visual instruction: the next level¶
Image–text pairs, and their extensions to other modalities, train representations but do not teach the model to follow instructions. For a system to answer "what anomalies are in this chart?" or "transcribe the text in this image and correct it," it needs additional training on visual-instruction data.
Visual-instruction data consists of triples: image, textual instruction and expected response. The model learns to condition jointly on the image and the instruction to generate the correct response, which is the format used by models such as LLaVA or InstructBLIP.
Generating high-quality instruction data is significantly more expensive than extracting pairs from the internet because it requires either human annotation, which is expensive and slow, or synthetic generation with powerful language models that receive a description of the image and produce plausible instructions and responses. Synthetic generation scales, but it also imports the generator model's biases: if the model generating the data has blind spots, the model trained on that data can inherit them.
LLaVA, published in 2023, showed that synthetic visual-instruction data generated with GPT-4 could produce strong visual instruction following at comparatively low data-generation cost LLaVA. The approach has become common practice for projects without the budget for large-scale human annotation, although the final quality remains bounded by the quality of the model generating the synthetic data.
4. Why data quality dominates¶
One of the most consistent lessons from multimodal systems is that training-data quality strongly determines representation robustness. A model with a suboptimal architecture trained on high-quality data tends to outperform a state-of-the-art architecture trained on noisy data, at least on tasks that the better data covers well.
This dependence on data quality has two consequences for how published results should be interpreted.
First, multimodal evaluation benchmarks are often diagnostically incomplete. A model can score highly on image-description tasks while failing on localization or verification simply because its training data emphasized the first and poorly covered the second. The training-data distribution therefore appears directly in the model's capability profile.
Second, weaknesses compound through the training pipeline. If image–text pretraining produces representations in which certain image types are only weakly associated with their correct descriptions, later visual-instruction tuning cannot repair that problem from scratch because it builds on the representations it receives, including their strengths and gaps.
Radford et al. documented this pattern when analyzing CLIP failures on image categories underrepresented in the training data CLIP: the model generalized well on common categories and systematically worse on infrequent ones, even when image quality was equivalent. Fixing the problem required rebalancing the data, not changing the architecture.
5. The role of alignment with human preferences¶
Beyond supervised training, more recent multimodal models include a phase of alignment with human preferences, analogous to RLHF in language models. Human evaluators compare model responses to visual questions and indicate which is better, so the model learns to generate responses that people consider useful, correct and aligned with their expectations.
This phase captures something that pure supervised training cannot measure directly: subjective preferences about how the model should describe what it sees, how much detail is appropriate for different questions, and how to balance precision with readability.
The risk is that evaluator preferences are not uniform and can introduce cultural, gender or aesthetic biases that become encoded in the model. If evaluators tend to prefer longer, more elaborate descriptions, the model will learn to produce longer answers regardless of whether that length is appropriate for the question. The bias is not in the architecture or the visual data, but in who evaluates and which criteria they apply, which makes it difficult to detect with standard benchmarks and easier to observe in real use.
Next chapter
Chapter 3 — Architectures → — The four multimodal architecture families, their differences in quality, cost and latency, and why multimodal embedding and multimodal generation are not the same layer of the system.
6. References¶
Core sources
| Key | Source | Short description |
|---|---|---|
| R1 | Radford et al. (2021) — Learning Transferable Visual Models From Natural Language Supervision (arXiv) | CLIP and contrastive learning at scale. |
| R2 | Liu et al. (2023) — Visual Instruction Tuning (arXiv) | LLaVA: synthetic visual-instruction data generation with GPT-4. |
| R3 | Li et al. (2023) — BLIP-2: Bootstrapping Language-Image Pre-training (arXiv) | Staged training strategy for vision–language systems. |
| R4 | Jain et al. (2023) — VCoder: Versatile Vision Encoders for Multimodal Large Language Models (arXiv) | Study of how visual-encoder choice determines a multimodal system's capability profile beyond the LLM architecture. |
| R5 | Girdhar et al. (2023) — ImageBind: One Embedding Space To Bind Them All (CVPR) | Alignment of six modalities using only pairs with image as the anchor. |
| R6 | Google DeepMind (2026) — Gemini Embedding 2 (blog) | Native multimodal embeddings across text, images, video, audio and documents. |
| R7 | Zhai et al. (2023) — Sigmoid Loss for Language Image Pre-Training (arXiv) | SigLIP: independent pairwise sigmoid loss that improves stability relative to CLIP. |
| R8 | Oquab et al. (2023) — DINOv2: Learning Robust Visual Features without Supervision (arXiv) | DINOv2: self-supervised visual encoder with dense representations and stronger spatial generalization. |
Frequently asked questions¶
Why is contrastive learning more efficient than teaching a model to describe images? Learning to match images with text that already exists on the internet is much cheaper than predicting every word of a generated description. The method forces visual and text encoders to build a shared space where similar concepts are represented by nearby vectors, without needing to generate new text or annotate images by hand.
What is the connector module between the visual encoder and the language model for? It acts as a bridge between two worlds with different representations. If the connector is too simple, it fails to transfer the richness of the visual signal to the language model. If it is too complex, it requires more data and more compute to train. The connector is the critical design point in this architecture because it determines how much visual information reaches reasoning.
What does the visual instruction tuning popularized by LLaVA involve? It trains the model on triples of image, textual instruction and expected response, so it learns to follow complex instructions about visual content instead of only describing what it sees. LLaVA showed that these data can be generated synthetically with a powerful language model, although the trained model inherits the blind spots of the model that generated them.
What is the difference between learning from image–text pairs and building a shared representation space across six modalities as ImageBind does? Image–text pairs align only those two modalities. ImageBind uses image as a common anchor and learns alignment between audio, depth, thermal signals and IMU without ever seeing direct pairs between those modalities: if audio and image are aligned, and text and image are also aligned, then audio and text become aligned transitively. The result is that a text query can retrieve audio even though they were never paired directly.