---
title: Multimodal System Architectures
seo_title: Multimodal System Architectures — video
description: Four multimodal architecture families, their differences in quality, cost and latency, and when each way of combining modalities makes sense.
keywords: multimodal architectures, early fusion late fusion, ViT, multimodal encoder, multimodal decoder, LLaVA, GPT-4V, generative AI architecture, vision transformer
date: '2026-04-02T00:00:00+00:00'
robots: index,follow,max-snippet:-1,max-image-preview:large,max-video-preview:-1
hide:
- toc
- navigation
video_watch_page: true
---


<div class="s5-video-watch" data-s5-video-watch data-video-id="videos-series-multimodalidad-iag-03-arquitecturas-md">
  <header class="s5-video-watch__header">
    <div class="s5-video-watch__crumbs"><a href="https://5sigmas.com/en/videos/">All videos</a><span>Multimodality</span><span>1:02</span></div>
    <h1>Multimodal System Architectures</h1><p>Four multimodal architecture families, their differences in quality, cost and latency, and when each way of combining modalities makes sense.</p>
  </header>
  <div class="s5-video-watch__player">
    <video controls crossorigin="anonymous" preload="metadata" poster="/en/series/multimodalidad-iag/03-arquitecturas.jpg" playsinline data-s5-watch-player><source src="/en/series/multimodalidad-iag/03-arquitecturas.mp4" type="video/mp4">Your browser does not support the video element.</video>
    <p>Links containing <code>?t=</code> open the video at a specific second.</p>
  </div>
  <section class="s5-video-watch__summary" aria-labelledby="video-summary-title">
    <div class="s5-video-watch__section-head"><span class="s5-eyebrow">Video summary</span><h2 id="video-summary-title">The ideas to retain</h2></div>
    <div class="s5-video-watch__snippet-grid"><article><span>01</span><h2>1. Visual encoder + connector + language model</h2><p>The most widespread approach in recent years consists of three chained components: a visual encoder that processes the image and produces a high-dimensional representation, a connection…</p></article>
<article><span>02</span><h2>2. Fusion through cross-attention</h2><p>Flamingo, published by DeepMind in 2022, introduced a different approach: instead of processing the image before the text and passing its representation as input, it inserted…</p></article>
<article><span>03</span><h2>3. Native multimodal tokenization</h2><p>The third approach is the most radical: instead of connecting a visual encoder to a language model through some type of connector, the system discretizes images or audio into tokens of the…</p></article></div>
  </section>
  
  
  <aside class="s5-video-watch__source"><div><span class="s5-eyebrow">Context and evidence</span><h2>Continue with the full article</h2><p>The chapter develops the mechanism, primary sources, limitations and connections to the rest of the series.</p></div><a class="s5-video-watch__source-link" href="https://5sigmas.com/en/series/multimodalidad-iag/03-arquitecturas/">Read the article →</a></aside>
  <section class="s5-video-watch__related" aria-labelledby="related-videos-title"><div class="s5-video-watch__section-head"><span class="s5-eyebrow">Next step</span><h2 id="related-videos-title">Related videos</h2></div><div class="s5-video-watch__related-grid"><article><a href="https://5sigmas.com/en/videos/series/multimodalidad-iag/05-riesgos/"><img src="https://5sigmas.com/en/series/multimodalidad-iag/05-riesgos.jpg" alt="" loading="lazy" width="1280" height="720"><span>Multimodality · 1:02</span><strong>Risks of Multimodal AI Systems</strong></a></article>
<article><a href="https://5sigmas.com/en/videos/series/multimodalidad-iag/04-evaluacion/"><img src="https://5sigmas.com/en/series/multimodalidad-iag/04-evaluacion.jpg" alt="" loading="lazy" width="1280" height="720"><span>Multimodality · 1:02</span><strong>Evaluating Multimodal Systems</strong></a></article>
<article><a href="https://5sigmas.com/en/videos/series/multimodalidad-iag/02-alineamiento/"><img src="https://5sigmas.com/en/series/multimodalidad-iag/02-alineamiento.jpg" alt="" loading="lazy" width="1280" height="720"><span>Multimodality · 1:02</span><strong>Alignment: From Pairs to Interactions</strong></a></article></div></section>
</div>
