All videosReasoning1:15

Reasoning Models

Five chapters on how LLMs reason: test-time compute, failure modes, latency, cost and risks when models use tools.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

Contents

02

1. What "reasoning" means for an LLM

Which definitions of reasoning can be useful in the context of language models. The emergence of explicit reasoning-model product families, with OpenAI o1 as an important inflection point.

03

2. How these systems fail

Failure modes are not purely random: shortcuts, systematic errors, objective drift and other recurring patterns. Methods for detecting and mitigating these failures.

Video text and visual description

This video has no speech. The text track reproduces the written content; the descriptions below explain the visuals.

0:00 — Reasoning is a process.

A complex problem can be broken into steps. Those steps consume time and compute. And they can fail.

Visual description: The illustrative multiplication 17 × 6 is split into 10 × 6 and 7 × 6: 60 + 42 = 102. An answer of 104 is marked incorrect. The grid represents the calculation being broken down, not a benchmark result.

0:15 — The budget is also a choice.

At inference time, you can spend more steps, samples, verification or tool interactions. On some tasks that can improve the answer, but it also increases cost and latency.

Visual description: Twelve units of compute are allocated to steps, samples and verification. This is an illustrative allocation: it shows that a budget can be distributed in different ways, not that one allocation is universally optimal.

0:35 — One series, five questions.

First: what reasoning means for an LLM. Second: how these systems fail. Third: how to use more inference-time compute. Fourth: what happens to latency and streaming. Fifth: when more reasoning introduces new risks.

Visual description: The map connects reasoning to five questions: what it means, how it fails, how to use test-time compute, what happens to latency, and what risks it introduces. Quality, cost and time are shown as related decisions.

0:58 — Thinking longer is not free.

More reasoning can improve quality on suitable tasks. It also consumes time, compute and money. And it can introduce overthinking, errors or new attack surfaces.

Visual description: A triangle connects quality, cost and latency, and risk around the task. This is a qualitative comparison: position does not represent a score or a performance measurement.