All videosReasoning1:22

What It Means for an LLM to Reason

What reasoning means for a language model, what o1 and DeepSeek R1 added, and why evaluating reasoning requires looking at steps, cost and failure modes.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

1. What "reasoning" means for a human

In cognitive psychology, reasoning is a deliberate process that allows people to solve problems outside the patterns they already know. It differs from pattern recognition—which is fast,…

02

2. What LLMs can do that looks like reasoning

LLMs can perform actions that, in practical terms, produce outputs similar to human reasoning in many contexts.

03

3. The emergence of reasoning models: OpenAI o1

In September 2024, OpenAI released o1, the first model explicitly designed to "think before answering." The difference from earlier models was not simply model size or training data, but…

Text context for search and agents

What this video covers

This video does not yet have a reviewed synchronized transcript. This editorial context describes its content without presenting it as spoken audio or literal on-screen text.

What reasoning means for a language model, what o1 and DeepSeek R1 added, and why evaluating reasoning requires looking at steps, cost and failure modes.

  • 1. What "reasoning" means for a human: In cognitive psychology, reasoning is a deliberate process that allows people to solve problems outside the patterns they already know. It differs from pattern recognition—which is fast,…
  • 2. What LLMs can do that looks like reasoning: LLMs can perform actions that, in practical terms, produce outputs similar to human reasoning in many contexts.
  • 3. The emergence of reasoning models: OpenAI o1: In September 2024, OpenAI released o1, the first model explicitly designed to "think before answering." The difference from earlier models was not simply model size or training data, but…

Video text and visual description

This video has no speech. The text track reproduces the written content; the descriptions below explain the visuals.

0:00 — What does it mean to “reason”?

Operationally, we care about a process that breaks problems into steps. It consumes resources and time. And it can fail systematically.

Visual description: A checkable arithmetic example shows 17 × 6 decomposed as (10 + 7) × 6 = 60 + 42 = 102. The video explicitly labels it as an exact example, not an observed internal chain of an LLM.

0:17 — Verify outcomes, reinforce strategies.

RLVR uses tasks with verifiable outcomes to turn correctness into a reward signal. DeepSeek R1 uses GRPO to compare attempts at the same problem relative to one another. At inference time, the model can spend more steps before producing the final answer. Training and inference are two different stages.

Visual description: A schematic separates TRAINING from INFERENCE. During training, a verifiable reward signal updates weights W and reinforces strategies; during inference, weights remain fixed while a reasoning budget leads to the final answer.

0:39 — More reasoning, more than one lever.

An answer can use a single sample or combine several and decide by consensus. Those are different strategies and should be measured separately. On AIME 2024, o1 scored 74% with one sample and 83% with consensus across 64 samples. That is evidence from that experiment, not a general guarantee.

Visual description: The animation separates an illustrative arithmetic set of attempts from reported o1 results on AIME 2024: 74% with one sample and 83% with 64 samples plus consensus. It explicitly notes that consensus and correctness are not equivalent.

1:02 — The label matters less than the failure curve.

Models can produce deliberate chains and still collapse when complexity or the available budget changes. Results also depend on benchmark and evaluation design. The useful practice is to measure where reasoning works, what it costs and how it fails.

Visual description: An experimental-design matrix crosses task, budget and verification with conditions A, B and C. The cells contain lines and question marks; the video states that these are experimental conditions, not measured results.