What It Means for an LLM to Reason
What reasoning means for a language model, what o1 and DeepSeek R1 added, and why evaluating reasoning requires looking at steps, cost and failure modes.
Links containing ?t= open the video at a specific second.
The ideas to retain
1. What "reasoning" means for a human
In cognitive psychology, reasoning is a deliberate process that allows people to solve problems outside the patterns they already know. It differs from pattern recognition—which is fast,…
2. What LLMs can do that looks like reasoning
LLMs can perform actions that, in practical terms, produce outputs similar to human reasoning in many contexts.
3. The emergence of reasoning models: OpenAI o1
In September 2024, OpenAI released o1, the first model explicitly designed to "think before answering." The difference from earlier models was not simply model size or training data, but…
What this video covers
This video does not yet have a reviewed synchronized transcript. This editorial context describes its content without presenting it as spoken audio or literal on-screen text.
What reasoning means for a language model, what o1 and DeepSeek R1 added, and why evaluating reasoning requires looking at steps, cost and failure modes.
- 1. What "reasoning" means for a human: In cognitive psychology, reasoning is a deliberate process that allows people to solve problems outside the patterns they already know. It differs from pattern recognition—which is fast,…
- 2. What LLMs can do that looks like reasoning: LLMs can perform actions that, in practical terms, produce outputs similar to human reasoning in many contexts.
- 3. The emergence of reasoning models: OpenAI o1: In September 2024, OpenAI released o1, the first model explicitly designed to "think before answering." The difference from earlier models was not simply model size or training data, but…
Video text and visual description
This video has no speech. The text track reproduces the written content; the descriptions below explain the visuals.
0:00 — What does it mean to “reason”?
Operationally, we care about a process that breaks problems into steps. It consumes resources and time. And it can fail systematically.
Visual description: A checkable arithmetic example shows 17 × 6 decomposed as (10 + 7) × 6 = 60 + 42 = 102. The video explicitly labels it as an exact example, not an observed internal chain of an LLM.
0:17 — Verify outcomes, reinforce strategies.
RLVR uses tasks with verifiable outcomes to turn correctness into a reward signal. DeepSeek R1 uses GRPO to compare attempts at the same problem relative to one another. At inference time, the model can spend more steps before producing the final answer. Training and inference are two different stages.
Visual description: A schematic separates TRAINING from INFERENCE. During training, a verifiable reward signal updates weights W and reinforces strategies; during inference, weights remain fixed while a reasoning budget leads to the final answer.
0:39 — More reasoning, more than one lever.
An answer can use a single sample or combine several and decide by consensus. Those are different strategies and should be measured separately. On AIME 2024, o1 scored 74% with one sample and 83% with consensus across 64 samples. That is evidence from that experiment, not a general guarantee.
Visual description: The animation separates an illustrative arithmetic set of attempts from reported o1 results on AIME 2024: 74% with one sample and 83% with 64 samples plus consensus. It explicitly notes that consensus and correctness are not equivalent.
1:02 — The label matters less than the failure curve.
Models can produce deliberate chains and still collapse when complexity or the available budget changes. Results also depend on benchmark and evaluation design. The useful practice is to measure where reasoning works, what it costs and how it fails.
Visual description: An experimental-design matrix crosses task, budget and verification with conditions A, B and C. The cells contain lines and question marks; the video states that these are experimental conditions, not measured results.


