All videosReasoning2:01

Reasoning Risks and Production Controls

Overthinking, prompt injection and agent hijacking in reasoning models with tools. Design criteria for bounding risk in production.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

1. Overthinking: when more reasoning degrades the answer

The intuition that more reasoning always produces better answers is incorrect. There is a documented phenomenon in reasoning models called overthinking (Apple Research, 2025): the model…

02

2. Quality vs cost vs latency in a real product

The three-way tension that every generative-AI system has to manage becomes sharper in reasoning models.

03

The cost profile

In models without extended reasoning, quality failures are relatively stable across the operating distribution: the model works well on the cases it was trained for and fails on cases…

Video text and visual description

This video has no speech. The text track reproduces the written content; the descriptions below explain the visuals.

0:00 — Thinking longer has a limit.

Apple Research found three regimes as task complexity increases: on simple tasks, standard models can outperform reasoning models; at medium complexity, extra reasoning helps; and at high complexity both can collapse. Reasoning effort also stops scaling indefinitely: it rises with complexity up to a point and then falls, even with token budget available.

Visual description: Three columns for simple, intermediate and complex tasks reveal relative advantage, improvement and collapse; a note shows reasoning effort can stop increasing.

0:25 — The attacker can hide inside the data.

In an application with RAG, a browser or tools, the model processes content the user did not write. A malicious instruction can travel inside a retrieved webpage, document or API response. Greshake et al. showed that indirect prompt injection can manipulate the behavior of LLM-integrated applications and influence when or how other APIs are called.

Visual description: External content crosses a trust boundary into instructions and tools; an untrusted command creates authority confusion and reaches the API path.

0:50 — Safety can also open attack surface.

TabooRAG shows another route: a retrievable document can inject risk-relevant context into an otherwise benign query and trigger a refusal. The attack exploits shared alignment criteria to transfer blocking behavior across models. The risk here is not a harmful answer, but loss of availability.

Visual description: A benign query retrieves ordinary or contaminated documents; risk context reaches safety criteria and ends in a refusal, illustrating an availability attack.

1:12 — Budget and risk must be decided together.

Conformal Thinking reframes the reasoning budget as risk control: stop when confidence is sufficient and also cut off instances that appear unsolvable within the budget. The goal is not to spend every available token, but to satisfy a target error rate while using compute where it improves reliability.

Visual description: A conceptual confidence/risk chart shows sufficient-confidence, continue-evaluating and unsolvable regions with upper/lower stopping rules and calibration.

1:36 — Capability without control is not useful autonomy.

A robust agent limits permissions and validates external content while separating untrusted data from authoritative instructions. Before an irreversible action, it adds an explicit verification boundary. It also sets hard budgets for time, tokens and tool calls, and keeps explicit fallbacks and can abstain. The goal is to bound damage when one layer fails.

Visual description: Three barriers, Permissions, Context and Verification, bound a trajectory; hard time/token/tool budgets, a lock before irreversible action and fallback/abstention are added.