All videosReasoning1:55

Latency, Streaming and Product Design

TTFT, streaming and perceived-latency thresholds in reasoning models. RouteLLM, design patterns and production session-cost management.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

1. Perceived-latency thresholds

User-experience research has studied for decades how latency affects perception. Jakob Nielsen established a scale in the 1990s that remains relevant:

02

Dynamic routing: RouteLLM

One relevant pattern is dynamic routing, exemplified by RouteLLM (Ong et al., 2024): instead of applying the most capable—and slowest—model to every request, a lightweight classifier…

03

2. Streaming and the perception of latency

Streaming—sending tokens to the client as they are generated instead of waiting for the complete response—is the most widely used tool for improving perceived latency without reducing…

Video text and visual description

This video has no speech. The text track reproduces the written content; the descriptions below explain the visuals.

0:00 — Waiting is also part of the system.

Three classic thresholds help design waiting: 0.1 seconds feels immediate, 1 second preserves flow and 10 seconds marks the approximate limit of continuous attention. These are interaction references, not universal laws. Above one second, the interface should show that work continues; above ten, it needs visible progress and a clear way to interrupt.

Visual description: A 0.1 s, 1 s and 10 s scale marks interaction references; activity and interruption controls then appear for longer waits.

0:27 — Starting early is not finishing early.

TTFT measures from sending the request until the first visible token appears. Total latency ends when the last token arrives. Two systems with the same total time can feel different if one starts responding sooner. Streaming reduces silent waiting; it does not remove the remaining work.

Visual description: Two bars compare systems with equal total duration but different TTFT, separating the first visible token from the last token.

0:49 — Make progress visible.

Before the first token, work can happen that the user cannot see. After that, streaming delivers the answer progressively while generation continues. The improvement is perceptual: the user gets a sign of life earlier. Total time can still remain unchanged.

Visual description: A response card moves from pre-output waiting to progressively visible text and a progress line, showing feedback without claiming lower total time.

1:09 — Not every request needs the maximum budget.

RouteLLM learns when to use a stronger model and when to use a weaker one to balance quality and cost. In its evaluation, some routers reduced cost by more than two times without degrading measured quality. In a product, that separation can map fast routes to simple queries and reserve more capacity for difficult ones. Latency policy remains a system decision.

Visual description: A router branches a query toward lower cost or more capability to illustrate dynamic model and budget allocation.

1:31 — Latency as policy, not as an accident.

The robust pattern is not to always think longer, but to classify the request, assign a budget and show progress before waiting looks like failure. When the budget is exhausted, the system needs an explicit fallback: partial answer, clarification, alternate model or recoverable error. Knowing when to stop is part of the design.

Visual description: A state machine separates Running, Complete and Budget exhausted, then lists explicit fallbacks so waiting does not become indefinite.