Video library

One technical idea per video. Full evidence one click away.

Explore 76 native-English explanations published on 5sigmas. Every watch page connects the short explanation to its original chapter and related material.

76 videos available

2:11

Multimodality · 2:11

Learning alignment

The contrastive objective distinguishes those batch correspondences without guaranteeing every pair is correctly labelled.

1. Image–text pairs: the foundation and its limits · 2. Beyond the pair: aligning multiple modalities
2:13

Multimodality · 2:13

Structure, resolution and time

If that position answered the question, the reduction already removed the needed evidence.

1. A modality is not only an input type · 2. The problem is not adding modalities, but crossing them without destroying them
2:15

LLM inference · 2:15

Measure the complete system

These are different experiments: the second can reduce arrivals precisely when latency increases.

A benchmark is a protocol, not a scalar · The workload should resemble the problem you are trying to solve
2:14

LLM inference · 2:14

Decide under constraints

Selection optimises within the eligible set. A cheap but incompatible alternative is not a valid option.

Policy starts with what it cannot choose · Model routing chooses before the first attempt
2:15

LLM inference · 2:15

Avoid and propose work

The continuation must still be generated. A prefix hit does not retrieve a complete answer.

Start from the latency budget again · Prefix caching reuses state from a previously computed prefix
2:15

LLM inference · 2:15

Precision and distribution

The error is 0.12. Lower precision requires task evaluation, not just memory accounting.

Quantization does not mean that the whole model has one precision · The basic operation introduces representation error
2:13

LLM inference · 2:13

The request clock

Long inputs and long outputs add work in different phases. One clock does not distinguish them.

The request changes shape after prefill · TTFT is a client-visible boundary, not one operation
2:14

Context engineering · 2:14

Retrieval and grounding

The assembler checks access, revision and authority before using them together.

1. Retrieval is not context assembly · 2. Lexical, semantic, and structured retrieval solve different problems
2:15

Context engineering · 2:15

Budget and provenance

That leaves 48 thousand for dynamic information. This is not a universal model property.

The maximum context window is not your operational budget · A context item needs more than text
2:11

Context engineering · 2:11

What reaches the model

The decision depends on this specific input, not everything held by the application.

First, what does “context” mean here? · Prompt engineering is one part of the problem
2:19

Voice agents · 2:19

Latency budgeting

On the same clock, total waiting is 300 plus 480. It is not the model’s TTFT.

A latency metric needs two explicit boundaries · The clock is part of the definition
2:18

Voice agents · 2:18

Turns and interruptions

VAD detects activity. End-of-turn needs evidence about the complete intent.

VAD answers “is there speech?”, not “is the thought complete?” · A timeout is a policy, not semantic proof
2:14

AI systems evaluation · 2:14

From incident to repair

After release, measure recurrence and exposure over the agreed observation window. Close the loop with production evidence and an accountable owner.

Observability and evaluation answer different questions · Preserve a reproducible system identity first
2:14

AI systems evaluation · 2:14

Evidence before expanding

A violation blocks; insufficient data keeps the rollout paused. Choose stages for the change and its risk; not every release needs them all.

Before exposing traffic, define the unit of change · Shadow: observe the candidate before giving it authority over the response
2:16

AI systems evaluation · 2:16

The result and the path

A hard violation blocks release even when other metrics improve. Save the trajectory and the verdict reason so the result is reproducible.

A trajectory is a sequence of transitions, not a list of tool names · Start with the outcome: was the task actually solved?
2:14

AI systems evaluation · 2:14

How to evaluate the evaluator

Freeze the judge version before testing it on a separate set. Accept, restrict or reject its use according to risk and evidence.

Define what "calibration" means first · Choose the narrowest grader that answers the question
2:14

AI systems evaluation · 2:14

Cases that test the system

Changing a label or grader can change the score without improving the model. Create a new version; do not silently rewrite the earlier result.

The object is an eval release, not "the dataset" · The smallest unit must be auditable
2:12

AI systems evaluation · 2:12

What are we evaluating?

A broader test checks the task, policies and real effects. Preserve versions, cases and protocol to make comparisons reproducible.

Start by defining the object, not the metric · Level 1 — Evaluate the model
1:30

AI security · 1:30

Production Controls

How least privilege, independent authorization, kill paths, observability, and evidence-bound release gates limit damage when the model fails.

Separate document reading from actions · Give every tool only the permissions it needs
1:30

AI security · 1:30

Red Teaming

How to test the full causal path from adversarial input to authorization, external effect, recovery, and a reproducible release regression.

The threat model comes before the benchmark · Separate what the model can do from what the system executes
1:30

AI security · 1:30

Poisoning

How untrusted input can become persistent memory, re-enter a later decision, and survive naive deletion through derived state.

Storing data does not make it trustworthy · Persistent memory is already a measurable attack surface
1:30

AI security · 1:30

Jailbreaks

How repeated and adaptive prompt attempts change the attack surface, and why authorization and attempt limits still matter after a refusal.

Refusing a request does not create a perfect boundary · Trying many variants changes the cost of the attack
1:30

AI security · 1:30

AI Security

How untrusted input can influence an AI system and which authorization boundaries limit whether that influence becomes an action.

1. An instruction hidden in a document can change what the system does · 2. Asking the model to ignore its limits
2:18

AI agents · 2:18

From demo to production

Reserve model decisions for the part that depends on the situation. Keep boundaries, verification and an operable fallback around that part.

Budgets before promises · Retries, idempotency, and terminal failures
2:12

AI agents · 2:12

Security and authorization

Tool restrictions must still constrain execution. Combine controls and auditing; this example does not guarantee immunity.

Direct and indirect prompt injection · Authorization must live outside the prompt
2:12

AI agents · 2:12

AI agents

Conversation can continue while the tool is working. Only a verified result supports reporting that the task finished.

Contents · 1. What an agent is—and is not
2:16

Infrastructure · 2:16

The real footprint of a data center

Reusing components can extend service without eliminating every impact. Compare alternatives for the same useful work over the complete lifecycle.

1. The comparison that calibrates the conversation · 2. Water: withdrawal, consumption and the technology that determines it
2:12

Infrastructure · 2:12

What orbital computing means

Include link coordination, failure handling and retirement at end of life. The proposal must explain who operates it, who is accountable and how it ends.

1. What it means to process data in orbit · Processing observation data
2:18

Infrastructure · 2:18

Power, heat and connectivity

Error correction and redundancy help within their limits. An unrecoverable failure needs another plan; maintenance is not free.

1. Why the cold of space does not mean free cooling · The real scale of radiators
2:14

Infrastructure · 2:14

Why now?

Change one assumption and calculate the budget again. Do not present the optimistic scenario as an achieved price.

1. The explosion in compute demand · 2. Terrestrial bottlenecks
1:30

AI security · 1:30

Prompt Injection

How an instruction hidden in a document can enter an AI system and which controls separate reading from action.

1. The problem starts when a document reaches the model · 2. The instruction can enter through retrieval
0:20

Systems engineering · 0:20

Proactive and reactive agents and tool calls

A technical walkthrough of Reactive / Proactive Agent and the conversational contract it encapsulates: immediate response, asynchronous work, and deferred completion.

The problem is not the tool, but time · The first turn accepts work; it does not promise results
2:09

Reasoning · 2:09

Test-Time Compute

Test-time compute as a second scaling axis. The three levers—more steps, more candidates and more structure—and their quality, cost and latency tradeoffs in reasoning models.

1. What test-time compute is and why it matters · 2. The three levers
1:02

Economics, energy and well-being · 1:02

AI and GDP Today

Why AI's macroeconomic impact takes time to appear in GDP, where it appears earlier, and which signals best indicate what is already changing.

1. Why macro impact takes time to arrive · The four mechanisms behind the lag
1:48

Reasoning · 1:48

How Reasoning Fails

Sycophancy, shortcut learning, specification gaming and cascading failures: the failure modes of reasoning models, how to detect them and how to mitigate them.

1. Failure types · 1.1 Shortcuts — shortcut learning
1:02

Economics, energy and well-being · 1:02

Measurement: GDP vs Well-being

Why GDP does not capture real well-being, which dimensions matter most, and when subjective well-being diverges from material well-being.

1. What GDP measures and what it leaves out · What GDP leaves out
1:22

Reasoning · 1:22

What It Means for an LLM to Reason

What reasoning means for a language model, what o1 and DeepSeek R1 added, and why evaluating reasoning requires looking at steps, cost and failure modes.

1. What "reasoning" means for a human · 2. What LLMs can do that looks like reasoning
1:15

Reasoning · 1:15

Reasoning Models

Five chapters on how LLMs reason: test-time compute, failure modes, latency, cost and risks when models use tools.

Contents · 1. What "reasoning" means for an LLM
1:02

Economics, energy and well-being · 1:02

Electricity and Well-being

Why reliable, affordable electricity enables real gains in health, logistics and industry, and why supply quality matters as much as kilowatt-hours.

1. The four main channels · Health
1:02

Economics, energy and well-being · 1:02

AI, GDP, Well-being and Energy

Quantitative analysis of AI's impact on energy, productivity and well-being using real World Bank, IEA and Penn World Table data rather than speculative projections.

Contents · 1. Electricity → well-being: the real mechanisms
1:03

Foundations · 1:03

What is Generative AI?

How generative AI works: from embeddings and the Transformer to foundation models. Scaling laws, LLMOps, and differences between LLMs, RAG, and agents.

1. Embeddings: Turning text into numbers · 2. The Transformer: The architecture that changes everything
1:03

Foundations · 1:03

What is AI?

What Artificial Intelligence is, how it works and how it evolved: from heuristics and Machine Learning to neural networks and foundation models.

1. The General Framework: AI, ML, DL and GenAI · 2. How do these systems learn?