All videosHistory of AI2:11

Scaling data, computation and reuse

Learning those filters builds representations from pixels.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

1. 2012: when scale became central

AlexNet won ILSVRC 2012 with a result that changed the field's perception: 15.3% top-5 error versus 26.2% for the runner-up. The system was trained on 1.2 million images using two GTX 580…

02

2. The Transformer and massive pretraining

The next major shift arrived with Attention Is All You Need in 2017. The Transformer was not merely another language architecture. It reorganized the problem around attention mechanisms,…

03

3. Scale became a methodology

The idea that performance improves relatively predictably as parameters, data and compute increase did not originate with LLMs, but LLMs made it central. Work such as Deep Learning Scaling…

Key moments

Jump directly to a section

  1. A filter shares parameters across positions
  2. Partitioning does not eliminate communication
  3. Doubling positions quadruples pairs
  4. A fixed budget imposes a trade-off
  5. One base can adapt to different tasks
  6. Instruction following adds a training objective
Read the reviewed transcript

This video has no narration. This transcript reproduces its on-screen text; it does not invent a spoken track.

A filter shares parameters across positions

The same local filter visits different image regions.

Its response changes with each region’s content while the weights remain shared.

Learning those filters builds representations from pixels.

Partitioning does not eliminate communication

Split a matrix product into four output blocks.

Each compute unit calculates a block using the operands it needs.

The output assembles those blocks; moving operands and results also consumes resources.

Doubling positions quadruples pairs

Unmasked dense attention considers sixteen pairs for four positions.

At eight positions, the grid contains sixty-four pairs.

This quadratic relation counts interactions; it does not itself predict kernel runtime.

A fixed budget imposes a trade-off

In this normalized example, cost is proportional to parameters times data.

Doubling parameters requires fewer data at the same fixed budget.

The best allocation needs a validated loss law, not an assumption that more parameters always wins.

One base can adapt to different tasks

Pretraining builds shared parameters from broad data.

The same base can receive instructions or adaptations for different tasks.

Reusing the base does not imply equal quality across all tasks.

Instruction following adds a training objective

Predicting continuations and responding appropriately to instructions are not identical objectives.

Response examples and preference signals can guide later training.

That training changes behavior; it does not guarantee every response is true or safe.