Scaling data, computation and reuse
Learning those filters builds representations from pixels.
Links containing ?t= open the video at a specific second.
The ideas to retain
1. 2012: when scale became central
AlexNet won ILSVRC 2012 with a result that changed the field's perception: 15.3% top-5 error versus 26.2% for the runner-up. The system was trained on 1.2 million images using two GTX 580…
2. The Transformer and massive pretraining
The next major shift arrived with Attention Is All You Need in 2017. The Transformer was not merely another language architecture. It reorganized the problem around attention mechanisms,…
3. Scale became a methodology
The idea that performance improves relatively predictably as parameters, data and compute increase did not originate with LLMs, but LLMs made it central. Work such as Deep Learning Scaling…
Jump directly to a section
Read the reviewed transcript
This video has no narration. This transcript reproduces its on-screen text; it does not invent a spoken track.
A filter shares parameters across positions
The same local filter visits different image regions.
Its response changes with each region’s content while the weights remain shared.
Learning those filters builds representations from pixels.
Partitioning does not eliminate communication
Split a matrix product into four output blocks.
Each compute unit calculates a block using the operands it needs.
The output assembles those blocks; moving operands and results also consumes resources.
Doubling positions quadruples pairs
Unmasked dense attention considers sixteen pairs for four positions.
At eight positions, the grid contains sixty-four pairs.
This quadratic relation counts interactions; it does not itself predict kernel runtime.
A fixed budget imposes a trade-off
In this normalized example, cost is proportional to parameters times data.
Doubling parameters requires fewer data at the same fixed budget.
The best allocation needs a validated loss law, not an assumption that more parameters always wins.
One base can adapt to different tasks
Pretraining builds shared parameters from broad data.
The same base can receive instructions or adaptations for different tasks.
Reusing the base does not imply equal quality across all tasks.
Instruction following adds a training objective
Predicting continuations and responding appropriately to instructions are not identical objectives.
Response examples and preference signals can guide later training.
That training changes behavior; it does not guarantee every response is true or safe.


