Scale — deep learning to foundation models
A short video explanation of Scale — deep learning to foundation models.
Links containing ?t= open the video at a specific second.
Video summary
The ideas to retain
1. 2012: when scale stopped being a detail
AlexNet won ILSVRC 2012 with a result that changed the field's perception: 15.3% top-5 error versus 26.2% for the runner-up. The system was trained on 1.2 million images using two GTX 580…
2. The Transformer and massive pretraining
The next major shift arrived with Attention Is All You Need in 2017. The Transformer was not merely another language architecture. It reorganized the problem around attention mechanisms,…
3. Scale became a methodology
The idea that performance improves relatively predictably as parameters, data and compute increase did not originate with LLMs, but LLMs made it central. Work such as Deep Learning Scaling…


