Learning parameters and representations
Maintenance grows as conditions change; automating a rule does not solve all its exceptions.
Links containing ?t= open the video at a specific second.
The ideas to retain
1. The age of rules: when intelligence was written by hand
The first symbolic systems
Early AI programs were built around a powerful intuition: if reasoning can be expressed as a sequence of formal steps, perhaps we only need to represent those steps and let the machine…
Expert systems: the mature symbolic paradigm
This approach reached its most solid form in expert systems. Rather than aiming for general intelligence, they tried to capture knowledge in a narrow domain through rules, facts and…
Jump directly to a section
Read the reviewed transcript
This video has no narration. This transcript reproduces its on-screen text; it does not invent a spoken track.
Exceptions expand the rule base
An initial rule handles a bounded set of cases.
A new case requires an exception and may conflict with another rule.
Maintenance grows as conditions change; automating a rule does not solve all its exceptions.
Fitting examples is not enough to generalize
A flexible curve can pass through all training points.
Held-out points test whether the fit preserves the pattern beyond the sample.
Low training error does not guarantee low error on new data.
The gradient gives a local direction
Use loss L equals w minus three squared, starting at w equals zero.
The gradient is minus six. A step size of zero point two moves w to one point two.
Loss falls from 9 to 3.24; the next step recomputes the gradient.
A representation changes what can be separated
In XOR, the two opposite corners with one active input have label one.
No straight line separates those corners from the other two in the original plane.
A nonlinear transformation can distinguish the case with exactly one active input.
Error propagates through derivatives
A network composes operations: x passes through two weights before producing an output.
The chain rule connects loss changes to each intermediate operation.
Backpropagation reuses those derivatives to update weights across layers.
A short window discards distant context
A trigram model conditions the next token on the previous two.
Two sentences with the same ending receive the same short context despite different earlier topics.
Limited memory simplifies the model and removes dependencies outside the window.


