Chapter 6 of 10All chapters
Chapter 6 of 10
Keeping it general
Fighting memorisation.
Techniques
Dropout randomly disables units during training. Weight decay penalises large weights. Early stopping halts when validation performance stops improving. Data augmentation invents plausible variations.
- Each adds a constraint that discourages memorising individual examples.
- More data remains the strongest regulariser of all.
Batch normalisation
Rescaling activations between layers stabilises training and allows higher learning rates. It became standard because it works, and the full explanation is still debated.