MEPX
Chapter 6 of 10All chapters

Chapter 6 of 10

Keeping it general

Fighting memorisation.

Techniques

Dropout randomly disables units during training. Weight decay penalises large weights. Early stopping halts when validation performance stops improving. Data augmentation invents plausible variations.

  • Each adds a constraint that discourages memorising individual examples.
  • More data remains the strongest regulariser of all.

Batch normalisation

Rescaling activations between layers stabilises training and allows higher learning rates. It became standard because it works, and the full explanation is still debated.