Chapter 7 of 10All chapters
Chapter 7 of 10
Training in practice
Why it is fiddly.
What goes wrong
Loss not decreasing, exploding to infinity, or falling on training while rising on validation. Each has a short list of usual causes, starting with the learning rate.
- Always overfit a tiny subset first: if it cannot, something is broken.
- Initialisation and normalisation of inputs matter more than beginners expect.
Compute
Training large models needs specialised hardware because the operations are enormous matrix multiplications. That is exactly what graphics processors were built for.