Chapter 4 of 10All chapters
Chapter 4 of 10
Activation functions
Small choices with large effects.
The common ones
ReLU passes positive values and zeroes negatives, and is the default for hidden layers. Sigmoid and softmax turn outputs into probabilities at the end.
- ReLU largely solved the vanishing gradient problem that stalled deep networks.
- Softmax is for choosing one class among several.
Why it mattered
Deep networks were known long before they worked. Better activations, initialisation and hardware are what made training them practical.