MEPX
Chapter 4 of 10All chapters

Chapter 4 of 10

Activation functions

Small choices with large effects.

The common ones

ReLU passes positive values and zeroes negatives, and is the default for hidden layers. Sigmoid and softmax turn outputs into probabilities at the end.

  • ReLU largely solved the vanishing gradient problem that stalled deep networks.
  • Softmax is for choosing one class among several.

Why it mattered

Deep networks were known long before they worked. Better activations, initialisation and hardware are what made training them practical.