MEPX
Chapter 6 of 10All chapters

Chapter 6 of 10

Common algorithms

A short tour.

The workhorses

Linear and logistic regression are interpretable baselines. Decision trees split on feature values. Random forests and gradient boosting combine many trees and win most tabular problems.

  • Gradient boosted trees remain the default for structured data.
  • Neural networks dominate images, audio and language, not spreadsheets.

Unsupervised

Clustering groups similar records without labels, and dimensionality reduction compresses many features into few. Both are exploratory rather than predictive.