Chapter 5 of 10All chapters
Chapter 5 of 10
Multiple comparisons
Testing until something works.
The problem
Every test carries a false positive chance. Run twenty at the 5 percent level and expect one significant result from pure noise.
- Corrections such as Bonferroni or false discovery rate control this.
- Dashboards with dozens of metrics are a multiple comparisons machine.
Researcher degrees of freedom
Choosing the outcome, the subgroup or the exclusion rule after seeing the data inflates false positives enormously. Pre-registering the plan is the fix.