Chapter 7 of 10All chapters
Chapter 7 of 10
Myths about data and training
Bigger is not automatically better.
Quality over quantity
Duplicated, biased or mislabelled data at scale produces those problems at scale. Curation now matters as much as volume for frontier models.
- Models can memorise and reproduce rare training examples.
- Training on machine-generated text uncritically degrades quality over generations.
Open and closed
Open weights let anyone run and adapt a model; open training data is much rarer. Open source in this field usually means something narrower than the term implies.