MEPX
Chapter 7 of 10All chapters

Chapter 7 of 10

Myths about data and training

Bigger is not automatically better.

Quality over quantity

Duplicated, biased or mislabelled data at scale produces those problems at scale. Curation now matters as much as volume for frontier models.

  • Models can memorise and reproduce rare training examples.
  • Training on machine-generated text uncritically degrades quality over generations.

Open and closed

Open weights let anyone run and adapt a model; open training data is much rarer. Open source in this field usually means something narrower than the term implies.