MEPX
Chapter 6 of 10All chapters

Chapter 6 of 10

The alignment problem

Getting what you meant.

Specification

Systems optimise what they are measured on. Any measurable proxy for a goal will be pursued literally, including in ways that defeat the intent behind it.

  • Engagement optimisation producing outrage is the widely cited example.
  • This is Goodhart's law, familiar long before machine learning.

Current methods

Training on human preferences, explicit rules and red teaming all help and none is complete. Systems still behave unexpectedly outside the situations they were tested in.