Chapter 6 of 10All chapters
Chapter 6 of 10
The alignment problem
Getting what you meant.
Specification
Systems optimise what they are measured on. Any measurable proxy for a goal will be pursued literally, including in ways that defeat the intent behind it.
- Engagement optimisation producing outrage is the widely cited example.
- This is Goodhart's law, familiar long before machine learning.
Current methods
Training on human preferences, explicit rules and red teaming all help and none is complete. Systems still behave unexpectedly outside the situations they were tested in.