Chapter 11 of 12All chapters
Chapter 11 of 12
Beyond text
Images, audio and agents.
Other kinds of model
The same training ideas generate images, transcribe speech, synthesise voices and drive robots. Multimodal models handle several of these at once, taking an image and text together.
- Convincing fake images, video and voices are now cheap, which changes what evidence means.
- Provenance markers and content credentials are an attempt to answer that.
Agents
An agent is a model given tools and allowed to act in a loop: search, run code, call an interface. Capability grows and so does the cost of a mistake, which is why permission and review matter.