MEPX
Chapter 11 of 12All chapters

Chapter 11 of 12

Beyond text

Images, audio and agents.

Other kinds of model

The same training ideas generate images, transcribe speech, synthesise voices and drive robots. Multimodal models handle several of these at once, taking an image and text together.

  • Convincing fake images, video and voices are now cheap, which changes what evidence means.
  • Provenance markers and content credentials are an attempt to answer that.

Agents

An agent is a model given tools and allowed to act in a loop: search, run code, call an interface. Capability grows and so does the cost of a mistake, which is why permission and review matter.