MEPX
Chapter 9 of 10All chapters

Chapter 9 of 10

Knowing what is happening

Logs, metrics and traces.

Three views

Metrics count and time things over the whole system. Logs record individual events. Traces follow one request across services. You need all three for different questions.

  • Alert on symptoms users feel, such as error rate and latency, not on CPU.
  • Percentiles matter more than averages: the slowest one percent is somebody every minute.

Before the incident

Dashboards and alerts written during a calm week are worth more than any amount of investigation during a broken one.