Chapter 9 of 10All chapters
Chapter 9 of 10
Knowing what is happening
Logs, metrics and traces.
Three views
Metrics count and time things over the whole system. Logs record individual events. Traces follow one request across services. You need all three for different questions.
- Alert on symptoms users feel, such as error rate and latency, not on CPU.
- Percentiles matter more than averages: the slowest one percent is somebody every minute.
Before the incident
Dashboards and alerts written during a calm week are worth more than any amount of investigation during a broken one.