Build the timeline from artifacts and record two clocks: when the system broke, and when a human understood it. Cloudflare's 2 July 2019 postmortem publishes both, to the minute:
| UTC | What happened |
|---|---|
| 13:31 | Engineer merges the pull request |
| 13:37 | CI builds the rules and the tests pass |
| 13:42 | Automatic deployment begins — global |
| 13:45 | First page fires |
| 14:00 | WAF identified as the culprit |
| 14:02 | Global WAF termination proposed |
| 14:07 | Termination executed |
| 14:09 | Traffic and CPU normal |
| 14:52 | WAF re-enabled after validation |
The number people quote is 27 minutes of outage. The number that tells you where to invest is different: 3 minutes to alert, 18 minutes to diagnosis. Detection was excellent; the expensive interval was 13:45–14:00, with the whole team awake and looking at the wrong thing. Optimising alerting would have saved nothing here.
The second clock also breaks the intuition that a big outage has a big cause. GitHub's 21 October 2018 analysis records a network partition lasting 43 seconds — and 24 hours and 11 minutes of degradation. Their detection was fast too (alerts at 22:54, two minutes in; database topology understood by 23:02). The duration of the cause tells you nothing about the duration of the impact, so a timeline that records only the trigger has recorded the least useful part.