Every note here needs an owner and the specific check that would have caught the thing. "Improve monitoring" is a wish; "the test suite must fail a rule that raises CPU beyond a threshold" is an action, because someone can verify next week whether it exists. Cloudflare's factor list is largely written this way — each names a missing mechanism rather than an intention.
The most instructive one is the rollback: "the rollback plan required running the complete WAF build twice, taking too long." Nobody would have written a design document saying recovery takes two full builds. It became true incrementally, and it was only measurable during an outage. The check is a rehearsal, not a code change — time the rollback while nothing is wrong, or you will discover its duration at the worst moment.
And an accepted cost, written down, is also an action. GitHub's decision is the clearest published example: with the East Coast primary lost, they refused to fail back, and said why —
"we decided that the 30+ minutes of data written to the US West Coast data center prevented us from considering options other than failing-forward in order to keep user data safe."
They chose 24 hours and 11 minutes of degradation over the risk of losing 30 minutes of writes, and published the reasoning. That is not a fix and it is not an omission — it is a decision, with the alternative named and the cost stated. A year later nobody has to ask why recovery took a day. An accepted risk that isn't written down becomes the next incident's contributing factor, and usually gets rediscovered as a surprise.