The Cloudflare outage is remembered as "a bad regex." Their own postmortem lists eleven contributing factors, and the regex is one of them. Several have no regular expression anywhere near them:
- A CPU-exhaustion protection existed and "was removed by mistake during a refactoring" — the guard for this exact failure had been deleted, by accident, before the incident.
- The regex engine in use "didn't have complexity guarantees."
- The test suite had "no way of identifying excessive CPU consumption" — which is why the tests passed at 13:37, five minutes before the outage.
- The standard operating procedure "allowed a non-emergency rule change to go globally into production without a staged rollout."
- The rollback plan "required running the complete WAF build twice, taking too long."
- The status page updated slowly.
- Responders had access problems reaching internal systems during the incident.
That last one is the archetype of the factor a technically-framed review has no slot for, and it recurs everywhere: the tools you need to fix an outage are often behind the thing that is down.
The practical consequence of naming one cause instead of eleven: you get one action item. Fixing only the regex leaves the missing CPU guard, the blind test suite, the unstaged rollout and the slow rollback all in place — and the next defect of any kind walks the same path. Note also that at least three of the eleven (rollout policy, rollback duration, internal access) would have reduced the impact of any incident, not just this one. Those are usually the cheapest and the most reusable, and they are invisible if the review stops at the trigger.