The first backfill attempt read the whole table in a single pass and wrote the new columns in batches. It was abandoned.
Why it looked right: it is the obvious approach, it is easy to reason about, and on a copy of the data it completed in a few minutes.
What killed it: the copy was a snapshot, so it had no concurrent writes. Against a live table the single long-running read held a transaction open for the duration, which prevented vacuum from reclaiming rows, which grew the table, which made the read slower. The failure is self-reinforcing, and it does not appear at all in the environment where it was tested.
The signal that misled us, which is still there: a staging environment with production-sized data but no production write volume. It will make the next long-running job look safe too. Sizing staging by rows is what made this look tested; the thing that mattered was concurrency, and nothing about the environment surfaced that.
Replaced by a chunked backfill with a bounded transaction per chunk.