I read a lot of postmortems and skim most of them. Inngest’s write-up on their September 18 outage is one I read twice, because it has exact timestamps for every stage of the failure, and the failure is one I’ve come close to causing myself: deleting an account with ON DELETE CASCADE set across dozens of tables, inside a single transaction, and trusting that “it’s just a delete” means it’s cheap.
It wasn’t cheap. It took down the service twice in one afternoon.
15:41 UTC — the transaction that held too much
A Vercel Marketplace account deletion kicked off a cascade across dozens of related tables, all inside one Postgres transaction. Cascading deletes at that scope aren’t a single fast operation — they’re a long-running transaction that holds row and table locks for as long as it takes to walk every foreign key relationship and remove every dependent row. The bigger the account’s footprint, the longer that transaction runs, and the longer everything else waits behind its locks.
Between 15:41 and 15:56 UTC, retried deletion attempts queued up behind the original — not instead of it, behind it. Nothing deduplicated the retries. Nothing serialized them against the transaction already in flight. Each retry just piled onto the lock queue, waiting its turn to also try to delete rows that were, at that moment, still mid-deletion.
The part that actually took the site down
Here’s the mechanism I hadn’t fully internalized until this postmortem: blocked queries don’t just sit there quietly. They hold their PgBouncer client connection slot for the entire duration of the lock wait. A query blocked for ninety seconds occupies a connection slot for ninety seconds, whether or not it’s doing any actual work.
Client connections climbed to roughly 14 times steady-state as the queued, retried, lock-blocked queries accumulated. By 16:18–16:19 UTC, both PgBouncer instances hit max_client_conn and started returning FATAL: no more connections allowed. At that point it doesn’t matter how healthy your database itself is — nothing new can talk to it. The outage wasn’t really a database failure. It was a connection-pool failure caused by a database problem, which is a distinction that changes what you’d actually fix.
The executor, the queue-proxy, and the CDC pipeline all crash-looped once connections were unavailable, which is what turned a slow query problem into a full outage. Two separate windows — 24 minutes, then another 22 minutes during the fix rollout itself — before it was actually stable.
Three separate things had to go wrong together
What I like about this postmortem is that no single decision was crazy on its own. Cascading FK deletes are a completely normal schema pattern. Retrying a failed or slow operation is completely normal client behavior. A connection pool with a max size is completely normal infrastructure. None of the three is a mistake in isolation.
Stacked together, they compound into a failure mode where the retry logic amplifies the exact problem it’s trying to route around. The retries weren’t caused by a bug in the retry logic — they were doing exactly what retry logic is supposed to do. The bug was that nothing above the retry layer knew the original operation was still in flight and holding locks, so “try again” meant “add another blocked query to a pool that’s already draining.”
The fix is duller than the failure, which is the right shape
Inngest’s fix wasn’t a clever one — a kill switch on Marketplace deletions to stop the bleeding immediately, converting hard deletes to soft deletes so the expensive cascading cleanup happens asynchronously instead of inside a user-facing transaction, batching and rate-limiting and deduplicating whatever hard deletes remain, and isolating PgBouncer pools so one workload’s connection exhaustion can’t starve unrelated traffic.
None of that is novel engineering. It’s the standard playbook for “an expensive operation is coupled to a synchronous user-facing path.” What made it necessary wasn’t a lack of sophistication — it’s that “account deletion” reads as a low-frequency, low-risk operation until you actually measure how many tables it touches and how long the transaction runs under load.
Where I checked my own systems after reading this
Two questions worth running against anything you own that does cascading deletes: does the delete run inside one transaction that holds locks for its full duration, and does your retry logic know whether the original attempt is still in flight before firing another one? If the answer to the second is “no, it just retries on timeout,” you have the exact shape of this bug sitting dormant, waiting for an account large enough to make the transaction slow enough to matter.
Pool isolation is the boring insurance policy underneath all of it — one workload’s failure mode shouldn’t be able to starve every other workload’s connections, and that’s true independent of whether you ever hit this specific cascading-delete scenario at all.