Capacity must already be there

A region that takes over from a failed one must carry its own traffic and the other's from the first minute. If each normally runs at 60% of what it can handle, the survivor is asked for 120% the moment traffic moves.

Scaling up on the day looks cheaper, but adding instances is a control plane action, and AWS's disaster recovery guidance says a recovery that depends on Auto Scaling is less resilient for it. Quotas in the recovery region must allow the scale-up, and AWS suggests capacity reservations where the machines have to be available when needed. Its own services go the other way: spread over three zones, they overprovision by half so that any two can carry the whole load without launching anything.

That headroom is the statically stable choice and the expensive one. Sarah Wells, describing the FT's two-region setup, says the one remaining region must handle all traffic, and that growth and architecture changes quietly erode the margin, so only load-testing the failover shows it is still there.

The load arriving at the survivor also includes every client retrying what just failed, which is how retries turn an outage into an overload.