Failover is a decision, not an event
Diagrams draw failover as an arrow: primary fails, traffic moves. In a running system the arrow is a decision made on incomplete information, under time pressure, and it can be hard to undo.
It rests on uncertain inputs. Is the region really gone, or only unreachable? How much would promotion lose, given that a replica is only as current as its lag? What does waiting cost, against a wrong call? AWS's multi-Region guidance says the time to answer when to fail over is long before it is needed, with criteria that the whole organisation understands.
Who decides is itself a design choice. AWS recommends that the failover steps be fully automated while a small group of predetermined people decide whether to start them, weighing likely data loss and what is known about the event. Automation acts only on what it can measure, and that can go badly; a person weighs more but spends minutes doing it.
GitHub in 2018 shows the decision from inside. With both sites holding writes the other lacked, failing back was unsafe, and the team chose data integrity over a quicker recovery. That kind of choice belongs in the runbook before the day.