It moved execution to a region that could not yet prove the current authorization state. Asynchronous replication has a recovery point that may precede recent writes. AWS's Aurora Global Database failover guidance explicitly notes the possibility that not all transactions reached the secondary before an unplanned promotion. That is a property to plan around, not a statement that this hypothetical service uses Aurora. If revocation is a safety boundary, “the workflow resumed” is not enough evidence to execute an approved action.

I would reconstruct the source commit sequence, replication watermark in B, the last approval and revocation events, workflow history and tool-call time. Distinguish a revocation that committed before A failed from one that was merely acknowledged by an unavailable client path. Define what acknowledgement meant. A regional store that has lost an acknowledged revocation cannot safely serve a fresh permission decision under the old contract. The action gateway should require a sufficiently current authorization epoch or consult an independent authority. If it cannot establish current state, pause the action rather than guess from the last replicated approval.

The recovery protocol needs a fence. Prevent A from executing after B becomes authoritative, and prevent B from replaying old approved work until the relevant policy and approval state is reconciled. A global or quorum-backed authority for high-impact actions may be justified, with availability and latency costs. Another design is to issue short-lived, scoped permits whose revocation semantics are explicit. Neither design can promise instantaneous revocation during a partition without paying for coordination. State the tradeoff and the product's maximum exposure window, then test failover between approval, revocation and action dispatch.

What if we wait for B to catch up? After A is lost, B may not have the missing log at all. Then the system needs an external source, a conservative hold, or a documented recovery process. The second region is healthy. Why does failover still take the AI service down? asks whether the secondary has capacity and independent dependencies. The agent's task was approved yesterday. The employee loses access today. Does the scheduled action run? asks whether a scheduled action still has authority after the employee loses access. Here replication lag during failover makes that authority question acute, even though the workflow itself resumed successfully.