Distributed Reliability · Principal
Both databases prepared the transaction. Why are writes now blocked?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Prepare is a promise from each participant that it can finish later. It is not the final commit. The coordinator first asks both databases to prepare, then decides commit or abort and tells them. If it dies after they prepare but before they learn the decision, the participants are in doubt. A prepared transaction can keep locks and other resources while waiting. PostgreSQL's PREPARE TRANSACTION documentation explicitly warns about prepared transactions left open and the locks they hold. That is why ordinary writers touching those rows can queue even though the coordinator process is gone.
The tempting fix is ROLLBACK PREPARED on both. I would stop there until we know the durable decision. The coordinator may have written a commit decision and sent it to one database before crashing. Unilaterally aborting the other would split the outcome. Recovery needs a stable transaction ID on both participants, a durable coordinator decision log, and a process that resends the recorded decision until all participants acknowledge. If the log records no decision, the coordinator's protocol must define what recovery can safely decide. Preserve the distinction between no decision was durably made and this operator has not found the decision yet. PostgreSQL's two-phase transaction overview makes clear that an external transaction manager owns the distributed protocol.
In the incident I would find prepared transaction IDs and ages, map them to application operations, inspect the coordinator log and each participant's outcome, and recover only with evidence for the transaction's decision. Alert on old prepared transactions long before they exhaust lock or transaction resources. Put a bound on prepare-to-decision time in normal operation and test coordinator crash at every transition, especially after the first commit message.
This is the availability price of atomic commit under a coordinator failure. If one side is an external API with no prepare and commit protocol, two-phase commit cannot make that API a participant. Then the design needs an application workflow with idempotency, durable intent and reconciliation, with explicit handling of a partial outcome.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →