Distributed Reliability · Principal
The second region is healthy. Why does failover still take the AI service down?
The question
Interview question
Region A handles most inference traffic. Region B handles a small live slice and passes health checks. When A fails, routing sends its load to B. B's model workers can scale, but a shared quota service and B's vector-search cluster saturate. The AI endpoint becomes unavailable in both regions. Was the failover design actually regional?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Only parts of it. An endpoint health check proves that a particular request path worked at current traffic. It does not prove capacity for the combined regional load, or independence from control planes, identity, keys, retrieval, telemetry and rate-limit stores. AWS's multi-region dependency guidance emphasizes avoiding shared fate across regional components, and its disaster-recovery testing guidance calls out inadequate secondary capacity and quotas. Those are concrete design risks, not a guarantee that a particular cloud failover will fail.
I would draw the user journey, not just the gateway and GPU fleet: admission, authentication, model routing, prompt or adapter loading, retrieval, tools, result storage, streaming and billing. For each dependency, record where it lives, what traffic it sees during failover, what data must already be replicated, and what happens if it is unavailable. Cross-region reads back into failed A defeat the point. B may have enough GPUs but insufficient model weights, KV warm-up capacity or search partitions. Shared rate limits may be safe only if their data plane is available independently of A.
Design to a specified recovery objective and degraded-service policy. Reserve or prearrange burst capacity where possible. Route essential requests first, limit expensive long-context and batch work, and reject early when admission cannot honor deadlines. Ensure per-tenant limits do not disappear under failover. Test traffic shift gradually with failure injection that removes A and then separately removes shared dependencies. Measure successful end-to-end tasks, p99, queue growth and data freshness in B. A partial failover can be worse than a decisive one if retries bounce between regions and double load.
Would I keep B at full idle capacity all year? Maybe not. A warm standby can be cheaper with a longer recovery time, but the product must accept that contract. The model provider recovered. Why did 50,000 agents take it down again? is a retry storm after a model provider recovers. This is a regional failure in which supposedly independent services and quotas share fate. A green health check is not evidence of a full-load recovery plan.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →