Retries look like reliability because the idea is simple. A request fails, so you try again. In most systems, that is reasonable. For normal services, it is one of the first things we add. Timeout, retry, backoff, jitter. Basic production hygiene.

But LLM systems behave differently.

A retry is not just another HTTP request. It may carry the same prompt, the same retrieved context, the same chat history, the same tool output and the same expensive model call. It adds load, latency and token cost. It can also push more traffic into a provider that is already unhealthy.

Fallback has the same problem. It sounds simple in a design review. If one model fails, send the request to another model. But the moment you do that, you may keep the system available while changing the product underneath the user. The response can become slower, more expensive, less accurate or just different in tone.

Retries look like reliability. In LLM systems, they can turn one failure into more load, more cost and different product behavior.

The old retry pattern assumes failures are short, recovery is cheap and fallback is equivalent. LLM systems break all three assumptions.

A retry storm is still a storm#

Most retry logic is designed for small transient failures: a network blip, a dropped connection, a random 500 or a temporary timeout. The assumption is that if you wait a little and try again, the dependency may recover and the user never sees the failure.

That assumption works when the failure is small and isolated. But LLM providers do not always fail like that. Sometimes the model endpoint is degraded for minutes. Sometimes the provider is capacity constrained. Sometimes rate limits are already kicking in. Sometimes latency is high but the API is not fully down, so from the outside it looks alive, but for your workload it is not healthy.

This is where normal retry logic starts doing the wrong thing. The first request times out, the retry goes out, that retry also times out, more requests enter the system, those requests also retry, the queue grows and the provider receives more traffic at the exact moment it needed less traffic.

At that point, retry logic is not protecting the system. It is adding pressure.

Retry storms existed long before LLMs. What changes with LLM systems is the shape of the dependency. Calls are slow, payloads are large, outputs can stream for a long time and many user flows are blocked on one expensive external model call.

If your timeout is too short, you may retry work that is still running. If your max retries are too high, you keep sending traffic into a dependency that is already struggling. If every service uses the same retry pattern, the whole system can learn to fail in sync.

The provider may have caused the first failure, but your retry policy can create the second one.

That is the part teams miss. Retries are supposed to reduce user visible failures, but during an LLM incident they can become part of the incident.

Every retry has a token bill#

In traditional systems, a retry costs CPU, network and latency. That still matters, but in LLM systems a retry can also mean paying again for the same request. The same user question, same system prompt, same retrieved documents, same conversation history and same tool context may get sent again.

That changes the economics of reliability. A retry is no longer just a recovery attempt. It is also a billable event.

Imagine a production LLM application doing 50 requests per second at peak. Each request has around 3,000 input tokens and 500 output tokens. Now imagine the provider has a degraded window of 15 minutes where 60% of requests fail or time out.

Without retries, you are sending 50 requests per second into the provider. With naive retries, if each failed request is retried 3 times on average, then 30 failed requests per second create 90 retry attempts per second.

So your 50 requests per second system is now behaving like a 140 request per second system against a dependency that is already unhealthy. That is almost 3 times the request pressure during the exact window where the provider is least able to handle it.

The token math gets ugly. Over 15 minutes, the original traffic is 45,000 requests. The retry traffic alone is 81,000 extra attempts. At 3,000 input tokens per attempt, that is 243 million extra input tokens sent during the incident window. This is before counting output tokens from partial successes, tool calls, long context windows or retry chains across multiple internal services.

The outage did not only create errors. It created traffic, cost and more pressure on the provider.

This is where normal retry thinking breaks down. In a normal HTTP service, a retry is often cheap enough that teams do not think about it as a financial event. In an LLM system, retries can show up directly on the bill. Even a short incident of 15 minutes can become a cost spike. If the system is high traffic and every request carries long context, your retry policy can spend money while making the user experience worse.

That is the strange part. You may be paying extra to fail harder.

This is why retry budgets matter, not just retry counts. The question is not only how many times should we retry. The better question is how much extra traffic are we allowed to create during failure, how many extra tokens are acceptable, when do we stop trying, when do we return a degraded response and when do we tell the user to try again later.

If your retry policy does not answer those questions, it is not a reliability policy. It is hope with exponential backoff.

Fallback is not the same product#

The common answer is simple: if one model fails, fall back to another model. It sounds reasonable. Sometimes it is reasonable.

But there is a hidden assumption inside that sentence. It assumes the fallback model is equivalent.

It is not.

LLM providers are not interchangeable replicas of the same database. They do not have the same behavior, cost, latency, refusal style or reasoning profile. So fallback may preserve availability, but it may not preserve the product.

This is something teams often learn late because the infrastructure dashboard can look healthy while the product experience has changed. Requests are no longer failing, but users are now seeing different answers. Maybe the fallback model is more verbose. Maybe it refuses more often. Maybe it is cheaper but weaker on reasoning. Maybe it reasons better but runs much slower. Maybe it follows the output format less reliably. Maybe the tone changes enough that customers notice something is different.

From the infrastructure point of view, fallback worked. From the product point of view, behavior changed.

The infrastructure dashboard stays green. The product does not.

When the primary model goes degraded, the fallback chain may handle the traffic exactly the way it was designed to. Requests no longer fail. From the SRE view, the incident looks contained.

But the fallback model is not the primary model. It costs differently, responds with different latency and produces output the primary would not have produced for the same prompt.

A workflow tuned for the primary model’s tool calling reliability may break on the fallback. A response format that worked on the primary may degrade on the fallback. A reasoning task the primary handled in one pass may need retries on the fallback.

That is the trap. The metric you watched, request success rate, said the system recovered. The metric you were not watching, response equivalence, had moved.

That is not true redundancy. That is degraded mode.

There is nothing wrong with degraded mode if you name it honestly. The problem is pretending it is invisible. If your primary model is used because it has the right quality, latency, tone and tool behavior for the product, then falling back to a different model changes those properties. Sometimes that is acceptable. Sometimes it is not.

A customer support assistant may tolerate a slightly slower fallback. A legal drafting workflow may not tolerate a model with weaker citation behavior. An internal coding tool may tolerate a tone change. A user facing assistant may not. A workflow that depends on strict JSON output may break if the fallback model is worse at formatting.

The fallback decision is not only an infrastructure decision. It is a product decision.

That means multi model fallback needs evals, cost tracking, latency tracking, behavior comparison and clear rules for what the fallback is allowed to do. The question should not only be whether fallback kept the system up. The real question is whether fallback preserved the behavior users depend on.

Availability can stay green while the product quietly changes.

Circuit breakers matter more than blind retries#

The answer is not to never retry. That would be wrong. Retries are useful. They handle transient failures, smooth out small network issues and reduce random user visible errors.

The problem is blind retries.

Production reliability is not about retrying harder. It is about knowing when to stop. That is where circuit breakers matter.

A circuit breaker is not failure. It is the system refusing to make the failure worse.

If the provider is returning too many errors, stop sending full traffic. If latency is rising beyond a threshold, stop stacking more requests behind it. If retry traffic is crossing the budget, stop retrying. If token spend is spiking during failures, trip the breaker. If the fallback model is overloaded or behaving differently, do not blindly push all traffic there.

The shift is simple: you do not only monitor success rate. You monitor the cost of trying to succeed.

For LLM systems, circuit breakers should not trip only on errors. They should consider latency, timeout rate, rate limit responses, retry volume, token spend and fallback quality. The system needs a few modes: normal mode, degraded mode and fail fast mode.

In normal mode, retries can be reasonable. In degraded mode, reduce context, lower volume, delay noncritical work, use cheaper paths or ask the user to retry later. In fail fast mode, stop pretending the dependency is healthy.

That last part matters. Sometimes the most reliable thing a system can do is stop making the bad call for a while. Not forever, just long enough to protect the rest of the system.

A retry policy tries to save the request. A circuit breaker tries to save the system.

LLM systems need both, but the second one matters more than teams think.

Retry less blindly#

Retries are not bad. Blind retries are.

In LLM systems, retries are a reliability decision, a cost decision and a product decision. That makes them more dangerous than they look. A normal retry policy assumes the dependency will recover quickly, the extra work is cheap and the fallback is close enough. Those assumptions are often false.

The provider may be degraded for minutes. The extra work may resend millions of tokens. The fallback model may keep the system available but change the product.

That is why retry logic in LLM systems needs boundaries: retry budgets, token budgets, circuit breakers, fallback evals, degraded modes and clear rules for when to stop.

The goal is not to make every request succeed at any cost. The goal is to stop one failure from becoming more load, more cost and a different product behavior.

Retries look like reliability, but in LLM systems reliability comes from controlling them.