Per request routing looks efficient until switching models breaks cache locality and makes the whole conversation cost more.

Most AI cost optimization starts with the obvious moves: use cheaper models by default, cache repeated prompts, route simple tasks to smaller models, put a gateway in front of vendors, track spend by team, and cap heavy users before finance gets surprised.

That is a decent version one.

The problem starts when the workload stops being a single request.

A one shot prompt is mostly independent. A multi turn chat, coding agent, internal copilot, research assistant, or tool heavy workflow is not. Each turn carries system instructions, tool schemas, conversation history, policy text, retrieved context, repo state, previous tool output, and some working set that keeps growing across the session.

Once that happens, the cheapest model for the next turn is not always the cheapest route for the whole conversation.

A router that optimizes every turn independently can destroy the cache locality that made the previous route cheap. Locally the decision looks correct because the current turn is small, easy, and fits a cheaper model. At session level it can be wrong because the previous model had a warm reusable prefix and the new model does not. You saved on price per token, then paid again for the same context.

This is not really a “how do I route between models” problem. That part is normal platform work. The real problem is routing when request cost depends on session state.

The naive router#

The first router most teams build is simple.

request comes in
classify task difficulty
pick cheapest model that can answer
send request
record cost

There is usually a model catalog behind it. Small model is cheap and fast, medium model is good enough for most work, large model is expensive but better, and maybe a reasoning model is reserved for harder or riskier tasks. The router scores the current request, checks policy, checks provider health, and picks the cheapest route that looks safe.

This works when requests are mostly independent.

That assumption breaks fast in LLM systems. A chat turn is not isolated. An agent tool call is not isolated. A coding assistant turn is not isolated. Each one may carry a stable prefix that was already processed on the previous turn and could have been reused if the router stayed on the same route.

So the router cannot only ask:

what is the cheapest model for this request?

It also has to ask:

what is the cheapest path for this session?

That is where a stateless gateway starts becoming a real control plane.

Prompt caching is where this gets weird#

Prompt caching is often explained too casually, which makes the routing problem look simpler than it is.

If a new request shares the same stable prefix as an earlier request, the serving stack may reuse the work already done for that prefix. That prefix can include system instructions, tool definitions, output schemas, policy text, product docs, repo context, or earlier conversation history. If the cache hits, the request avoids paying full price for repeated input, and time to first token can improve too.

This is not semantic caching.

It is not “these two prompts are kind of similar.”

It is closer to “this serialized prefix matches in the way the model and provider expect.”

That means small things matter. A timestamp near the top can break it. A request ID inside the stable prompt can break it. Tool ordering can break it. Prompt template changes can break it. Model changes can break it. Provider changes can break it. Moving dynamic content before static content can break it. Redaction can break it too if the safe replacement is not stable.

This is why prompt construction becomes infra. If every app team builds prompts differently, the gateway may look like a router, but it does not have a stable cost surface to optimize.

Prompt cache and KV cache are the same idea at different altitudes#

This section has to be precise because inference engineers will notice if prompt cache and KV cache are treated like two unrelated features.

They are not parallel systems. One is the lower level mechanism. The other is the API surface you can touch.

During generation, a transformer caches the key and value tensors it computed for previous tokens, so the next output token does not require recomputing attention over the entire sequence from scratch. That is the KV cache. It is internal to inference, normally always on, and application code does not opt into it. Without it, decoding would be too expensive to serve normally.

Prompt caching is cross request prefix reuse exposed at the API layer.

If a new request shares a stable prefix with an earlier request, the serving stack can reuse previously computed state for that prefix instead of recomputing it. OpenAI exposes this through prompt or prefix caching behavior and cached token accounting. Anthropic exposes it through prompt caching controls, cache reads, cache writes, and TTL options. The product surface is different, but the instinct underneath is the same: reuse old prefix work instead of recomputing it.

Clean mental model:

KV cache:
  intra request reuse during decoding
  internal to serving
  part of normal inference

Prompt cache:
  cross request prefix reuse
  exposed through provider API behavior
  controlled indirectly through prompt shape, breakpoints, model choice, and provider choice

When you call an external provider, you do not manage KV blocks directly. You see the abstraction: cached input tokens, lower cached token price, cache read and write behavior, lower TTFT when the prefix hits, and usage metadata after the call. You control the inputs to that abstraction: stable prefix, cache breakpoints where supported, model choice, provider choice, prompt ordering, and session stickiness.

When you self host, the abstraction disappears and you manage the serving problem directly. Now the question becomes which replica has the prefix blocks, how PagedAttention or a similar block manager packs them, what gets evicted under memory pressure, whether to offload cold blocks to CPU, and whether restoring KV state is cheaper than recomputing prefill.

Same instinct. Different altitude.

External provider mode is mostly a gateway and control plane problem. Self hosted mode becomes gateway plus inference scheduler plus memory management.

Worked cost example: when the cheaper model costs more#

Take a coding assistant session.

The prompt contains system instructions, security policy, tool schemas, repo summary, selected files, conversation history, and earlier tool output.

Roughly:

system instructions                 2k tokens
security and policy text            3k tokens
tool schemas and output rules        5k tokens
repo summary and selected files     20k tokens
conversation and tool history       10k tokens

That gives you a 40k token stable or mostly stable prefix before the new user message.

Now the user says:

Can you also update the test and fix the failing assertion?

That might add only 100 new input tokens. Suppose expected output is 1,200 tokens.

Route A: stay on the current model#

Assume the conversation is already on Model A. Model A is more expensive per token, but it has a warm reusable prefix for the 40k token context.

Rough pricing, only for intuition:

Model A uncached input: $3 / 1M tokens
Model A cached input:   $0.30 / 1M tokens
Model A output:         $12 / 1M tokens

Cost:

40,000 cached input tokens   = 40,000 × 0.30 / 1,000,000 = $0.012
100 uncached input tokens    = 100 × 3 / 1,000,000       = $0.0003
1,200 output tokens          = 1,200 × 12 / 1,000,000    = $0.0144

Total:

$0.0267

Route B: switch to the cheaper model#

Now assume the router looks only at this turn, decides it is easy, and switches to Model B.

Rough pricing:

Model B input:  $1 / 1M tokens
Model B output: $4 / 1M tokens

Looks cheaper, but the prefix is cold on Model B.

Cost:

40,100 uncached input tokens = 40,100 × 1 / 1,000,000 = $0.0401
1,200 output tokens          = 1,200 × 4 / 1,000,000  = $0.0048

Total:

$0.0449

The cheaper model is about 68 percent more expensive for this turn because the router destroyed prefix locality.

Two LLM routing paths comparing a warm cache route on the current model with a cold cache route on a cheaper model, showing that switching models can increase total cost.

The cheaper model wins on price per token, but loses when switching destroys the warm prefix cache and forces the system to pay for the same context again.

One caveat: the 40k prefix did not become warm for free. The first turn had to establish that cache. Some providers charge a write premium for that, and most providers have a minimum cacheable prefix size, so tiny prompts do not benefit from this at all. The first turn is more expensive. The later turns become cheaper. Over a real multi turn coding session, the write cost gets amortized across many reads, which is why the warm route wins.

Do not get stuck on the exact prices. Change the model prices, output length, and prefix size, and the shape still holds. Once the stable prefix is large enough, the router cannot optimize only token price. It has to optimize effective cost.

effective_cost =
  uncached_input_tokens * uncached_input_price
+ cached_input_tokens * cached_input_price
+ output_tokens * output_price
+ expected_fallback_cost
+ expected_retry_cost
+ cache_write_or_rebuild_cost
+ latency_risk
+ quality_risk
+ switch_penalty

Most naive systems optimize this:

token_cost = input_tokens * input_price + output_tokens * output_price

That is incomplete once session level caching exists.

The system is not just a router box#

If the design is just one router box and three model boxes, that is not really a control plane. It is just an API multiplexer.

The actual system needs a control loop.

End to end cache aware LLM gateway architecture showing app or coding harness, prompt builder, gateway router, provider model layer, session state store, response reconciler, and spend ledger.

Cache aware routing is not just a model picker. The gateway needs prompt metadata, session state, provider usage, and post call reconciliation before it can make the next routing decision well.

The hot path is app, prompt builder, gateway, model, response. The control loop is session state, budget reservation, reconciliation, and spend ledger. Leave the control loop out and you can still send traffic to models, but you are not really optimizing cost.

The diagram is the shape of the system. The next question is what each part actually owns.

What each layer owns#

The app or coding harness owns the workflow. It knows whether the user is chatting, planning a code change, debugging a test, reviewing a diff, or running a background agent. The gateway should not infer all of that from raw text if the harness already knows it.

A better request includes metadata like:

task_type: code_edit
phase: test_debug
risk: medium
latency_sensitivity: interactive
session_id: abc
max_output_budget: 2000
expected_context_stability: high

The prompt builder owns prompt shape. It keeps stable content before dynamic content, canonicalizes tool definitions, versions prompt templates, makes redaction stable where policy allows, produces prefix fingerprints, and tells the gateway what changed.

The gateway owns shared controls: auth, tenant policy, redaction, provider abstraction, routing, budget checks, fallback, logging, and cost controls. The router inside it should know current model affinity, estimated cache warmth, policy constraints, task risk, provider health, expected switching cost, and run level budget.

The session state store holds routing metadata, not raw prompt text.

session_id
current_model
current_provider
prefix_fingerprint
prompt_template_version
tool_schema_version
stable_prefix_token_count
last_seen_at
estimated_cache_expiry
last_cached_tokens
last_uncached_tokens
failure_count
route_confidence

The reconciler corrects the router’s guesses after the provider returns usage data. It records actual cached input tokens, uncached input tokens, output tokens, latency, retry behavior, fallback behavior, and billed cost. Without this, the router keeps believing its own estimates.

The spend ledger should not only show cost per request. It should show cost per conversation, cost per workflow, cached versus uncached token share, cache miss after model switch, fallback cost, retry cost, model switches per session, and prompt version changes that affected cost.

Cache state is not a boolean#

A clean design doc says:

if cache is warm, stay
if cache is expired, switch

Production is not that clean.

With an external provider, the gateway may not know exactly where the cache lives, whether the next request lands on the same backend, whether memory pressure evicted it, whether provider side routing changed, or whether the provider’s retention behavior matched your expectation for this request. The gateway usually learns the truth only after the call, from usage metadata.

So the router should model cache hit probability, not a true or false cache flag.

cache_hit_probability =
  f(last_seen_at,
    provider retention behavior,
    prefix fingerprint match,
    model/provider affinity,
    prompt template version,
    historical hit rate for this route,
    recent observed misses,
    traffic rate,
    route confidence)

A simple heuristic is enough to start.

if prefix changed:
  P(hit) = 0

else if model or provider changed:
  P(hit) = low

else if last_seen_at is very recent:
  P(hit) = high

else if last_seen_at is inside expected TTL:
  P(hit) = medium

else:
  P(hit) = low

Then route on expected cost.

expected_input_cost =
  P(hit)  * cached_route_cost
+ P(miss) * cold_route_cost
Chart showing cache hit probability decaying over time after a session becomes inactive, with routing becoming safer once the cache is likely cold.

Cache state is not a clean warm or cold flag. The router has to estimate hit probability and switch only when the expected benefit is larger than the cache loss.

If the current model has a 70 percent chance of cache hit and the cheaper model has 0 percent chance, the cheaper model must be much cheaper to justify switching. If the current cache is probably expired, the router has more freedom. If the provider returned misses recently for the same route, lower confidence and stop pretending the cache is warm.

The router should also record why it made the call.

selected_model: model_a
reason: preserve_probable_cache
cache_hit_probability: 0.73
alternative_expected_saving: low
switch_penalty: high

This is how you debug the router when it stays on an expensive model and the cache misses anyway.

Stickiness is not rigidity#

Keeping a session on the same model does not mean never switch. It means switching has a cost, so it needs a threshold.

A good router should prefer the current route while the cache is probably warm, but it should still switch when the task becomes harder, the context window no longer fits, policy requires another model, provider health drops, budget is exhausted, the workflow enters a new phase, or quality risk becomes too high.

A simple rule:

switch only if:
  expected_benefit > cache_loss + latency_risk + quality_risk

Without a threshold, the router thrashes. One turn goes to Model A, the next goes to Model B, then back to A, then back to B. Each decision looks rational locally, and the full session gets worse.

External provider versus self hosted inference#

The design changes depending on whether you call external APIs or run inference yourself.

With external providers, the gateway mostly controls prompt shape, model choice, provider choice, cache breakpoints where supported, budget, fallback, and observability. It does not control provider side cache placement, eviction, batching, or backend routing. The gateway learns what happened after the call, then adjusts future routing.

External provider routing is mostly a control plane problem.

preserve model/provider affinity
keep prompt prefix stable
estimate cache hit probability
watch actual cached token usage
avoid unnecessary switches
track fallback cost

With self hosted inference, you own more of the data plane. Now the router can know which replica has which prefix blocks, how much GPU memory is available, whether KV blocks were evicted, whether blocks were offloaded to CPU, and whether restoring KV state is cheaper than recomputing prefill.

Self hosted routing may care about:

which replica has the prefix
GPU KV cache occupancy
block level reuse
LRU eviction pressure
prefill queue length
decode queue length
batch shape
offload bandwidth
replica locality
tenant isolation

In external provider mode, “stay on same model/provider” may be the best approximation you have. In self hosted mode, that is not enough. You may need “stay on the replica or cache tier that has the KV blocks,” or “route to the engine where restoring KV cache is cheaper than recomputing prefill.”

Same mental model: do not route only by model price. Route by effective cost of the path.

Failure modes that actually show up#

Router thrashing#

If the routing score sits near a boundary, the session bounces between models. That is bad for behavior consistency and bad for cache locality.

Use hysteresis. Do not upgrade and downgrade at the same threshold.

if currently on medium:
  downgrade only when score is clearly low

if currently on small:
  upgrade only when score is clearly high

Fallback destroys locality#

Cross provider failover is good for availability, but it can be expensive for large contexts because the fallback route may be cold. A tiny request and a 100k token coding session should not use the same failover strategy.

The ledger should track fallback extra cost, not just fallback count.

fallback_provider
fallback_extra_uncached_tokens
fallback_extra_cost
fallback_latency
fallback_quality_change

Prompt rollout creates a cache cliff#

A prompt change can behave like cache invalidation. Teams add a system instruction, reorder tools, change redaction, or insert dynamic metadata near the top, and suddenly cost and TTFT spike.

Version everything that can affect the prefix.

system_prompt_version
developer_prompt_version
tool_schema_version
policy_version
redaction_version
prompt_builder_version
model_version

Otherwise the bill jumps and everyone blames the model.

Redaction breaks stable prefixes#

If sensitive data is redacted differently every time, the safe prompt shape changes and the cache misses. Privacy transforms are part of serialization. They need stable replacement strategy where policy allows it.

This does not mean preserving sensitive data. It means the safe representation should not randomly change.

Cache hit rate becomes a vanity metric#

High cache hit rate can still hide bad economics if the hits are on tiny prompts and expensive workflows are cold. What matters is cached token share weighted by cost and tied to useful work.

Better questions:

how much cost did cached tokens actually save?
which model switches caused cold prefixes?
which prompt version broke cache?
which workflows are cheap per request but expensive per completed task?

Stickiness keeps you on the wrong model#

Cache locality matters, but not more than quality in high risk work. A cheap model with a warm prefix is still the wrong route if the task moved into architecture reasoning, code migration, financial analysis, legal review, or an irreversible tool action.

The router needs upgrade rules, not just cache rules.

The main tradeoffs#

Stickiness versus quality is the first tradeoff. Staying on the same route preserves cache, but sometimes the task becomes harder and quality matters more. Coding agents hit this all the time. Repo lookup can run on one model. Debugging a nasty test failure may need another. A risky code change may need a stronger review path.

Cache locality versus provider diversity is the second. A multi provider gateway gives availability, pricing leverage, regional routing, and policy flexibility. Cache locality pushes toward staying on the same provider and model. There is no one global rule. Interactive sessions should usually prefer locality. Batch jobs can optimize cost more aggressively. Regulated workflows should let policy beat cache. Coding agents should preserve affinity inside a phase and reconsider at phase boundaries.

Long stable prefix versus context hygiene is the third. Caching rewards stable prefixes, which creates a bad incentive to keep stuffing more context into them. More context is not always better. Latency can grow, attention gets diluted, instructions conflict, and the model can get worse. Sometimes the right move is not preserving the prefix. It is compacting, summarizing, dropping stale context, or rebuilding the working set.

Request level cost versus run level budget is the fourth. A single request may look cheap while the full agent run burns money across 30 steps. Budget control should exist at the run or workflow level, not only the request level.

Coding agents make this worse#

Coding agents are where naive routing breaks hardest.

They have long system prompts, tool definitions, repo maps, file contents, test results, stack traces, diffs, and subagent workflows. A normal chat may carry a few thousand tokens. A coding agent can carry tens of thousands across many turns.

Now imagine per turn routing inside that workflow.

Planning goes to one model, code edit goes to another, test debugging goes to a reasoning model, formatting goes to a small model, then review goes back again. Each choice may make sense in isolation, but the workflow becomes unstable in behavior and cost.

For coding agents, routing should often happen at phase boundaries.

planning
implementation
test/debug
review
summarization

Subagents create another cost trap. If every subagent gets the same massive context with slightly different instructions, you multiply cost and weaken prefix reuse. Scoped context is better. The review subagent gets the diff and policy. The test debug subagent gets failure logs and relevant files. The search subagent gets query and constraints.

Do not pass the whole conversation everywhere just because it is convenient.

Subagents are not free workers. They are cost multipliers unless the harness controls context shape.

Gateway hot path can become the bottleneck#

Once the gateway owns routing, policy, spend, logging, redaction, and failover, it becomes critical infrastructure. That also means it can become the bottleneck.

Keep the hot path small.

auth
policy check
lightweight token estimate
candidate selection
cache aware routing decision
budget reservation
provider call
stream response

Move heavy work async.

exact cost reconciliation
analytics
dashboard events
cache effectiveness analysis
eval sampling
historical reports

Do not turn the gateway into a data warehouse. Do not make every policy check call another model. Do not synchronously write every analytics event before streaming starts. Do not do exact token accounting across every candidate model if a good estimate is enough for the decision.

Another bottleneck is the session state store. If every request does heavy reads and writes against a central state service, that service becomes part of your p99. Keep the record compact. Store routing metadata, not prompts.

Budgeting has to happen before the request#

Spend tracking after the request is accounting. Cost control needs admission.

Before sending a call, the gateway should estimate max uncached input cost, expected cached input cost, expected output cost, fallback budget, tenant remaining budget, session remaining budget, and agent run remaining budget.

Then it can choose:

allow
allow with cheaper model
allow with lower output cap
ask for context compaction
queue
reject
require approval

This matters more for agents because each step can look acceptable while the full run quietly burns money. A serious agent budget needs limits like:

agent_run_budget
max_steps
max_tool_calls
max_output_tokens
max_uncached_input_tokens
max_wall_clock_time

Again, the control boundary should match the workload, not the API call.

What the router is really optimizing#

The first generation of AI cost optimization is model selection.

Which model should answer this?

The next generation is workload shaping.

Why are we sending this much context?
Why is the stable prefix unstable?
Why are we switching models here?
Why are subagents copying the whole state?
Why are retries paying full price again?
Why is the dashboard counting requests when cost lives in sessions?

That is where the real savings are.

A router cannot fix bad prompt layout. A cheaper model cannot fix an agent that drags the whole repo through every turn. A spend dashboard cannot fix a gateway that only records cost after the money is already gone.

LLM cost is not only a pricing problem.

It is a systems problem.

Final mental model#

Do not route LLM calls like normal API requests.

A normal API request is mostly independent. An LLM request may be one step in a longer computation with cache state, growing context, model affinity, workflow phase, fallback risk, and budget behind it.

While the cache is warm, model affinity has value. When the task changes, quality has value. When context gets too large, compaction has value. When subagents fan out, scoped context has value. When fallback happens, availability has a cost.

And when the cheap model kills cache locality, the cheap model is not cheap anymore.

That is the whole point.

Your LLM router can make cheap models expensive.

References#

  1. OpenAI API documentation: Prompt caching
  2. OpenAI Cookbook: Prompt Caching 101
  3. OpenAI Cookbook: Prompt Caching 201
  4. Anthropic Claude Platform Docs: Prompt caching
  5. Anthropic Claude API Reference: Messages
  6. vLLM Documentation: Automatic Prefix Caching
  7. NVIDIA TensorRT-LLM: KV cache reuse
  8. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
  9. Don’t Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks