Most people start LLM inference from the API.

Put a model behind HTTP. Add a router. Add GPU workers. Stream tokens. Add autoscaling.

That diagram is fine for a product review, but it does not explain why the system fails.

The pressure starts inside the inference engine. A request is not a normal request anymore. It is a sequence that grows one token at a time. The prompt can be small or huge. The output length is unknown when the request arrives. Two requests that look similar at the API layer can have completely different cost. One finishes in 40 tokens. Another holds KV memory for thousands of decode steps.

So request count is a weak unit.

The system has to think in tokens, KV memory, batch slots, tenant share, queue deadline, cache locality, and GPU time. If the platform only counts requests, it will admit the wrong work, rate limit the wrong tenant, autoscale late, and wonder why GPUs are busy while users still see bad latency.

LLM inference is a scheduler wrapped around an expensive model.

The model generates tokens. The serving layer decides whether those tokens arrive on time.

Workload shape comes first#

Chat, long context RAG, coding agents, reasoning models, and offline batch jobs can use similar model runtimes, but they stress different parts of the system.

Chat cares about first token latency and smooth streaming. Long context RAG stresses prompt processing and KV memory. Agents repeat the same growing prefix, cancel often, and need safe retry behavior around tools. Reasoning models can decode for a long time, so output length variance becomes painful. Offline inference can wait, so the system should pack by length and maximize throughput.

For a serious online assistant, assume this shape:

large model or a small set of large models
streaming responses
multi tenant traffic
some long context requests
agent sessions with repeated growing prefix
strict first token latency
strict time between tokens
mixed prompt and output lengths

This workload fails in the runtime, not at the HTTP layer. GPUs are expensive. Decode needs batching. KV memory is limited. Output length is not known early enough. The system survives only if the scheduler protects those resources.

Prefill and decode pull the system in different directions#

Each request has two main phases.

Prefill reads the prompt, builds KV cache for prompt tokens, and produces the first output token. It is parallel across prompt tokens and uses large matrix operations. A long prompt is expensive, but the GPU can process a lot of that work in parallel.

Decode generates one token at a time. Token 200 cannot exist before token 199. Every decode step reads model weights, reads KV cache, runs attention, and emits one token per active sequence.

At small batch size, decode is economically bad. The GPU reads large model weights and produces very few useful tokens. With a larger batch, the same weight read produces tokens for many active sequences. Batching is the cost model of decode.

But decode batching is limited by KV memory. Every active sequence carries KV. As the sequence grows, KV grows. The scheduler is always between two bad states. If the batch is too small, decode is expensive. If the batch is too full, KV pressure creates tail latency, preemption, or memory failure.

Time to first token means how long the user waits before the first streamed token appears. Time between tokens means how smooth the stream feels after it starts.

A benchmark can show high tokens per second and still produce a bad product if p99 time between tokens spikes. Goodput means useful tokens delivered inside the latency SLO. Raw tokens per second can look good while the product is broken.

KV memory is the first serious capacity number#

KV cache stores key and value tensors for previous tokens so decode does not recompute the whole prompt on every step. It saves compute, but it consumes GPU memory and grows with active context.

Roughly:

KV bytes per token =
2 times layers times KV heads times head dim times bytes per value

For an 8B class model with grouped query attention, this can be around 128 KB per token. Exact numbers change by architecture, but the shape stays the same. KV grows linearly with active tokens.

Take an 80 GB GPU. Suppose weights use 16 GB and runtime overhead leaves around 58 GB for KV.

58 GB divided by 128 KB gives around 475K KV tokens

Now divide by context length:

2K context gives around 237 active sequences
8K context gives around 59 active sequences
32K context gives around 14 active sequences

A few long context requests can eat more capacity than many short chats. A tenant can hurt the system without sending many requests. It only has to hold a lot of KV and decode for a long time.

Long context also changes decode cost. At shorter context, decode is mostly about reading model weights. As context grows, attention over KV becomes more visible. Each decode step reads more historical state. KV reads and attention kernels become part of the bottleneck. Long context is not just a larger request body. It changes the serving profile.

The serving path#

Diagram showing an LLM inference request flowing from client to API gateway, model router, admission controller, engine scheduler, GPU workers, and streaming layer, with side control plane systems for prefix cache, KV blocks, tenant quotas, metrics, autoscaling, and failure detection.

The request path in an LLM inference system. The gateway handles coarse policy, but the scheduler and KV manager decide what can safely run on the GPUs.

The gateway enforces coarse policy. The scheduler makes the runtime decision. It sees live batch shape, free KV blocks, prefix cache state, preemption pressure, queue delay, and current latency.

At fleet scale, the control plane should not try to be perfectly consistent in the hot path. Tracking every KV block movement with strong consistency will turn the router into the bottleneck. The router needs useful locality information, not perfect knowledge. Fresh enough to route well. Cheap enough to stay out of the way.

Continuous batching keeps the batch alive#

Static batching collects a group of requests and runs the group until every request finishes. It works for fixed shape inference. It fits badly with LLM generation.

A batch starts with 32 requests. After 30 decode steps, 20 requests are done. After 200 steps, 5 remain. After 1000 steps, maybe one long generation is still alive. The GPU still runs decode, but useful batch size has collapsed.

The model weights are still read. Fewer useful tokens come out.

Continuous batching rebuilds the batch every decode iteration. After each forward pass, the scheduler checks what finished, what got cancelled, what hit max tokens, how many KV blocks were freed, and which waiting requests can enter. A short request that finishes after 30 tokens releases capacity immediately. A new request can join the next iteration instead of waiting for the longest old request.

The batch becomes a rolling set of active sequences.

That is the foundation of online LLM inference, but the engine still needs to execute mixed work efficiently. Different sequences can sit at different positions. Their attention lengths differ. A naive implementation that treats every operation as one uniform tensor shape will waste work or add overhead.

The usual trick is selective batching. Token wise operations such as projections, feed forward layers, layer norm, and many logit operations can be batched across tokens. Attention is different because each sequence has its own KV length and mask. So the engine batches the linear algebra aggressively, then uses variable length attention kernels for the per sequence attention part.

In simple language: batch what can be flattened, handle attention with sequence aware kernels.

That detail matters because continuous batching is not only a scheduling idea. The kernels have to support the strange batch the scheduler creates.

Chunked prefill keeps streams from freezing#

Imagine 100 users are already streaming, and a new request arrives with a 32K token prompt. Running the whole prefill in one shot may cause visible streaming gaps for everyone else. Delaying that prefill for too long makes the new user wait.

Chunked prefill splits the prompt into smaller pieces and mixes prefill with decode. Each scheduler iteration gets a token budget. Decode tokens for active sequences are reserved first, then remaining budget is used for prefill chunks.

Diagram showing one scheduler iteration where active decode sequences each receive one token while remaining token budget is used for bounded prefill chunks, preventing long prompts from blocking active streams.

Chunked prefill keeps active streams moving while long prompts make progress in bounded pieces.

The long prompt progresses, but existing streams keep moving.

Chunk size is an operating knob. Too small, and overhead becomes visible. Too large, and time between tokens spikes return. The right value comes from production p99 stream latency, queue delay, batch size, and GPU behavior. It is not copied blindly from another system.

There is a useful resource match here. Decode is often bandwidth heavy and can leave compute underused. Prefill is compute heavy. A good scheduler uses some of that compute room without hurting streaming. If it gets greedy, users see pauses.

This is why the scheduler needs latency budgets, not just a queue.

Engine overhead also matters#

Decode repeats tiny steps many times. If the batch is small or the latency target is strict, host overhead and kernel launch overhead start to matter. The GPU can be waiting for the CPU side to launch the next set of kernels.

CUDA graphs help by capturing a repeated decode step and replaying it with lower launch overhead. The catch is that graphs like stable shapes, while continuous batching creates changing shapes. Engines usually bucket shapes or use piecewise graph capture so common decode shapes run fast without giving up all scheduling flexibility.

This detail never appears in the API diagram, but it shows up in latency. A serving engine can have the right scheduling idea and still lose p99 because the runtime path is too dynamic, too host driven, or too shape unstable.

Attention kernels matter too. Prefill uses long prompt attention. Decode uses one new query token per active sequence attending over growing KV. Long context decode needs kernels that parallelize over KV length and avoid bad memory movement. FlashAttention style kernels help prefill. FlashDecoding style kernels help decode at long context. Paged attention kernels gather KV from blocks instead of contiguous memory.

These are not separate from system design. Kernel behavior decides which scheduler choices are actually affordable.

Worked decision one: admission under KV pressure#

This is where the design becomes real.

Assume the engine is in this state:

active decode sequences = 160
decode batch is near the stream latency knee
free KV blocks = 18 percent
tenant A uses 70 percent of its KV quota

new request from tenant A
prompt length = 32K tokens
max output = 2K tokens
prefix cache hit = likely first 24K tokens
current stream latency p99 = close to limit
current first token queue = moderate

A naive system sees free memory and admits the request.

That is how serving systems get unstable.

The request has only 8K uncached prefill after prefix reuse, which helps, but it can still grow by 2K output tokens. Tenant A already uses a large share. Stream latency is near the knee. Free KV exists, but not a lot. If the scheduler admits it aggressively, prefill chunks may hurt active streams, or output growth may push KV into preemption.

A better decision looks like this:

reuse cached 24K prefix only if tenant policy allows it
reserve KV for uncached prompt and expected output growth
keep chunk size small because stream latency p99 is close to limit
charge tenant A for KV and token service
admit only if tenant A remains inside budget
otherwise queue, reduce output budget, or reject early

If free KV were 5 percent and stream latency already broken, the answer changes. Reject or degrade before doing expensive prefill.

A queue stores work. A scheduler decides which work should exist inside the engine right now.

KV cache management decides concurrency#

Diagram showing two LLM sequences mapped to fixed size KV cache blocks, where early prefix blocks are shared by both sequences and later suffix blocks diverge with separate ownership.

Paged KV stores sequences in fixed size blocks, which reduces fragmentation and allows safe prefix reuse.

A naive KV allocator reserves a huge contiguous buffer per request based on max context. Most requests do not use the max. Memory fragments. Effective batch size drops. Cost goes up.

Paged KV fixes this by splitting KV memory into fixed size blocks, often something like 16 tokens per block. Each sequence has a block table from logical token position to physical block. Blocks do not need to be contiguous. When the sequence grows, it gets more blocks. When it finishes or is cancelled, blocks return to the pool.

This is virtual memory thinking applied to inference. vLLM made this idea widely known with PagedAttention.

Paged KV also enables prefix sharing and copy on write. Shared prompt blocks can be reused across requests until the sequences diverge. This matters for chat, agents, beam search, parallel sampling, and repeated tool loops.

The correctness bar is high. Block tables, reference counts, copy on write, and tenant ownership have to be right. A KV bug can become a privacy bug, not just a crash.

The nasty KV failure is not always out of memory. That failure is obvious. A worse version is when the engine stays alive but keeps preempting and recomputing. GPU utilization looks high, yet useful throughput drops because the machine is busy doing work twice. That dashboard fools teams.

Busy is not the same as healthy.

Reducing KV is a serving decision#

Since KV controls concurrency, model architecture choices show up directly in serving.

Multi head attention stores separate K and V for every attention head. Multi query attention shares K and V more aggressively. Grouped query attention sits in the middle and is common because it cuts KV without the full quality cost of sharing everything. If KV heads drop by 4x, KV memory drops by 4x. That can mean more active sequences, larger decode batches, and lower cost.

MLA style attention goes further by caching a smaller latent representation and reconstructing what attention needs. The point is the same from the serving side: reduce KV per token so more context or more users fit in memory.

KV quantization is another lever. If KV can move from 16 bit to 8 bit with acceptable quality loss, active KV capacity roughly doubles. More aggressive formats can help more, but long context quality and model behavior have to be validated. KV errors compound over many decode steps, so this is not a free switch.

The serving engineer should care about attention architecture and KV format because they decide how many live sequences the scheduler can carry.

Prefix caching changes routing#

Prefix caching reuses KV for repeated prompt prefixes. If a request shares a cached prefix, the engine skips prefill for that part and only processes the suffix.

Agents make this valuable. The same system prompt, tool definitions, policy text, history, and previous tool results appear again and again with a small delta.

A simple agent loop looks like this:

system plus tools plus history
system plus tools plus history plus tool result one
system plus tools plus history plus tool result one plus tool result two
system plus tools plus history plus tool result one plus tool result two plus tool result three

Without prefix caching, every step pays prefill on the full growing context. With prefix caching, most steps pay only for the new delta.

Routing now becomes cache aware. If replica A has the warm prefix and replica B does not, sending the request to B wastes prefill and hurts first token latency. If every request goes to A because A has the best cache match, A gets overloaded and the cache benefit disappears in queueing.

A useful router scores locality and load together:

prefix overlap
queue delay
free KV blocks
active KV tokens
tenant pressure
replica health

Round robin throws away locality. Pure cache routing creates hot replicas. The production version blends both.

Prefix caching also has boring correctness traps. Tokenization has to match. Chat templates have to match. Special tokens have to match. Even small differences can destroy cache hits or, worse, create unsafe reuse if keys are computed incorrectly. Cache keys should be based on token sequences, not text that only looks similar. For multi tenant systems, cache scope and ownership matter as much as hit rate.

Sticky sessions help, but hard pinning is dangerous. Prefer the warm replica when it is healthy and not overloaded. Move when pressure is too high.

Worked decision two: cache locality versus load#

Assume a new agent step arrives. The request shares a large prefix with previous steps.

The router sees three replicas:

Diagram showing a router choosing between three LLM replicas by comparing prefix overlap, queue delay, free KV blocks, tenant pressure, and stream latency instead of using simple round robin routing.

Cache aware routing is a tradeoff between prefix reuse and current load. The warmest replica is not always the best replica.

A pure cache router picks A. That can be wrong. A has the best prefix hit, but the queue is already high and stream latency is close to the limit. The cache hit may save prefill, then the request still waits behind a hot replica.

A pure least loaded router picks B. That can also be wrong. B is healthy, but it will recompute most of the prefix. For an agent loop with a large shared context, that extra prefill can hurt first token latency and burn GPU.

A better router compares saved prefill against queue delay and KV pressure. C may be the best choice if its partial prefix hit avoids most recompute while its queue is not as bad as A. B may win if A and C are both close to overload. A wins only if the cache saving is large enough and the queue is still inside the latency budget.

This is the routing version of the same scheduler problem.

Locality is useful until it creates a hot spot. Load balance is useful until it destroys expensive cache state.

Admission control protects the engine#

An LLM server should not admit work only because an HTTP request arrived.

Admitted work consumes KV. It can increase queue delay. It can force preemption. It can hurt other tenants. If it is already too late to meet SLO, accepting it only burns GPU and creates retries.

Admission should look at:

prompt tokens
uncached prompt tokens
requested max output
expected output range
current KV pressure
free block watermark
tenant token budget
tenant KV budget
queue deadline
active batch size
prefix cache probability

Output length is unknown, so the controller cannot be perfect. It needs headroom for running sequences to grow. Too much headroom wastes GPU. Too little headroom creates preemption thrash.

Preemption can drop KV and recompute later, or swap KV to CPU memory and bring it back. Recompute is often simpler and better for shorter contexts. Swap can help for very long contexts if transfer is cheaper than recompute. If preemption is frequent, the admission policy is already sick.

Early rejection is not rude. Serving doomed work is worse.

Fairness should follow tokens and KV#

Request rate fairness does not work well here.

One tenant can send a small number of huge requests and dominate the engine. Another tenant can send many tiny requests and use less capacity. If the limiter sees only requests per second, it protects the wrong resource.

Fairness should use weighted token service and KV occupancy.

A practical system needs tenant level controls like:

token rate budget
KV memory budget
active sequence limit
queue per tenant
priority class
weighted scheduling
agent fanout limit

Batch jobs can use spare capacity. Interactive traffic should win when SLO is tight. Agent traffic needs fanout control because one user action can create many model calls.

The gateway can enforce coarse policy. The engine scheduler needs runtime fairness because it sees actual cost.

Billing and reliability should agree. If customers are billed in tokens but protected by request count, the system has a mismatch.

Overload needs a clear behavior#

A serving system without a rejection path will eventually build a retry storm.

Queue grows. First token latency rises. Clients timeout. Clients retry. Now the system is handling original traffic plus retry traffic. Even after the spike ends, stale work keeps the service unhealthy.

Use bounded queues and deadlines. Reject early if the request is unlikely to meet SLO. Return retry guidance. Apply retry budgets. Shed low priority work before it reaches expensive GPU paths.

Some requests can degrade:

lower max output
route to a smaller model
disable speculative decoding
skip optional reranking
reduce tool fanout
use a cached response where product allows it
return partial output

The timing matters. Degrade before prefill and KV allocation. After the GPU has already burned work, the saving is smaller.

Clean benchmark traffic does not show this. Production traffic has bursts, retries, cancellations, tenant imbalance, and long tails. The overload policy is part of the architecture.

Speculative decoding depends on load#

Speculative decoding uses a cheaper draft path to propose multiple tokens, then verifies them with the larger model. If acceptance is high, the system emits more than one token per expensive model step.

It can reduce latency. It can also hurt throughput.

At low batch size, decode may be bandwidth limited and compute may be underused. Speculation can use that compute room. At high batch size, the GPU may already be closer to compute saturation. Rejected draft tokens then compete with real user work.

So speculation should be controlled by the scheduler. Use batch size, acceptance rate, request priority, and SLO pressure. It is a latency lever, not a free throughput lever.

Scaling to a fleet#

Replication is the first scale out step. Many model replicas sit behind a router.

The router should not use request count as load. It should look at queue delay, active decode batch, active KV tokens, free blocks, prefix overlap, tenant pressure, model availability, and health.

For chat and agents, prefer a replica with warm prefix KV, but do not pin forever. Sticky but movable routing keeps cache benefit while avoiding one hot replica becoming a failure domain.

At fleet scale, the prefix index becomes part of the serving system. It should know where useful prefixes are warm, but it cannot track every block movement synchronously in the hot path. Approximate and cheap is usually better than perfect and slow.

The scale problem is not only adding GPUs. It is placing requests near useful KV without creating hot spots or a central router bottleneck.

When prefill and decode should split#

Start with aggregated serving per replica if the workload allows it. Continuous batching, chunked prefill, paged KV, prefix caching, and good admission control are already a lot of machinery.

Splitting prefill and decode makes sense when one engine cannot balance them well enough, or when the two phases need independent scaling. Long prompts, tight first token and streaming latency together, large fleet scale, repeated long prefixes, and reasoning heavy traffic can push the system there.

The split looks like this:

request

prefill pool
process prompt
write KV

KV transfer
fast network path

decode pool
stream tokens
protect time between tokens

The win is cleaner resource separation. Prefill pool can optimize prompt processing. Decode pool can protect streaming. The pools scale independently.

The cost is also real. KV transfer must beat recompute. Placement gets harder. Failure handling gets harder. Prefill can succeed and decode can fail. KV transfer can fail in between. Overload can move from one pool to another if admission is not coordinated.

A split system needs admission across both pools. If prefill admits more work than decode can absorb, the queue only moved downstream. The global controller has to understand prefill queue, decode queue, KV transfer capacity, and decode KV pressure together. Otherwise disaggregation becomes a more complicated way to miss the same SLO.

Mooncake, Dynamo, and LMCache are worth studying here because they treat KV movement and KV placement as first class serving problems, not as afterthoughts.

Use disaggregation when the workload pays for the complexity.

KV tiering turns memory into a cache hierarchy#

GPU HBM is best for active KV, but it is limited. For long context and agents, reusable KV may be worth keeping outside GPU memory.

A tiered setup can look like:

GPU HBM
CPU DRAM
local NVMe
remote storage

Do not read KV from slow storage during every decode step. That would be terrible. The point is to keep reusable prefix KV around and move it back when transfer is cheaper than recompute.

For short prompts, recompute can be cheaper. For long shared prefixes, reuse can be much cheaper. For strict latency paths, even a cheaper transfer may miss the SLO unless it is already warm.

Once KV is tiered across machines, the system needs ownership, tenant scope, eviction, hot prefix tracking, and transfer scheduling. KV becomes a distributed cache whose hottest tier is GPU memory.

Parallelism follows topology#

Large models may need multiple GPUs. Even models that fit on one GPU may need parallelism for latency or throughput.

Tensor parallelism splits matrix operations across GPUs, usually within a fast NVLink node. It reduces per GPU memory and can improve latency, but it adds communication every layer. The group moves together. If one GPU stalls, the group stalls.

Pipeline parallelism splits layers across GPUs. It can work across slower links, but bubbles appear if the pipe is not filled. It often fits throughput better than low latency interactive serving.

Expert parallelism matters for mixture of experts models. Experts live on different GPUs, and tokens route to experts. The pain is all to all communication and hot experts. If too many tokens hit one expert, other GPUs may be idle while that expert limits throughput.

Context parallelism matters for very long context because a single sequence can exceed what one GPU can handle comfortably.

There is no generic add GPUs answer. Interconnect, batch shape, model architecture, KV placement, and scheduler policy decide whether extra GPUs turn into useful tokens.

Streaming and cancellation are correctness paths#

Streaming is not just frontend behavior.

Once tokens are sent, failure semantics get messy. If a worker dies after 500 tokens, the system cannot unsend them. Restart, resume, or fail all have product impact. For chat, failing may be acceptable. For an agent, partial output may already have triggered a tool action.

Cancellation also has to reach the engine quickly. Users close tabs. Agents stop once they parse a valid tool call. Clients timeout. If the sequence keeps running after the client is gone, it holds KV and steals decode capacity.

Retries need idempotency. The serving layer may retry a failed generation, but the application layer must make tool side effects safe. A repeated tool call should not send two emails, charge twice, or delete twice.

A clean serving contract includes:

client disconnect cancels sequence
cancel frees KV immediately
partial stream has a terminal state
tool calls carry idempotency keys
retries are bounded and deadline aware
partial generations are accounted

Model serving and product correctness meet here.

Security includes KV#

Multi tenant isolation is not only about prompt logs.

Prefix cache can leak through timing. If tenant B gets faster first token latency because tenant A warmed the same prefix, that can reveal information. Incorrect block ownership is worse because one tenant can reuse KV from another tenant.

Logs are another problem. Prompts, retrieved documents, tool outputs, and generations may contain sensitive data. Traces should not casually store raw payloads.

Availability is also isolation. A tenant that consumes all KV or decode slots can deny service to others.

So cache scope matters. Use tenant scoped prefix cache where confidentiality matters. Salt and verify cache keys. Enforce block ownership. Keep raw prompt logging off by default. Treat tenant KV budget as a security and reliability control.

KV is part of the trust boundary.

Determinism is not only temperature#

Temperature 0 removes sampling randomness, but serving can still introduce variation.

The same prompt can be batched with different other prompts. Different batch shapes can change GPU kernel reduction order. Floating point math is not associative. Tiny numeric changes can sometimes change the selected token.

For chat, this may be fine. For evals, debugging, regression tests, and reinforcement learning rollout generation, it matters.

Thinking Machines Lab recently framed this as a batch invariance problem. If the same request produces different numerics depending on who else is in the batch, the serving engine is part of the nondeterminism. SGLang and vLLM now expose deterministic or batch invariant paths, with the usual tradeoff that reproducibility costs performance.

Use deterministic serving mode where needed. Fix model version, prompt template, seeds where sampling exists, and kernel behavior where possible. It should not be the default for every interactive request, but it should exist for paths where reproducibility matters.

Many teams blame the model when the serving stack is part of the nondeterminism.

Failure modes worth designing around#

Long prompt stalls active streams. The model is alive, GPUs are alive, but users see pauses between tokens because prefill was allowed to dominate. Fix it with chunked prefill, token budgets, and stream latency alerts.

KV pressure causes silent waste. The system does not run out of memory. It preempts and recomputes so much that GPU utilization stays high while useful throughput drops. Fix it with watermarks, memory aware admission, tenant KV limits, and preemption rate alerts.

One tenant dominates decode. It may not send many requests. It sends heavy ones. Fix it with weighted token fairness, KV budgets, active sequence caps, and priority queues.

Cache locality fights load balance. The router keeps choosing the warm replica until it becomes hot. Fix it by blending prefix overlap with queue delay, KV pressure, and health.

Autoscaling arrives late. New workers need time to load weights, initialize memory, warm runtime paths, and register. Fix it with warm pools, minimum capacity, and scaling from SLO headroom and token rate, not only GPU utilization.

Deploys kill streams. A worker is terminated while users are receiving tokens. Fix it with draining, no new admissions, finish or migrate active sequences, then remove the worker.

Retries create overload loops. The queue grows, clients timeout, retries add load, and stale work keeps the system sick. Fix it with bounded queues, deadlines, retry budgets, and early rejection.

These are normal failures. They show up when traffic becomes real.

Observability should follow the scheduler#

A dashboard centered on request count and GPU utilization will miss important failures.

Track the metrics the scheduler cares about:

time to first token p50 p95 p99
time between tokens p50 p95 p99
queue time
prefill time
decode time
tokens per second
goodput inside SLO
batch size distribution
active sequences
active KV tokens
free KV blocks
KV allocation failures
preemption rate
prefix cache hit rate
cache miss cost
tenant token usage
tenant KV usage
rejection rate
cancellation rate
retry rate
GPU memory pressure
GPU health errors
router decision reason

The useful view is usually a timeline. Batch size, KV pressure, prefix hit rate, preemption rate, first token latency, and stream latency together.

If stream latency spikes when prefill chunks grow, chunking is too aggressive. If first token latency spikes while prefix hit rate drops, routing or cache eviction may be the issue. If GPU utilization is high but goodput is low, the system may be recomputing, serving doomed requests, or running past the latency knee.

Ask what the GPU is busy doing.

Serving useful tokens inside SLO is capacity. Recomputing after bad admission is waste.

Autoscaling has to respect cold start#

New inference workers are not cheap to start. They may load large weights, allocate memory, warm kernels, capture optimized runtime paths, and register with routing. Reactive scaling can arrive after users already see bad latency.

Autoscaling should use leading signals:

queue growth
first token latency headroom
stream latency headroom
incoming prompt token rate
expected decode token rate
KV pressure
prefix cache hit rate
tenant burst signals

Interactive systems usually need warm pools and a minimum capacity floor. Scale to zero can work for batch or internal low SLO jobs. It is usually bad for a user facing assistant.

Autoscaling should also understand model mix. Adding capacity to a small model pool does not help if the large reasoning model is saturated. Adding prefill capacity does not help if decode is the bottleneck.

Cost is inside the scheduler#

GPU time is the bill.

Continuous batching increases useful tokens per weight read. Paged KV increases active sequences. Prefix caching avoids repeated prefill. Chunked prefill protects streaming while keeping compute useful. Quantization reduces weight bandwidth and KV size. Cache aware routing avoids avoidable prefill. Admission control avoids spending GPU on work that already missed deadline.

The objective is:

maximize useful tokens per GPU hour
while respecting first token latency, stream latency, correctness, and isolation

Raw throughput alone is a trap. A benchmark can push max tokens per second by allowing terrible p99 latency. Production usually runs near the throughput latency knee, where batching is large enough to be economical but not so large that streams feel broken.

The cost model is embedded in the scheduler.

A design path that survives review#

Start with aggregated serving per replica unless the workload proves otherwise.

Each replica should have continuous batching, chunked prefill, paged KV, prefix caching, cancellation, and memory aware admission. The router should blend load with prefix locality and use sticky but movable session routing. The gateway should enforce tenant identity, max tokens, idempotency keys, and coarse token limits. The engine scheduler should enforce runtime fairness with token and KV accounting.

Add bounded queues and early rejection from the start. Add observability from the start. Chunk size, watermarks, admission thresholds, and tenant weights cannot be tuned from guesses.

Add deterministic serving mode for evals and debugging paths, while keeping the main path optimized for interactive goodput.

As traffic grows, add a better global prefix index, cache aware routing, warm pools, predictive scaling, and KV tiering. Split prefill and decode only when one engine cannot manage interference well enough or when independent scaling gives clear cost or latency wins.

That is the shape worth defending. Start with the real bottlenecks. Add complexity only when the workload earns it.

The model gives you tokens.

The serving system decides whether those tokens show up on time.

References#

Orca, OSDI 2022 Iteration level scheduling and selective batching for transformer serving.

vLLM and PagedAttention, SOSP 2023 Paged KV cache, memory efficient batching, KV sharing, and practical high throughput LLM serving.

Sarathi Serve, OSDI 2024 Chunked prefill, stall free scheduling, and the throughput latency tradeoff in online LLM inference.

DistServe, OSDI 2024 Disaggregated prefill and decode serving, with separate resource allocation for first token latency and decode latency.

Mooncake, FAST 2025 KV cache centric disaggregated serving architecture behind Kimi, with KV aware scheduling and early rejection under overload.

NVIDIA Dynamo KV aware routing, KV block management, KV offload, and distributed inference orchestration for large scale LLM serving.

LMCache KV cache offload and sharing across inference engines, including vLLM and SGLang integration.

VTC, OSDI 2024 Virtual Token Counter scheduling for fair LLM serving across tenants using token based service accounting.

FlashAttention and FlashDecoding Attention kernels for efficient long prompt prefill and long context decode.

Medusa and EAGLE Speculative decoding techniques that trade extra compute for lower decode latency when acceptance is high.

S LoRA and Punica Serving many LoRA adapters on a shared base model with specialized batching and adapter memory management.

Thinking Machines Lab, Defeating Nondeterminism in LLM Inference Batch invariant kernels and the argument that serving systems can introduce nondeterminism even at temperature zero.

vLLM and SGLang deterministic inference docs Practical deterministic and batch invariant inference paths in modern open source serving engines.