The input is known after tokenization. The output is a random variable bounded by the configured maximum, stop conditions and the model's context limit. Each generated token needs another decode step and extends live KV state. So a request that fits the prompt budget can still occupy decode slots and memory for a long time. A per-iteration scheduler token budget, such as the one described in vLLM's scheduler configuration, controls work scheduled in an iteration. It is not automatically a tenant-level promise across the request's whole lifetime.

For a hard tenant quota, I would reserve capacity for a defensible upper bound at admission, debit actual tokens as they arrive, and release unused reservation on EOS, error or cancellation. The exact reservation need not always equal an enormous client-supplied maximum. Product limits can cap that maximum, and the scheduler can use bounded incremental leases with enforcement at every decode step. If the quota cannot be extended, stop cleanly at the agreed output boundary rather than silently exceeding it. API policies differ. OpenAI's rate-limit guide says a configured maximum can affect rate-limit estimates, while Anthropic's current API guidance says its output-token limit counts actual tokens in real time and does not charge that limit against max_tokens in advance. A gateway must define its own admission contract instead of assuming either provider's rule.

Quota accounting and fleet fairness need separate controls. Reserve per-tenant budget without granting unlimited concurrent live sequences. Limit active decode slots or weighted work per tenant, and account for a request's accumulated output, not only its prompt. Reconcile on every terminal path, including disconnected streams and engine failures. A leaked reservation blocks future traffic. An uncharged stream lets a tenant exceed its share.

If the interviewer proposes reserving every request's full maximum forever, I would ask what happens to utilization when most answers stop early. Hard bounds protect the quota, but excessive static reservation can reject useful work. Compare admission rejects, quota overshoot, live KV, per-tenant inter-token latency and unused reservations under a realistic length distribution. The router saves money until hard requests exhaust the premium-model quota. What now? covers a router spending its scarce premium model quota. This page is about output growth during one admitted request and the accounting boundary between gateway and decoder.