Hedging spends spare capacity to escape an isolated slow replica. If the tail comes from shared saturation, a duplicate enters the same queues and consumes another prefill, KV allocation and perhaps many decode steps. It can push the fleet further into saturation, which triggers more hedges. The Tail at Scale studies when hedging helps large distributed services, while gRPC's hedging guide warns that naive hedging adds load and describes limits and throttling. Neither says to duplicate expensive generation indiscriminately.

I would start with a latency decomposition. Is the request queued, pre-filling a long prompt, decoding slowly, or waiting on a remote tool? Hedging at the gateway after a total-latency threshold may start another expensive prefill when the first replica is already generating. For a streamed answer, switching winners after the user saw tokens is not a transparent hedge. If only pre-first-token work is eligible, state that boundary. Requests with external side effects should not be duplicated unless their tools have their own idempotency and authorization contract.

An effective design sets a small hedge budget, waits until a measured tail threshold, sends the copy to a genuinely independent and healthy capacity pool, and cancels the loser through the actual engine scheduler. Closing the client stream alone may leave GPU work running. Suppress hedges when utilization, queue age or admitted decode load says the fleet is saturated. Compare the extra compute against the user latency saved, and use one deadline for both copies.

I would run open-loop load tests across low, medium and peak arrival rates. Count logical requests and physical attempts separately, track both copies' GPU time and KV residency, and inspect p95, p99, error rate and completed tasks per GPU-hour. A hedging policy that looks excellent at 30% utilization but destabilizes 90% utilization is not ready to ship. The condition that made hedging useful is also the condition the admission controller must prove still holds.