Memory footprint and execution time are different measurements. The quantized weights need scales and possibly zero points, and a kernel must unpack or dequantize them as it computes. Depending on hardware and implementation, that can be fused efficiently or can add instructions, temporary buffers and less favorable memory access. A kernel optimized for a different group size, hidden dimension, batch shape or GPU may underperform a well-tuned higher-precision kernel. vLLM's quantization documentation describes multiple formats and hardware support. It does not imply that every 4-bit path beats every higher-precision path.

I would first split the latency. Is the regression in prefill, decode, first-token time, scheduling or weight loading after cold start? Profile actual kernels, occupancy, bandwidth, launch count and time spent in conversions. Compare on identical prompts, output lengths and concurrency, with the same number of GPUs and the same batching policy. A switch from two GPUs to one also changes tensor-parallel communication and available compute. That can help or hurt. If the model now admits more concurrent requests because weights occupy less memory, p99 can worsen from queueing or KV pressure even when an individual kernel is faster. Measure per-request service time separately from waiting time.

Weight-only quantization generally does not shrink the KV cache unless KV uses its own lower-precision representation. On long contexts, KV reads and attention can dominate parts of decode. Check whether saved weight memory was actually reassigned to KV blocks, and whether that changed batch sizes. Compare quality too, especially on the difficult cohort that motivated the model. A fast but degraded completion that triggers more retries can lose at cost per successful task. Quantization frees memory but hurts rare enterprise queries. Ship it? asks whether to ship a quantized model after rare-query quality loss. This asks why the serving performance premise failed even before that decision.

Would I revert immediately? If the latency SLO is broken, route traffic to a known-good configuration while measuring the new path. Then test supported quantization kernels and hardware, group sizes and serving shapes under a realistic mix. A lower bit width is a storage property. The deployment win is an empirical result across throughput, p99, capacity, cost and task quality.