Model and Inference Engineering · Staff
The 4-bit weights fit on one GPU. Why did tokens per second get worse?
The question
Interview question
An inference team switches from a higher-precision model to a 4-bit weight-only checkpoint. It fits on one GPU, but decode throughput is lower and p99 rises for real request sizes. An engineer expected fewer weight bytes to guarantee faster service. What is missing from that argument?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Memory footprint and execution time are different measurements. The quantized weights need scales and possibly zero points, and a kernel must unpack or dequantize them as it computes. Depending on hardware and implementation, that can be fused efficiently or can add instructions, temporary buffers and less favorable memory access. A kernel optimized for a different group size, hidden dimension, batch shape or GPU may underperform a well-tuned higher-precision kernel. vLLM's quantization documentation describes multiple formats and hardware support. It does not imply that every 4-bit path beats every higher-precision path.
I would first split the latency. Is the regression in prefill, decode, first-token time, scheduling or weight loading after cold start? Profile actual kernels, occupancy, bandwidth, launch count and time spent in conversions. Compare on identical prompts, output lengths and concurrency, with the same number of GPUs and the same batching policy. A switch from two GPUs to one also changes tensor-parallel communication and available compute. That can help or hurt. If the model now admits more concurrent requests because weights occupy less memory, p99 can worsen from queueing or KV pressure even when an individual kernel is faster. Measure per-request service time separately from waiting time.
Weight-only quantization generally does not shrink the KV cache unless KV uses its own lower-precision representation. On long contexts, KV reads and attention can dominate parts of decode. Check whether saved weight memory was actually reassigned to KV blocks, and whether that changed batch sizes. Compare quality too, especially on the difficult cohort that motivated the model. A fast but degraded completion that triggers more retries can lose at cost per successful task. Quantization frees memory but hurts rare enterprise queries. Ship it? asks whether to ship a quantized model after rare-query quality loss. This asks why the serving performance premise failed even before that decision.
Would I revert immediately? If the latency SLO is broken, route traffic to a known-good configuration while measuring the new path. Then test supported quantization kernels and hardware, group sizes and serving shapes under a realistic mix. A lower bit width is a storage property. The deployment win is an empirical result across throughput, p99, capacity, cost and task quality.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →