First find out what was quantized. FP8 weights, activations and KV cache have different error paths. Suppose this deployment quantizes weights and activations using scales estimated from a small calibration set. A rare prompt can create activation values outside the range that set represented. The values may clip or lose useful resolution when mapped into FP8. A small error at one layer can change a near-tie next-token choice, and generation then follows a different path. TensorRT-LLM's FP8 guide describes calibration data for determining scales when quantizing an FP16 checkpoint. That does not mean every FP8 implementation uses the same scale or every changed answer is a calibration fault.

I would run the same prompts through the reference and quantized models with identical tokenization, template, decoding and weights. Start at the first divergent logits, not the final paragraph. Compare activation ranges and quantization error by layer and by phase. Does the error appear on prefill, after a long prefix, at a particular channel, or only once the KV cache is read? Record clipping rates and the scale used, along with the actual precision of each path. If the scales look fine, check kernel or cache conversion before increasing the calibration dataset blindly.

Average answer accuracy is a poor release gate for a rare but costly class. Build slices for long inputs, unusual scripts, tool syntax, numerical answers and known activation outliers. Recalibrate on representative traffic without leaking private prompts, or use a finer-grained scale or higher precision in the sensitive layers if the engine supports it. Quantization-aware training is another, much more expensive option. Compare quality at the same latency and memory budgets, including the throughput loss from the proposed fix. Keeping one narrow layer at higher precision can beat throwing away the entire speedup.

There is a second trap here. If we change the prompt set until this FP8 build passes, that set is no longer a clean final test. Preserve a separate holdout and track disagreement severity, not only total exact-match rate. A tiny average movement can hide a large regression in a small group of high-consequence prompts.