Model and Inference Engineering · Principal
The LoRA adapter works in serving. Why did merging it into 4-bit weights change answers?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The unmerged path effectively applies a frozen quantized base weight plus a separate low-rank update during the forward pass. A merged artifact must represent their sum in some chosen weight format. If that format is 4-bit, many small adapter updates cannot be represented exactly after requantization. Q(W) + Δ and Q(Q(W) + Δ) generally differ, and Q(W + Δ) can differ from both. Here Q is a quantizer and Δ is the adapter's effective weight update. QLoRA's original paper describes training an adapter through a frozen 4-bit base. PEFT's quantization guide describes that arrangement. Neither makes an arbitrary low-bit merge numerically exact.
The deploy pipeline needs to say what it is merging into. Starting from the original higher-precision base, adding the adapter, then quantizing once gives a different artifact from dequantizing an already quantized base, adding the adapter and quantizing again. The latter may be the only available path if the original base was not retained. Some serving stacks support merging in a higher-precision format without requantizing. Check what the actual library, quantizer, group size, scale handling and target modules do rather than assuming one universal merge routine.
To debug the regression, hold prompt formatting, tokenizer, sampling and base revision fixed. Compare logits and outputs from the live base-plus-adapter path, a higher-precision merged reference, and the proposed 4-bit artifact. Measure error by layer and by tasks the adapter was trained to affect. A global average quality score can hide one rare tool schema or customer vocabulary that moved across the decision boundary. Verify tied weights, adapter scaling and whether every target module was merged, since those failures can look like quantization loss.
If the loss comes from requantization, keeping the adapter separate or serving a larger merged artifact may be the right decision. A recent native 4-bit adapter-merging study explores format-specific ways to avoid destructive code changes, but its result is not a blanket promise for every NF4 or other 4-bit model. Quantization frees memory but hurts rare enterprise queries. Ship it? asks whether quantization itself is acceptable for rare queries. This question asks whether a post-training artifact conversion preserved the specific adapted model users tested.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →