Model and Inference Engineering · Staff
We added tokens to the vocabulary. Why did old prompts change without using them?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
First prove that the old prompt still has the same token IDs. Adding a token can change how surrounding text is segmented by some tokenizers, so “the prompt contains no new tokens” needs an actual token-ID comparison. But suppose the IDs truly are unchanged, the old weights are frozen, and the hidden state is identical. Generation can still change because the output vocabulary now has extra candidates.
The output head gives every vocabulary row a logit. Softmax normalizes over all those rows. In a toy model with two old tokens whose logits are both zero, their probabilities are 0.5 each. Add one new token with logit zero, and all three are about 0.333. Nothing about the old prompt or old logits moved. The denominator did. Top-p membership, sampled tokens and sequence likelihoods can change, and a newly initialized row with a large logit can become an unexpectedly frequent output. Hugging Face's model documentation explains that resize_token_embeddings adds newly initialized vectors and describes mean-based initialization intended to reduce the initial distribution shift for causal models. That mitigation does not promise exact equivalence.
I would compare tokenization of a fixed old-prompt suite, the old-token logits before and after resizing, the new-token logits, and the full normalized distribution. If the old logits themselves changed immediately, inspect whether the embedding/output head is tied, whether a resize retied or copied weights correctly, and whether any configuration or chat template changed. If only probabilities changed, the new rows are the leading cause. If we then fine-tune the expanded model, gradients can alter old shared weights too, so the “same hidden state” argument no longer applies.
We can initialize new rows carefully, mask them from generation until trained if the product allows it, and release the tokenizer, model weights and serving configuration as one compatible version. Do not claim that merely appending unused input tokens is behavior-preserving for an autoregressive model. The tokenizer changes midway through training, but the weight files still load. What has been mixed? asks what happens when a tokenizer changes midway through a training run. This case is narrower: even if old input IDs remain identical, widening the output head changes the sampling distribution.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →