The model consumes token IDs, not the prompt as a human sees it. Chat templates may insert role markers, BOS, EOS or an assistant-start prefix. If a second tokenization step inserts the same special token again, the checkpoint sees a sequence it was not intended to see for this path. Hugging Face's chat template guidance explicitly warns to use add_special_tokens=False when tokenizing a separately rendered template that already contains such tokens. That is a concrete implementation contract, not proof that every duplicate BOS has the same effect on every checkpoint.

I would dump token IDs and their decoded special-token names at the exact model boundary for one fixed conversation, from both the evaluation path and the serving path. Include the assistant generation prefix and any tool-role markers. Compare the first divergent token, not just the rendered text. Check whether the tokenizer treats a string like <s> as a special token or as ordinary characters in this configuration. Then compare prefill logits under the two sequences with the same weights, positions, attention mask and numerical path. If the double BOS is the cause, fixing tokenization should remove that input difference. A different divergence could still come from a template revision, pad handling or stop configuration.

The repair is to have one owner for serialization. Either call the chat-template API with tokenization enabled, or render once and tokenize without adding special tokens again. Pin the template and tokenizer revision with the checkpoint. Test system, developer, user, assistant and tool turns, including empty assistant prefixes and multi-turn history. A simple single-user prompt may pass while tool use breaks because a role boundary token moves. Log token counts and a safe special-token trace so a future rollout can detect a second BOS without storing private prompt text.

The interviewer may ask if two BOS tokens are syntactically invalid. Usually the tokenizer can represent both. That is exactly why this error survives basic health checks. The tokenizer changes midway through training, but the weight files still load. What has been mixed? covers changing a tokenizer midway through training. The DPO reference and policy score different prompt tokens. What does the margin mean? covers policy and reference scoring different prompt serializations during DPO. This case is a serving pipeline that gives one checkpoint the wrong initial sequence even though the model and tokenizer files themselves are unchanged.