An instruction fine-tune follows the examples well, but at inference it keeps talking until max_new_tokens. The dataset includes an end-of-sequence token after every assistant response. The training loss went down. Someone says the model simply needs more examples of short answers. Maybe, but I would inspect the labels first.

In next-token training, the model only learns to emit a token when that token is a target in the loss. Many pipelines use -100 to ignore a position. Suppose the tokenizer uses the EOS token ID as its padding ID, and a data collator masks every label with that ID to -100. It just removed the true EOS targets along with the padding. The EOS still appears in input_ids, so the model may learn to continue after it, but gets no direct gradient to predict it at the end of an answer. A Transformers collator implementation has used token-ID-based padding masks, and a project issue documents this exact failure when pad and EOS IDs coincide. Check the specific library version and your preprocessing rather than assuming every trainer has the bug.

I would print a few fully rendered training examples as token IDs, tokens, attention masks and labels after all collation. Mark the final assistant token, real EOS, added padding, user spans and tool spans. How many genuine EOS positions have a nonignored label? Is the loss computed on them after the model's shift? Does a chat template use a distinct end-of-turn token instead of EOS? The expected target depends on the model's generation and stopping contract. Do not add a generic EOS where the architecture expects a different terminator.

Fix the mask by distinguishing actual padding positions from a real EOS token, or by using a separate padding token where supported and updating the tokenizer and embeddings deliberately. Verify on one tiny batch that the probability of the correct terminator rises after a few optimization steps. Then evaluate stop reason, output-length distribution and task success, including multi-turn and tool-call examples. Hard truncation at a token budget can hide this issue while still billing for unnecessary decode and sometimes leaking into the next turn.

What if EOS labels are intact? Then inspect whether the serving template appends a start-of-assistant marker different from training, whether stop IDs are configured correctly, and whether decoding suppresses the terminator. The point is to separate a learned stopping failure from a serving stop-rule failure. Ranks process different token counts. Why does averaging their losses change the objective? covers loss weighting across unequal rank token counts. This question concerns one semantically critical target that the loss may never see.