Maybe, but the two displayed numbers are not automatically comparable. A usual language-model loss averages negative log probability per target token. Different tokenizers cut the same text into different numbers of tokens, with different prediction problems at each step. If one tokenizer makes 50 tokens and another makes 25 from the same passage, their per-token losses use different units. Hugging Face's perplexity documentation explicitly notes that tokenization affects perplexity comparisons.

Take a toy 100-byte passage. Model A's tokenizer creates 50 scored tokens with average loss 1.5 nats per token. Model B's creates 25 scored tokens with average loss 2.0 nats per token. Reporting just the averages makes B look worse. Their total negative log probabilities for the passage are 75 and 50 nats respectively, so under consistent conditions B assigns higher probability to that exact text. Divide by the same byte count to get 0.75 versus 0.50 nats per byte, or convert to bits per byte. This example assumes each tokenizer represents the same text reversibly and that both likelihoods include comparable boundaries and targets.

In a real comparison, use the same held-out raw documents, comparable context budgets in content, and explicit treatment of BOS, EOS, whitespace, byte normalization and ignored spans. A fixed token-length window can expose one model to less raw text than the other. Report bits per byte or another common text unit alongside downstream task results. For chat models, also check that different templates do not turn the comparison into two different instruction-following tasks.

This does not mean bits per byte settles every product decision. Tokenization changes inference length and cost, and a model can do better on text likelihood while doing worse on tool calls or rare names. The tokenizer changes midway through training, but the weight files still load. What has been mixed? covers changing tokenizer in the middle of one training run, which mixes incompatible token identities. Here each run can be internally sound. The mistake is reading two per-token metrics as if the token meant the same amount of text.