Yes, but a token ID from one model is not a token ID from the other. Even matching decoded text may span one draft token and several target tokens. Standard speculative sampling proposes a token sequence, scores it with the target, accepts or rejects using both models' probabilities, and samples a correction after rejection. Those probabilities need to describe the same event in a compatible token space. A direct ID remap or a comparison of unrelated token probabilities breaks the claim that the output has the target model's distribution. The original speculative sampling paper establishes that distribution-preserving goal.

There are principled ways across vocabularies. One approach restricts proposals to tokens that can be matched unambiguously across the two vocabularies, maps them, and still verifies against the target. Another works at the string level with a rejection scheme that accounts for different token boundaries. The heterogeneous-vocabulary research develops lossless algorithms for this setting. A production implementation has its own narrower contract. For example, vLLM's cross-vocabulary TLI documentation describes an intersection of normalized token strings and currently limits that path to greedy draft proposals. I would not silently turn on temperature sampling and assume that particular implementation now has the same guarantee.

The hard case is a proposal such as one draft token for a whole word where the target needs several subword tokens. Which prefix does the target verify, when does it sample a replacement, and what probability mass belongs to the rejected continuation? A correct answer must name an algorithm that handles those boundaries. Retokenizing the final text and calling it verified is too late if the accept/reject step used incompatible probabilities.

For an engine launch I would compare greedy output token by token against target-only decoding on crafted boundary cases, including Unicode, spaces and special tokens. For stochastic mode, compare empirical output distributions on a tiny vocabulary where the target probabilities can be calculated, and test stop-token behavior separately. Then measure accepted target tokens per verification pass and end-to-end latency. An exact but low-acceptance cross-vocabulary drafter may be slower than no speculation. The target distribution is the correctness contract, speed is the reason to use it.