The grammar and the engine can disagree about what a token means. A grammar-constrained decoder masks next-token IDs so the visible output can still satisfy its grammar. The serving engine also has stop-token IDs that terminate generation. If the engine recognizes an end token the grammar backend was never told about, the backend may treat that ID as ordinary allowed vocabulary while the engine treats it as stop. Generation ends before the closing quote and braces. A vLLM issue with a concrete multi-EOS reproduction shows exactly this contract mismatch. It is an issue report and reproduction, not proof that every backend or current release has the bug.

I would first capture the raw generated token IDs, the stop reason, grammar mask at the last step, model generation config and tokenizer's special tokens. Did generation hit an EOS token, a configured stop string, a length limit, or a transport cancellation? Those failure modes need different fixes. If it is the extra EOS case, enumerate all token IDs the engine can use to stop, pass that same set to the grammar backend, and allow those IDs only when the parser is in an accepting state. Also verify that padding or tool-control special tokens cannot slip through the grammar path. vLLM's structured-output documentation describes the public JSON and grammar modes, while the exact stop-token wiring is an implementation detail.

The regression test should put generation inside an open JSON string, make each possible stop ID highly likely in turn, and assert that it is masked until the document is complete. Then check legitimate termination immediately after a complete document, plus max-token exhaustion and client cancellation. A JSON parser at the response boundary should still reject incomplete output. Retrying may help a transient length cut, but it does not repair a mismatch between the grammar and the engine. The deeper lesson is that structural validity depends on the tokenizer, parser state and engine stop rules agreeing on the same token stream.