The reranker cannot score a distinction it never receives. A cross-encoder jointly encodes the query and passage, and the combined pair has a length limit. If a preprocessing policy trims the query, the model may see a broader question and prefer the generic policy. Sentence Transformers' CrossEncoder guidance documents the max_length behavior for longer inputs, while Transformers' truncation guidance explains options for paired sequences. The exact truncation side and allocation between query and document depend on the tokenizer configuration. Inspect actual token IDs and decoded text instead of guessing.

I would trace the candidate set before and after reranking for this query. Log safely the query length, passage length, truncation flag and token budget split, with source handles. Reconstruct the specific serialized pair the cross-encoder scored. Did it lose the negation, account type, date qualifier, or exception paragraph? Sometimes the query survives but the passage's decisive sentence is at the end and gets cut. Either way, NDCG over broad queries can stay fine while hard qualifier cases fail. Test ranking on matched pairs that differ only in a crucial word such as “before” versus “after” and on real policy exceptions.

For a short query, reserve enough budget to keep it whole and use remaining tokens for a focused passage. If passages are too long, split them along clause boundaries and score the relevant spans with heading and product context. Blindly raising the token limit can make latency and cost unacceptable and may still bury the decisive qualifier. Preserve a first-stage path that does not let a lossy reranker discard all exception candidates, such as a diversity or mandatory-candidate guard when a query explicitly names a clause. That guard should be evaluated, not hard-coded from this one incident.

An interviewer might say retrieval recall@20 is high. That tells us the right candidate entered the pool, which is valuable, but it does not prove it survived the final ranking and context budget. The reranker never sees the policy exception asks what happens when an exception never reaches the reranker. Here it arrived, and the model was asked to judge it under a different question than the user wrote.