An asymmetric retriever may put a different prompt or task marker before query text and document text, or route them through different modules. Training teaches the query representation to find the document representation under that contract. A generic encoder call may omit the query instruction or route. Sentence Transformers' semantic search guidance recommends encode_query() and encode_document() for asymmetric search. It also says that for models without special query/document behavior, the methods can produce identical embeddings. So do not claim every encode() call is wrong. Check the actual model and configuration.

I would replay a held-out query through both code paths with identical text and preprocessing. Inspect rendered input or task prompt, tokenizer output, selected module and resulting vector. Compare norms and cosine similarity between the query embeddings, then compare ranks against the same frozen document index. This should reveal whether the refactor changed the vector, whether a later normalization step changed it, or whether a different problem caused the drop. Also check the indexed documents were truly generated by the expected document path and model revision. Vector dimensions and a shared model name are not enough to establish compatibility.

The fix is a versioned embedding contract with explicit roles: model revision, tokenizer, query/document prompts or routes, pooling, normalization, output dimension and similarity metric. Store the document-side contract with the index. A query service should refuse or flag an incompatible contract instead of silently sending vectors into an existing collection. If the query path alone changed, restoring it may recover ranking without rebuilding documents. If the document encoder or pooling changed, a coordinated reindex or dual-index migration may be required. Measure recall on exact, semantic and tail queries, with downstream answer support, before cutover.

Could we compensate by tuning a score threshold? Not reliably. The query vector may have moved in different directions for different intents, changing ranks, not just scaling scores. Half the index uses the new embedding model. Are the search scores still meaningful? asks whether scores remain meaningful when half an index uses a new model. Here the document index is uniform. The mismatch is between the model's two roles at query time, and the bug can hide behind perfectly valid vector shapes.