A reranker that reads the query and each candidate together can rescue passages that a cheaper retriever placed in the wrong order. It also does real model work for every pair. If first-stage retrieval returns 100 passages per request and peak load is 500 requests a second, the reranker faces 50,000 query-passage pairs a second before any batching or length variation. Long passages cost more than short ones. A median traffic benchmark with ten candidates is not a capacity test of this path. Sentence Transformers' CrossEncoder documentation exposes candidate scoring and batch-size choices, while its retrieve-and-rerank guide describes the more expensive second stage.

I would split request latency into first-stage search, queue wait before reranking, tokenization, reranker compute, and generation. Watch candidate count and total pair tokens, not just QPS. If utilization is high and queue wait rises sharply near saturation, adding a larger GPU batch may improve throughput but also delay a short interactive request behind long pairs. A shared GPU pool can make the model serving side worse too. Look at p95 and p99 under the real mix of short questions, long policies and image-derived passages.

The obvious fix, cutting every request from 100 candidates to 10, may erase the very exception the reranker was added to rescue. Measure answer-bearing recall at the entry to the reranker as k changes. Then choose a budget by query class, perhaps more candidates for ambiguous or high-risk queries, fewer for exact IDs. Batch by compatible length when that reduces padding, set a queue deadline, and decide what the product does if reranking cannot finish. A first-stage-only fallback can be useful for low-risk questions, but it must not silently make an unsupported “no policy exists” claim.

Would a faster cross-encoder solve it? Maybe. Quantization, ONNX or a smaller model can change the throughput curve, and the quality tradeoff must be tested on the same candidates. A reranker can be perfectly accurate on its own examples and still be the reason the whole search product misses its deadline.