In multiple-negatives ranking loss, each query has one labeled positive document. The other documents in the batch become negatives for that query. Bigger batches give more comparisons and often harder ones. They also increase the chance that an “other” document actually answers the query.

Think of two support tickets asking the same thing in different words. One has the policy page as its labeled positive. The other's positive is a newer FAQ with the same correct answer. If they land in one batch, the first query is trained to push down the FAQ, and the second may push down the policy. This is a false negative. Sentence Transformers' loss documentation describes the in-batch-negative assumption and false-negative filtering options. Its duplicate-batch sampler guidance helps with exact duplicates, but equivalent answers with different text need more than string deduplication.

I would inspect the highest-scoring “negatives” for a sample of queries and have someone judge whether they are truly irrelevant. Measure the rate by topic, customer and document family. Compare recall against a judged set with multiple acceptable passages per query, not an evaluation set that marks only one document correct. Otherwise retrieval can improve for users while the metric punishes it, or the model can genuinely learn to reject useful alternatives and the metric can hide the harm.

Possible fixes depend on what labels we have. Group equivalent queries or documents away from one another during batch construction, mark all known relevant documents as positives in a multi-positive objective, or mask known false negatives from the denominator. Hard-negative mining needs a check for relevance before treating high similarity as proof of a mistake. Sampling only easy negatives avoids this conflict but can weaken learning.

If the interviewer says, “More negatives should always help,” ask whether those items satisfy the loss's negative assumption. The same large batch that gives a stronger signal can also make the wrong supervision more frequent.