A distributed search can return some results while one shard or remote cluster fails. The HTTP request succeeded in delivering a response, but the query did not cover everything the caller asked to search. Elasticsearch's search API allows partial results on shard failures or timeouts when configured to do so, and exposes timed_out and _shards failure information. Its distributed read documentation notes that partial results can still carry HTTP 200.

If a relevant incident lives on the missing shard, returning zero hits from the other shards says nothing about its absence. Even if the healthy shards return ten useful hits, the assistant cannot confidently claim those are all incidents. Positive claims backed by a returned document may still be defensible after normal permission and freshness checks. Universal or negative claims need complete coverage for the scope they describe. The answer contract should carry search completeness metadata into the model-facing tool result, not flatten a partial response into an ordinary list of hits.

I would inspect attempted, successful and failed shard counts, timeout flags and failures for every branch of a hybrid or federated query. Tag each source as complete, partial or unavailable. For a high-stakes negative answer, fail the retrieval step or retry with a bounded deadline and a healthy replica. If the product offers a degraded answer, state its scope honestly: “I found no incident in the sources searched successfully, but one source is unavailable.” Do not turn that sentence into an assurance that no incident occurred anywhere.

The interviewer could ask whether setting allow_partial_search_results=false solves it. It makes incomplete shard searches visible as errors in that API, which is valuable, but the application still must propagate the failure instead of converting an error to an empty list. Federated search, policy filtering and index freshness remain separate completeness questions. The search returned results, but the policy source timed out handles a policy source timeout despite ordinary search hits. This page is about incomplete coverage inside the search service itself and why a 200 status is not evidence of an exhaustive answer.