The field's type defines the comparison. As strings, "9" sorts after "80", while "100" sorts before it by character order. As integers, 9 is below 80 and 100 is above. Elasticsearch's numeric mapping guidance distinguishes numeric fields used for range queries from keyword fields, and its range-query documentation describes field-dependent behavior. A successful query response proves the search engine evaluated its mapped field. It does not prove it implemented the user's intended numeric predicate.

I would inspect the mapping in the live index, the ingestion transform and the query the agent tool sent. Do not rely on an application schema saying risk_score: number if a connector serialized it as text and the index dynamically mapped it to keyword. Replay boundary values such as 9, 79, 80, 81, 100 and 800, plus null, negative and decimal values if allowed. Check whether the score is actually a percentage, a 0-to-1 fraction, or an integer bucket. The CDC field is still called amount. When did its units change? covers a field retaining its name while its units change. This case is a comparison-type error, though unit checks belong in the same ingestion contract.

The durable fix is to index a validated numeric field and query it as numeric. An alias or new index version can support a reindex while old traffic continues, but avoid mixing partly repaired shards or documents under one silent field name. Keep the original text only if it has separate display or exact-match value. Validate conversion and reject or quarantine malformed records rather than coercing them to zero. A tool response can include the interpreted predicate and mapping version so the agent does not imply a numeric guarantee when the index used a string comparison.

One pushback: could the application parse all returned strings and filter them itself? It can as a temporary defense if it retrieves the complete candidate set. It cannot recover the "100" record that the search engine already excluded before returning top results. Rebuild the index or run an authoritative numeric query. This is why a fluent answer over correctly retrieved wrong rows is still wrong.