Ask what “99%” measures. If it means 99 of 100 independent attempts are blocked, and an attacker can make 100 independent attempts with the same 1% success chance, the chance of at least one success is 1 - 0.99^100, about 63%. That arithmetic is an illustration, not a forecast. Real attempts are correlated, the attacker adapts to failures, and the defender may change state or rate-limit. But one-request block rate does not describe the risk of a session with retries and feedback.

For an agent, success may also be more than a bad sentence. Did it call a forbidden tool, disclose private bytes, commit an external action, or merely produce an answer that a separate check caught? Define the attacker budget, tool access, number of turns, and stopping condition. A human or automated adversary can take a partial refusal, change the prompt, and try another route. The multi-turn jailbreak study is evidence that defenses evaluated on single-turn attacks can behave differently under human multi-turn attempts. Its measured rate belongs to its own models, prompts and defense settings, not our service.

I would report both per-attempt and per-attacker-session success, with the number of attempts and cost to reach success. Keep the full attack trajectory and severity, including attempts stopped by a gateway. Run an adaptive holdout with attackers who have realistic feedback, plus a fixed regression suite so changes are comparable. Slice by tool action, document exfiltration and answer-only behavior. For a severe action, put authorization at the tool boundary and make a model refusal only one layer of defense.

Someone may suggest a tighter rate limit. It reduces the number of tries per identity, which helps, but distributed identities and legitimate users complicate the budget. More importantly, a single successful prohibited action can remain unacceptable. For the most serious outcomes, I want to know the chance that one determined attacker gets through under a realistic budget, and exactly which boundary prevents the action if the model fails.