Evaluation and Quality · Principal
The defense blocks 99% of attacks. What happens when one user can try a hundred times?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Ask what “99%” measures. If it means 99 of 100 independent attempts are blocked, and an attacker can make 100 independent attempts with the same 1% success chance, the chance of at least one success is 1 - 0.99^100, about 63%. That arithmetic is an illustration, not a forecast. Real attempts are correlated, the attacker adapts to failures, and the defender may change state or rate-limit. But one-request block rate does not describe the risk of a session with retries and feedback.
For an agent, success may also be more than a bad sentence. Did it call a forbidden tool, disclose private bytes, commit an external action, or merely produce an answer that a separate check caught? Define the attacker budget, tool access, number of turns, and stopping condition. A human or automated adversary can take a partial refusal, change the prompt, and try another route. The multi-turn jailbreak study is evidence that defenses evaluated on single-turn attacks can behave differently under human multi-turn attempts. Its measured rate belongs to its own models, prompts and defense settings, not our service.
I would report both per-attempt and per-attacker-session success, with the number of attempts and cost to reach success. Keep the full attack trajectory and severity, including attempts stopped by a gateway. Run an adaptive holdout with attackers who have realistic feedback, plus a fixed regression suite so changes are comparable. Slice by tool action, document exfiltration and answer-only behavior. For a severe action, put authorization at the tool boundary and make a model refusal only one layer of defense.
Someone may suggest a tighter rate limit. It reduces the number of tries per identity, which helps, but distributed identities and legitimate users complicate the budget. More importantly, a single successful prohibited action can remain unacceptable. For the most serious outcomes, I want to know the chance that one determined attacker gets through under a realistic budget, and exactly which boundary prevents the action if the model fails.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →