The answer can be factually correct while the route taken to produce it violated an access rule. A tool read is an action with consequences even when its content does not appear in the final response. ATBench evaluates safety at the agent-trajectory level, reflecting the difference between output-only assessment and intermediate behavior. For this product, define two separate outcomes: task quality and constraint compliance. A release gate for protected data should require both. Do not average a perfect answer score with an unauthorized read and call the result acceptable.

I would build evaluation cases with realistic permission boundaries and instrument tool requests and effects, not just model messages. The grader should know which account was in scope, which objects were touched, what authorization decision the tool made, and what data the agent actually received. Distinguish an attempted forbidden call that the server rejected from a successful unauthorized read. Both reveal something about policy behavior, but they have different exposure. Test whether the agent can recover after a denied access without inventing an answer.

The evaluator cannot rely only on an LLM reading a natural-language trace. Use deterministic checks for resource IDs, permissions, side effects and policy epochs where possible. Human review can resolve ambiguous user intent or harm, and an LLM judge can help triage the remaining traces, but neither should be the source of truth for whether a specific database read was authorized. Keep the replay environment and permissions versioned, since the same call can be allowed under one role and forbidden under another. Report quality pass rate and safety-violation rate by task type, plus severity and confidence intervals for rare failures.

What if the final answer never used the private content? The read still happened. If the policy has a narrower goal of preventing disclosure to the user, measure that separately, but do not silently relabel an unauthorized access as safe. How do you evaluate an agent when several trajectories are correct? asks how to score agents when several trajectories can be correct. How do you replay an agent eval when tools change the world? asks how to replay actions that alter the world. This question asks for a non-negotiable trajectory constraint that a good final answer cannot erase.