I would stop exposure to that path now and inspect the two cases. A global average cannot offset an actual restricted-field disclosure. It is also possible that a reviewer mislabeled a permitted field, so the investigation has to separate “credible incident, contain it” from “proven model regression, decide permanent release.” Containment does not require a precise p-value. If traffic routing can pin this cohort to the old system, pause it there and keep unaffected cohorts on their current canary only if the policy boundary and shared components make that safe. If routing is not reliable, stop the broader rollout.

Check what the model actually received and what the user was allowed to see. A disclosure could come from retrieval authorization, a cache scope error, a changed prompt, a renderer, or the model restating material it should have ignored. Compare old and new under the same source revisions, identities, tool schemas, and policy. An offline replay with redacted fixtures can isolate generation behavior, but it cannot reproduce every production permission race. Record model snapshot, prompt, retrieval and policy versions, request cohort, assignment, and the exact field exposed. If the problem is an upstream access leak, rolling back the model alone will not close it.

For the rollout decision I want a cohort-specific gate with two sorts of evidence. Restricted disclosure is an action-level or claim-level rule that can block on a confirmed case. Disclaimers may be a different severity and require a severity-weighted review, including whether the omission changes the customer's decision. Sample representative work from this cohort and targeted boundary cases. Count exposures, not just reviewed answers, and show how cases were selected. With two observed failures and a tiny labeled sample, do not present a stable numerical risk estimate. OpenAI's evaluation guidance recommends production-shaped evaluation and careful human calibration. The slice and release thresholds still need to be set by the product's risk owners.

The release owner should include the regulated workflow owner, the security or privacy authority for the restricted field, and the model platform owner. The platform can keep an option available without setting it as this cohort's default. A sales commitment is not a substitute for the person accountable for the field and user experience. If the old model is also unsafe, pinning it is not automatically the answer. Restrict the data path, add human review, or reduce functionality until the exposure is controlled.

One label is disputed a week later. Reopen that case with a reviewer who has the policy, the user's actual authorization, and the output as displayed, not a cleaned-up transcript. If it was a false positive, correct the incident count and the grader. The other disclosure still needs an explanation. If both turn out to be allowed, we may restart a cohort canary with a repaired evaluation rule, but the dashboard did not become trustworthy retroactively. Check whether similar false negatives escaped review.

I would set an exposure budget for the new cohort canary, a reviewer turnaround, and an automatic pause on a credible restricted-field incident. The comparison uses a stable assignment and reports cohort results alongside global results, with label latency visible. If a cohort is too small to estimate a rare harm from live traffic, use adversarial tests and a deterministic permission boundary. Waiting until the global chart moves means waiting for many small customers to be hurt.