No, because the outcome of an A user can depend on a B user's artifact. Random assignment still happened, but the usual comparison assumes one user's treatment does not change another user's outcome. Here it does. The observed difference mixes B's direct effect on creators with the benefit or harm B spreads to viewers in both arms. It may understate a full rollout if A users got access to B summaries. It could also overstate one kind of benefit if B users receive a higher fraction of B-created work. Google research on experiments with shared exposure paths studies interference through a bipartite relationship. The specific workspace graph and estimand here are ours to define.

I would map exposure before choosing a new test. Who creates artifacts, who can view or edit them, how long do they persist, and which outcomes depend on them? Track artifact creator assignment, model version, publication time, viewers and edits, with privacy limits. Then distinguish the question we actually need answered. Is it the direct effect on a creator holding everyone else's artifacts fixed? The effect of switching an entire workspace to B? Or the effect of a 20 percent rollout where mixed workspaces continue? Those are different interventions. One user-level A/B number does not estimate all three.

If the product decision is workspace-wide rollout, assign whole workspaces where possible, and keep shared artifacts within those clusters during the measurement window. Balance workspace size, prior activity and project type, and use enough independent workspaces. Existing cross-workspace viewers or copied documents can still create spillover, so record those edges and report exposure. If a company has one giant shared space, workspace randomization gives almost no independent units. Consider an organization-level design, a staged switchback only if artifacts and their effects can be reset or modeled, or a limited experiment with an explicitly narrower claim. Do not present a tiny user-level confidence interval as proof when the actual independent units are few.

The analysis should include all assigned users, including those who never created an artifact, and separate creators, viewers and mixed exposure. Compare task completion, quality, time and downstream correction, not only how many summaries were generated. Record when an artifact was created under A and later edited under B, because its label cannot be inferred from the current owner alone. A straightforward “remove all users who saw the other arm” analysis conditions on a post-assignment event and can select a different population. Treat that as a diagnostic, not the headline causal estimate.

If the interviewer suggests isolating the GPU pool, that fixes a different interference path. The model A/B test shares one GPU pool. Can you trust its latency result? is about latency and capacity contamination through shared compute. This question remains even with perfectly separate GPUs because a generated document is a persistent treatment channel. The right experiment follows the people and artifacts through which the product creates value.

Treatment can reach control users through a shared artifact
Shared artifacts can expose control users to the treatment.