Before arguing about the GPU, look at which requests entered each measurement. Say Model A saw 98% short prompts at 100 ms and 2% long prompts at 10 seconds. Model B saw 99.5% short prompts at 120 ms and 0.5% long prompts at 12 seconds. B is slower in both buckets, yet its overall p99 is around 120 ms while A's is 10 seconds. At the 99th percentile, A's tail includes long prompts and B's does not. These are simplified point masses, but the reversal is real. You cannot average bucket p99s to recover the overall p99 either.

Why did the mixes differ? A randomized A/B test should make them similar in expectation, though small samples can still differ. Maybe routing only sent easy requests to B, a compatibility filter excluded long prompts, timeouts disappeared from one dashboard, or the two arms ran during different hours. Compare the assigned population, not just completed requests. Check prompt and output lengths, image and tool use, tenant, region, cache hit, batch size, admission failures and deployment window. If assignment is truly randomized and the distribution is comparable, the paradoxical aggregate deserves a closer look at quantile estimation and sample size.

For a decision, keep user-visible p99 as an SLO metric, but compare models on a matched workload or reweight to a declared target traffic mix. Report latency distributions within useful request classes and completion/failure rates. A single scalar p99 of a shifting population cannot tell us whether B is faster at the same work. Google's SRE guidance discusses separate objectives for heterogeneous workload classes. It supports separating classes, while the numeric example here is just arithmetic.

The follow-up is that traffic mix can be part of the product effect. If a model's behavior causes users to make more long requests, “hold the mix fixed” estimates a direct serving comparison, not the full system outcome. Report both questions clearly. A flattering aggregate is sometimes a real change in user behavior, sometimes an assignment bug, and sometimes a dashboard that lost difficult requests. The data lineage decides which.