“Higher rank means more capacity” is true, but incomplete. First write down what is actually added to a frozen linear layer: y = Wx + sBAx. A projects into a rank-r space and B maps back. In conventional LoRA, s = α/r. If you change r from 16 to 64 with the same α, the explicit multiplier becomes one quarter of its old value. PEFT documents that convention and its rank-stabilized alternative.

This does not prove the trained adapter's output is exactly one quarter as large. A and B have different shapes, their initialization and learned values change, and the optimizer responds to the new parameterization. With the usual zero initialization of B, both variants start as a no-op. But the scaling also changes gradients flowing into trainable factors, especially early on. More representational directions do not automatically buy a better solution at the same learning rate and training budget.

Rank-stabilized LoRA uses s = α/√r. Under a fourfold rank increase with fixed α, that explicit multiplier halves instead of falling to one quarter. That is a different experiment, not a free fix for every bad run. Keep the target modules, initialization, data order, effective batch size and optimizer settings visible. Check whether the implementation actually uses conventional LoRA, rsLoRA, or layer-specific rank and alpha values.

I would compare validation behavior at equal tokens and roughly comparable resource budgets, plus the norms of BA and the scaled sBA by layer. Look at whether the higher-rank run is under-updating, unstable, or simply fitting a different function. Adapter parameter count, optimizer memory and serving cost rise with rank as well.

If someone asks, “Should alpha always rise with rank?”, I would push back on “always.” Keeping α/r constant is one controlled comparison, and switching to α/√r is another. Neither holds training dynamics fixed. State which property you are trying to preserve, measure the actual update and quality, and choose the rank for the task rather than assuming the largest one wins.