Model and Inference Engineering · Staff
A fused RMSNorm kernel is faster. Why do small activations change the answer?
The question
Interview question
A model checkpoint specifies RMSNorm with epsilon `1e-6`. A fused inference kernel uses `1e-5`, and its reductions run in a lower precision than the reference path. Most prompts look fine, but a narrow slice changes. The team calls epsilon an inconsequential stability constant. Is it?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
RMSNorm scales a vector approximately as x / sqrt(mean(x²) + epsilon), then applies a learned per-channel weight. The RMSNorm paper gives the normalization idea, and PyTorch's RMSNorm documentation shows where epsilon enters the denominator. If mean(x²) is large relative to both epsilon values, the difference is small. If the activation is tiny, epsilon dominates. The denominator changes from roughly sqrt(1e-6) to sqrt(1e-5), a factor of about 3.16 in that limiting case. That is not a harmless rounding difference. Later layers may amplify or suppress it, so we should measure rather than claim every output will fail.
I would compare the reference and fused kernels with fixed input tensors at realistic shapes and dtypes. Include tiny values, ordinary values, a large outlier and mixed signs. Compare mean(x²), reciprocal square root, weighted output and gradient if the same kernel is also used in training. Find out whether the reference accumulates squares in FP32 and when it casts back. A lower-precision reduction can change the statistic even with the same epsilon, especially at extreme magnitudes. Confirm the weight is applied in the same order and precision. A test that only checks shape or average error over random normal inputs can miss the boundary where epsilon matters.
The checkpoint's normalization parameters belong to its architecture contract. Do not let a serving library silently substitute its own default epsilon. Pass the checkpoint value explicitly and assert it when loading. Then run layer-level equivalence tests and logit comparisons across the full model at the intended precision and batch shapes. Set tolerances based on expected numerical differences, not on a desire to make the test pass. If a quality regression remains after the kernel matches the reference operation, investigate other fused-path assumptions.
An interviewer may ask whether any epsilon is mathematically acceptable. We could train or fine-tune a different architecture with a different epsilon. That is not the same as swapping the constant after training and claiming identical behavior. The checkpoint loads after a SwiGLU kernel rewrite. Why did model quality fall? catches gate/up weights swapped in a gated MLP. Here every weight can be in the right place while one small constant and one accumulation dtype quietly change the computation.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →