Model and Inference Engineering · Principal
Eight GPUs make decode slower than four. Did tensor parallelism cross the node boundary?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Splitting a matrix across twice as many GPUs reduces some local weight work, but the shards have to exchange partial results. In common tensor-parallel Transformer layouts, collectives sit inside the repeated layer computation. Decode produces a small amount of work per token, so even a modest communication startup cost paid across many layers can eat the saved compute time. Moving from four GPUs on one node to eight spread across nodes also changes the link. NVIDIA's NeMo performance guidance recommends keeping tensor parallelism within the fast intra-node interconnect when possible. That is a topology recommendation, not a universal rule that TP across nodes cannot work.
I would not decide from GPU utilization or the number eight. Measure per-token time as local matrix work, collective launch and wait, KV reads, sampling, and scheduler gaps. Check the rank placement and the collective algorithm. Is one cross-node rank late, or do all ranks wait on an expensive all-reduce? Trace both prefill and decode because their compute-to-communication ratios differ. A larger batch can give the sharded matrix work enough volume to pay for communication, while an interactive batch of one may not. Keep output length, prompt length, cache state and actual link topology fixed when comparing.
For a model that fits in four GPUs, two independent four-GPU replicas may serve more requests and keep a lower latency than one eight-GPU tensor-parallel replica. That option consumes two copies of the weights and needs enough aggregate memory. If the model does not fit four GPUs, place TP inside each node and consider pipeline parallelism between nodes, accepting stage bubbles and more complex scheduling. Quantization or a different model size may change the fit calculation. NVIDIA's multi-node deployment guide describes this TP-within-node, pipeline-between-nodes layout as a common configuration, not a guarantee of best performance.
A useful follow-up is whether eight GPUs can ever help. Yes, when memory fit, local bandwidth savings, batch size and interconnect outweigh collective cost. Benchmark tokens per second and first-token and inter-token latency under the intended SLO. The answer may differ between a busy batch fleet and an interactive one.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →