Model and Inference Engineering · Staff
The 7B weights take 14 GB. Why won't Adam training fit on a 40 GB GPU?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Fourteen GB counts one BF16 copy of seven billion weights. Training also needs gradients and optimizer state, and backprop needs saved activations. Adam keeps a first and second moment per trainable parameter. In a common mixed-precision recipe, each moment is FP32. That alone is 8 bytes per parameter, or about 56 GB for 7B. Add roughly 14 GB for BF16 gradients and the 14 GB model copy. We are already near 84 GB before activations. Some recipes also keep FP32 master weights, another 28 GB, which brings those persistent model states to about 112 GB. These are decimal GB estimates, not a universal PyTorch allocation trace. PyTorch's Adam implementation defines the two moving averages, and torchtune's memory guide explains precision and activation storage.
I would ask what actually lives on that GPU. Is this full parameter training or LoRA? Are parameters BF16 or FP32, are gradients reduced or retained in FP32, is there a master copy, is the optimizer fused, and is the state already initialized? A profiler snapshot after the first optimizer step matters because Adam state is often allocated lazily. Also account for activation peaks as sequence length and microbatch grow, temporary buffers during optimizer step or collective communication, and allocator reserve. Loading the weights successfully says little about whether backward plus step fits.
There are different fixes for different terms. Shard optimizer state, gradients and parameters with ZeRO or FSDP across GPUs, remembering that all-gather and communication still have a peak and a throughput cost. PyTorch's FSDP guide describes that division. Reduce microbatch size or checkpoint activations if activation memory is the blocker. Offload state to CPU if the bandwidth and step-time cost is acceptable. Train adapters if the task permits fewer trainable parameters. Quantizing an inference checkpoint alone does not make a full Adam training run fit. The useful answer is a measured per-state budget and a peak trace for the exact training recipe.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →