Model Loading and Memory
Goal
Estimate serving memory across model weights, KV cache, runtime workspace, and request state so admission decisions use explicit budgets rather than guesswork.
Budget the whole GPU, not just the weights
Picture a desk with limited surface area. A large textbook stays open all day. Each student request adds sticky notes that must remain available while that request is active. The program also needs scratch paper for temporary calculations, and you deliberately leave part of the desk empty so a small surprise does not knock everything onto the floor.
GPU memory works in a similar way. Model weights are the large, mostly fixed textbook. The KV cache and other request state are sticky notes that grow with active sequences. Runtime workspace is scratch space used while operations run. A safety reserve is memory you choose not to fill during normal admission.
For a simplified example, imagine a 24 GB device where weights use 14 GB, runtime workspace uses 2 GB, and you reserve 2 GB for safety. That leaves about 6 GB for request-dependent state. If one long active request is estimated to need 1.5 GB of that state, four such requests fit the simplified budget but a fifth would cross it. Real runtimes have more details, but the reasoning habit is the same: add the memory pools before deciding what can start.
24 GB total
- 14 GB weights
- 2 GB workspace
- 2 GB safety reserve
= 6 GB left for request-dependent state
Separate fixed and request-dependent memory
Model weights are the large, mostly fixed part. A rough estimate is parameter count × bytes per stored value. Real systems also need some extra space for metadata, sharding, and runtime buffers.
Generation adds the KV cache. It stores attention keys and values for tokens that were already processed. Longer prompts, more generated tokens, and more active requests make this cache grow. The cache saves compute, but it competes with the other memory pools.
Runtime code may also need temporary workspace. Some systems reserve memory in pools. Moving data between host memory and GPU memory can be expensive, so placement matters.
Free memory is not guaranteed capacity
A "free memory" number can also be misleading. Memory may be reserved or split into pieces that are hard to reuse for one large allocation. Serving runtimes often manage blocks or pools to make reuse more predictable.
Admission control decides whether a new request may start now. If the request could push memory above the safe budget, the service should queue or reject it. That is safer than letting one request crash the shared model process.
Plan loading and offloading
Model loading time also affects operations. A replacement process may take seconds or minutes before it becomes ready, depending on model size, storage, compilation, and device initialization. Rollout plans should account for warm-up and not assume a new replica is ready immediately after process start.
Offloading moves some weights or cache between GPU and host memory. It can reduce GPU pressure, but transfers add latency and use bandwidth. It is a tradeoff, not free capacity.
Ask four concrete questions: Which memory pools exist? Which are mostly fixed? Which grow with active requests or sequence length? How much reserve keeps the process healthy? Those answers are more useful than one headline number for "GPU memory used."
A simple budget makes this concrete. Suppose the weights consume most of a GPU but leave several gigabytes free. That remaining space is not automatically safe concurrency. Active sequences still need KV cache and temporary runtime memory, and the process needs reserve for variation. If ten requests fit but the eleventh crosses the safe limit, admission control should queue or reject the eleventh instead of treating the remaining bytes as guaranteed capacity.
Predict
Run the local Lab
Run:
python3 labs/notebooks/level-14/l14-04-model-memory.py
The Lab separates fixed model memory from request-dependent KV-cache memory.
- Run it unchanged with
active_sequences = 8. Recordparts_gband the remaining headroom. - Before editing, predict which memory component changes if the active sequence count doubles. Weights, workspace, and reserve must stay fixed.
- Change only
active_sequences = 8toactive_sequences = 16, then rerun. - Confirm only
kv_cachedoubles. The total now exceeds the 24 GB device budget, so the script reports that the workload does not fit.
Loading lab…
Quick Check
Explain it back
Create a memory budget with weights, KV cache, workspace, and reserve. Explain which parts are mostly fixed and which grow with active requests.
Key Takeaways
- Serving memory includes more than model weights.
- KV cache grows with active generation state.
- Admission should use a safe memory budget.
- Model loading and warm-up affect rollout readiness.
- Offloading trades device capacity for transfer cost.
Next Lesson
Next, L14.5 — Quantization for Serving examines how lower-precision representations change memory, compatibility, and quality.
References
- vLLM, Documentation.
- NVIDIA, CUDA Documentation.
Completion is stored locally on this device.