Quantization for Serving
Goal
Compare serving precisions by memory use, hardware support, measured quality, and speed instead of assuming fewer bits are always better.
Use fewer bits to reduce weight memory
Suppose you need to store many measurements. Writing 3.14159265 keeps more detail than writing 3.14, but the shorter form needs less space. You lose some numerical detail in exchange for a more compact representation.
Quantization applies a related idea to numbers inside a model. Instead of storing every value in a high-precision representation, the serving system may use fewer bits or a more compact format.
Remember from L8.8 — Quantization for Fine-Tuning: fewer-bit representations trade numerical detail for lower memory use. Here the goal changes from fitting adaptation work into memory to choosing a serving format that also has to perform well on real hardware. Smaller weights use less device memory. They may also run faster when the hardware and serving runtime support that format well.
A rough weight-only example shows why people care. One billion parameters stored at about 2 bytes each need roughly 2 GB just for those parameter values. At 4 bits, the raw parameter storage is about 0.5 GB. That does not mean the whole service becomes four times smaller or four times faster: metadata, cache, workspace, kernels, and hardware support still matter. The example only isolates the storage effect.
1B parameters × 2 bytes ≈ 2.0 GB raw weights
1B parameters × 4 bits ≈ 0.5 GB raw weights
Measure speed and quality separately
But half the bytes does not mean half the latency. Speed also depends on hardware, kernels, model shape, batch shape, and conversion overhead. Measure the real deployment instead of guessing from bit width.
Quality needs measurement too. Quantization changes numbers inside the model and can change outputs. Some tasks tolerate heavy compression; others do not. Test the real task, not only one demo prompt.
Check runtime compatibility
Compatibility is part of release safety. Serving runtimes support different quantization formats on different hardware and versions. Record the runtime, hardware, and quantization settings with the deployment.
QLoRA uses low-bit quantization during model adaptation. Serving has a different goal: choose an inference format and execution path. A training technique does not automatically tell you the best serving format.
Capacity changes when weights shrink
Lower weight memory can change capacity too. A GPU may fit more active generation state, which increases concurrency. Then another limit may appear, such as KV-cache pressure, memory bandwidth, or scheduler overhead.
Compare both quality and operations during rollout. A quantized candidate may pass functional tests but change first-token latency, throughput, unused memory reserve, or errors under load.
Record the serving format
Record the precision, estimated weight memory, supported hardware/runtime, measured quality change, speed metrics, and rollback version. The right choice depends on the workload.
Imagine two serving candidates for the same model. Candidate A uses a higher-precision format and needs more GPU memory, but it already has fast, well-tested kernels on your hardware. Candidate B uses fewer bits and leaves much more memory for cache, yet its runtime path is slower for your batch sizes and slightly lowers accuracy on an important task. The smaller representation is not automatically the better deployment. You would compare task quality, time to first token, throughput, unused memory reserve, supported kernels, and failure behavior under the same workload. Then you can decide whether the extra capacity is worth any quality or latency tradeoff.
Predict
Run the local Lab
Run:
python3 labs/notebooks/level-14/l14-05-quantization.py
The Lab compares serving profiles using hardware support, quality, memory, and measured throughput.
- Run it unchanged.
int8should be the selected eligible profile;int4is excluded for both unsupported hardware and quality below the declared minimum. - In the
int8profile, find"hardware_support": True. Before editing, predict which profile should be selected if that one representation loses hardware support while every quality/throughput value stays fixed. - Change only
int8's"hardware_support": TruetoFalse, then rerun. - Confirm
int8becomes ineligible andfp16is selected. Memory size alone did not determine eligibility.
Loading lab…
Quick Check
Explain it back
Build a comparison table for two precisions. Include memory, measured quality, throughput, compatible hardware/runtime, and the rollout evidence required before changing production.
Key Takeaways
- Quantization can reduce weight memory significantly.
- Lower precision does not guarantee proportional speedups.
- Measure task quality on the serving candidate.
- Compatibility depends on runtime and hardware.
- Quantization changes should use normal rollout and rollback discipline.
Next Lesson
Next, L14.6 — KV Cache and Generation Throughput connects active sequence memory to decoding efficiency and concurrency.
References
- vLLM, Documentation.
- Dettmers et al., QLoRA.
Completion is stored locally on this device.