Skip to main content
L14.5

Quantization for Serving

Goal

Compare serving precisions by memory use, hardware support, measured quality, and speed instead of assuming fewer bits are always better.

Use fewer bits to reduce weight memory​

Suppose you need to store many measurements. Writing 3.14159265 keeps more detail than writing 3.14, but the shorter form needs less space. You lose some numerical detail in exchange for a more compact representation.

Quantization applies a related idea to numbers inside a model. Instead of storing every value in a high-precision representation, the serving system may use fewer bits or a more compact format.

Remember from L8.8 — Quantization for Fine-Tuning: fewer-bit representations trade numerical detail for lower memory use. Here the goal changes from fitting adaptation work into memory to choosing a serving format that also has to perform well on real hardware. Smaller weights use less device memory. They may also run faster when the hardware and serving runtime support that format well.

A rough weight-only example shows why people care. One billion parameters stored at about 2 bytes each need roughly 2 GB just for those parameter values. At 4 bits, the raw parameter storage is about 0.5 GB. That does not mean the whole service becomes four times smaller or four times faster: metadata, cache, workspace, kernels, and hardware support still matter. The example only isolates the storage effect.

1B parameters × 2 bytes ≈ 2.0 GB raw weights
1B parameters × 4 bits ≈ 0.5 GB raw weights

Measure speed and quality separately​

But half the bytes does not mean half the latency. Speed also depends on hardware, kernels, model shape, batch shape, and conversion overhead. Measure the real deployment instead of guessing from bit width.

Quality needs measurement too. Quantization changes numbers inside the model and can change outputs. Some tasks tolerate heavy compression; others do not. Test the real task, not only one demo prompt.

Check runtime compatibility​

Compatibility is part of release safety. Serving runtimes support different quantization formats on different hardware and versions. Record the runtime, hardware, and quantization settings with the deployment.

QLoRA uses low-bit quantization during model adaptation. Serving has a different goal: choose an inference format and execution path. A training technique does not automatically tell you the best serving format.

Capacity changes when weights shrink​

Lower weight memory can change capacity too. A GPU may fit more active generation state, which increases concurrency. Then another limit may appear, such as KV-cache pressure, memory bandwidth, or scheduler overhead.

Compare both quality and operations during rollout. A quantized candidate may pass functional tests but change first-token latency, throughput, unused memory reserve, or errors under load.

Record the serving format​

Record the precision, estimated weight memory, supported hardware/runtime, measured quality change, speed metrics, and rollback version. The right choice depends on the workload.

Imagine two serving candidates for the same model. Candidate A uses a higher-precision format and needs more GPU memory, but it already has fast, well-tested kernels on your hardware. Candidate B uses fewer bits and leaves much more memory for cache, yet its runtime path is slower for your batch sizes and slightly lowers accuracy on an important task. The smaller representation is not automatically the better deployment. You would compare task quality, time to first token, throughput, unused memory reserve, supported kernels, and failure behavior under the same workload. Then you can decide whether the extra capacity is worth any quality or latency tradeoff.

Predict

A 4-bit representation cuts weight memory sharply. Can you conclude latency will also be cut by the same fraction?

Run the local Lab​

Run:

python3 labs/notebooks/level-14/l14-05-quantization.py

The Lab compares serving profiles using hardware support, quality, memory, and measured throughput.

  1. Run it unchanged. int8 should be the selected eligible profile; int4 is excluded for both unsupported hardware and quality below the declared minimum.
  2. In the int8 profile, find "hardware_support": True. Before editing, predict which profile should be selected if that one representation loses hardware support while every quality/throughput value stays fixed.
  3. Change only int8's "hardware_support": True to False, then rerun.
  4. Confirm int8 becomes ineligible and fp16 is selected. Memory size alone did not determine eligibility.

Loading lab…

Quick Check

1. What is a primary serving benefit of quantization?
2. How should serving quality be checked after quantization?
3. Why record runtime and hardware with quantization results?

0 of 3 questions answered.

Explain it back​

Build a comparison table for two precisions. Include memory, measured quality, throughput, compatible hardware/runtime, and the rollout evidence required before changing production.

Key Takeaways

  • Quantization can reduce weight memory significantly.
  • Lower precision does not guarantee proportional speedups.
  • Measure task quality on the serving candidate.
  • Compatibility depends on runtime and hardware.
  • Quantization changes should use normal rollout and rollback discipline.

Next Lesson

Next, L14.6 — KV Cache and Generation Throughput connects active sequence memory to decoding efficiency and concurrency.

References

Lesson actions

Completion is stored locally on this device.

View progress