Quantization for Fine-Tuning
Goal
Explain quantization as a lower-precision representation of model weights, distinguish quantized frozen base weights from trainable adapters, and measure approximation error on a small example.
Imagine recording temperatures when your measuring device can store only a small set of values. A reading of 19.8°C might be stored as 20°C. The stored value is close and cheaper to represent, but it is an approximation rather than the original measurement.
Quantization applies this general idea to model numbers. A model can keep the same number of weight positions while representing those values with fewer bits or a restricted numerical format. That distinction matters: fewer bits is not the same thing as fewer parameters. In quantized adapter training, the compact base can stay frozen while separate adapter parameters remain trainable.
Quantization approximates values
Suppose the original weights are:
[-1.0, -0.2, 0.3, 0.95]
A very small toy quantizer might allow only:
[-1.0, -0.5, 0.0, 0.5, 1.0]
Then the values become approximately:
[-1.0, 0.0, 0.5, 1.0]
The representation uses a smaller set of possible values, but introduces error. For each value:
error = quantized - original
Compression is therefore a trade-off, not a free transformation.
Lower precision is not the same as fewer parameters
A model can still have the same number of weight positions. Each position is represented with fewer bits or a quantization scheme. Do not confuse:
- parameter count — how many learned scalar positions exist;
- precision/storage — how many bits or what representation stores them.
LoRA reduces the number of trainable adaptation parameters. Quantization reduces memory/storage/computation cost of represented weights.
QLoRA-style training separates roles
A common idea in quantized adapter fine-tuning is:
quantized base weights → frozen
LoRA adapter weights → trainable at suitable precision
The adapter learns updates while the compressed base provides the pretrained computation. The exact formats and kernels are implementation details that vary across libraries/hardware. The durable concept is the separation between a memory-efficient frozen base and a trainable adapter.
Quantization can change outputs
If weights are approximated, logits can move. A small weight error may be harmless in one case and change a boundary decision in another. That is why evaluation should include:
- before/after task metrics;
- sensitive boundary cases;
- numerical sanity checks;
- memory/resource measurements.
Do not infer behavior quality only from “the model loaded successfully.”
Memory claims need a denominator and scope
If you report a memory saving, specify what is included:
- base weights only?
- optimizer state?
- adapter weights?
- activations?
- gradients?
- runtime overhead?
A four-bit weight representation does not mean the entire training process uses exactly one quarter the memory of a 16-bit full fine-tune. System memory includes many components.
Quantization changes representation, not the number of model decisions
Quantization approximates values with a smaller or lower-precision representation. That can reduce the memory used by frozen weights, but it introduces approximation error and does not automatically shrink every other part of training memory. Activations, adapter parameters, gradients, optimizer state, and temporary buffers still matter.
This is why “4-bit means exactly four times less training memory” is too simple. Measure the system you actually run. In QLoRA-style training, a quantized base can remain frozen while higher-precision adapter parameters are optimized. The learner should be able to point to which values are quantized, which parameters are trainable, and which evaluation detects quality loss caused by approximation.
Predict
Run the Browser Lab
The Browser Lab implements a tiny nearest-level quantizer.
- Click Run once. The quantized values are
[-1.0, 0.0, 0.5, 1.0], butmean absolute error: 0.0is wrong, so one check fails. - Complete
mean_absolute_error: average the absolute difference between each original value and its quantized value. - Click Run again. You should see
mean absolute error: 0.1125, and every check should pass. - Make the levels finer: change
coarse_levelsto[-1.0, -0.75, -0.5, -0.25, 0.0, 0.25, 0.5, 0.75, 1.0]. - Before running, predict whether the error should shrink.
- Click Run. The error drops to
0.0375. The Lab should reportResult: experiment ran: the coarse-level baseline changed, but the invariant checks still confirm nearest-level quantization, correct error computation, and unchanged parameter count. - Notice that the number of weights never changed—still four. Finer levels need more bits to store each weight: five levels fit in 3 bits, nine levels need 4. Quantization trades storage per weight against approximation error, not parameter count.
- Press Reset afterward, after copying your function if you want to keep it.
Loading lab…
Quick Check
Explain it back
Explain the difference among parameter count, trainable parameter count, and weight precision. Then describe why quantization and LoRA can be used together without meaning the same thing.
Key Takeaways
- Quantization approximates weights with lower-precision representations.
- Approximation introduces numerical error that should be measured.
- Parameter count and numerical precision are different quantities.
- Quantized base weights can remain frozen while adapters are trained.
- Resource claims should state what memory components are included.
Next Lesson
Next, move from demonstration targets to pairwise preference evidence: chosen and rejected responses to the same context.
References
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs.
Completion is stored locally on this device.