Skip to main content
L8.5

Parameter-Efficient Fine-Tuning

Goal

Explain parameter-efficient fine-tuning as adapting a small trainable component around a frozen base model, compare trainable-parameter counts, and state what freezing does and does not guarantee.

Imagine a huge reference map that is already useful. One way to adapt it for a school event is to redraw the entire map. Another is to keep the map fixed and add a small transparent overlay that marks only the event routes and locations. The overlay works only together with the original map, but it is much smaller to store and change.

Parameter-efficient fine-tuning (PEFT) follows that second pattern. Full fine-tuning updates every trainable parameter, which can make optimizer memory, checkpoint storage, and experiment management expensive for a large model. PEFT instead keeps the pretrained base in the computation while updating a much smaller trainable component. Frozen therefore means "do not update these parameter values during optimization," not "remove this part of the model."

Count what actually changes​

Imagine a toy base model with 1,000,000 parameters. Full fine-tuning:

trainable = 1,000,000

A PEFT method adds or exposes 20,000 trainable parameters:

trainable = 20,000
frozen base = 1,000,000
trainable fraction ≈ 1.96% of total loaded parameters

The exact accounting convention may report the fraction relative to base-only or base-plus-adapter parameters. Record which denominator you use. The important engineering fact is that gradients and optimizer state are required for far fewer parameters.

Frozen does not mean unused​

A frozen base model still participates in every forward pass. Its weights produce hidden representations and outputs. They simply do not receive parameter updates from the optimizer. This distinction matters:

frozen ≠ removed
frozen ≠ zero
frozen ≠ ignored

The adapter learns in the context of the existing base computation.

PEFT changes storage and experiment structure​

With a shared base checkpoint, separate tasks can store separate small adapters:

base model
├── adapter-support
├── adapter-classification
└── adapter-style

That can make versioning and distribution easier than storing one complete copy of the base model for every adaptation. But the adapter alone is not a complete deployable identity. You also need the exact compatible base model and adapter configuration. An adapter trained for one base revision may not be meaningful on another.

Fewer trainable parameters does not mean zero risk​

PEFT can still:

  • overfit the target examples;
  • learn bad labels;
  • create behavior regressions;
  • depend on incorrect formatting;
  • fail to adapt enough capacity for the task.

It is a resource and modularity strategy, not an automatic quality guarantee.

Compare under the same evaluation​

If you want to compare full fine-tuning and PEFT, hold the important evidence fixed:

  • same base model;
  • same train/validation split;
  • same task metric;
  • compatible optimization budget reporting;
  • same behavioral regression suite.

Then compare both quality and resource cost. A smaller trainable fraction is useful only if the behavior is good enough for the use case.

“Frozen” does not mean “irrelevant”​

In PEFT, most base parameters are not updated, but they still perform nearly all of the forward computation. The adapter changes how that existing representation is used. This is why an adapter must be paired with a compatible base checkpoint: the small trainable artifact does not contain a complete model.

Count trainable parameters against the full loaded model-plus-adapter parameter count so the fraction has a clear denominator. Also separate training-memory savings from storage and inference claims. A smaller trainable set can reduce optimizer state substantially, yet activations, base weights, and runtime overhead still consume memory.

Predict

A base parameter is frozen during PEFT. What does that mean?

Run the Browser Lab​

The Browser Lab compares trainable-parameter counts for several adaptation strategies.

  1. Click Run once. The starter reports trainable fraction: 100.0 %, as if the whole model were being trained, so one check fails.
  2. Complete the TODO in trainable_fraction: the adapter is the only trainable part, and the loaded model contains the base plus the adapter.
  3. Click Run again. You should see trainable fraction: 1.961 %, and every check should pass.
  4. Read the storage lines: two full copies: 2000000 versus shared base + two adapters: 1040000. Two tasks can share one frozen base and keep only their small adapters.
  5. Change adapter_parameters = 20_000 to adapter_parameters = 100_000. Before running, predict the new trainable fraction. Optimizer state (the extra numbers an optimizer such as Adam keeps for each trainable parameter) grows in the same proportion.
  6. Click Run. The fraction becomes 9.091 %, and the shared-base total becomes 1200000.

Loading lab…

Quick Check

1. What is the main idea of PEFT?
2. What must usually accompany a saved adapter for reproducible use?
3. Does a small trainable fraction guarantee no regressions?

0 of 3 questions answered.

Explain it back​

For a 10-million-parameter base and a 100,000-parameter adapter, calculate the trainable fraction and explain what is frozen, what still runs in the forward pass, and which artifact identities must be saved.

Key Takeaways

  • PEFT adapts a smaller trainable parameter set around a pretrained base.
  • Frozen base weights still participate in forward computation.
  • Small adapters reduce optimizer/storage cost and improve modularity.
  • Adapter identity depends on base-model compatibility.
  • PEFT still requires held-out and regression evaluation.

Next Lesson

Next, inspect one specific PEFT method: LoRA, where a low-rank update is learned for selected weight matrices.

References

Lesson actions

Completion is stored locally on this device.

View progress