본문으로 건너뛰기
L8.7

Build a LoRA Adapter

Goal

Implement a tiny LoRA update, verify that its output shape matches the frozen base layer, and test initialization and scaling invariants.

Before tracing the multiplication, keep one picture in mind: the base path produces the model's original answer, and the adapter path produces a correction. The final layer output adds those two pieces together. If the adapter contribution is zero, the layer behaves like the unchanged base; as training changes the adapter factors, the correction can grow.

Use tiny vectors to turn that picture into arithmetic. Read each multiplication as answering one concrete question: first, "what small adapter signal does this input create?" and then, "how is that small signal expanded back into the layer's output shape?" The equation is:

y = W x + s · B(Ax)

This Lesson turns that equation into a small implementation. We will use plain Python lists so the data flow is visible. Real training libraries perform the same operations with tensor kernels and automatic differentiation.

Trace one vector​

Let:

x = [2, -1]
W = [[1, 0],
[0, 1]]

The frozen base output is simply:

W x = [2, -1]

Use rank 1:

A = [[0.5, 0.5]]
B = [[1.0],
[-2.0]]

First:

A x = 0.5×2 + 0.5×(-1) = 0.5

Then:

B(Ax) = [0.5, -1.0]

With scale s=0.2:

adapter contribution = [0.1, -0.2]
adapted output = [2.1, -1.2]

The adapter did not replace the base output. It added a correction.

Initialization can preserve the starting model​

A common design initializes one LoRA factor so the initial product BA is zero. Then the adapted layer initially behaves like the base model. That gives training a controlled starting point:

step 0: adapter contribution = 0
later: learned update becomes non-zero

The exact initialization convention varies by implementation. Test the actual invariant rather than memorizing one library's code.

Shape tests are necessary but not enough​

A broken adapter might still return a vector with the right length. Also test:

  • zero-update initialization reproduces base output;
  • changing adapter factors changes output;
  • changing frozen base weights is not required for adapter output to change;
  • scale zero disables adapter contribution;
  • parameter count matches the configured rank/dimensions.

These behavioral checks make the implementation easier to trust.

Merge versus separate adapter​

Some deployment flows keep base and adapter separate. Others merge the low-rank update into a weight matrix for inference:

W_merged = W + s·BA

Once merged, you must still preserve provenance: which base and which adapter produced the merged artifact? The arithmetic can be simple while artifact identity remains important.

Trace what training would change​

Suppose the current adapter update is too small for one training example. Backpropagation would produce gradients for A and B because they influence the final loss. If W is frozen, the optimizer should not change W. A useful training-step audit is therefore:

before_W == after_W
before_A != after_A (when gradient/update is non-zero)
before_B != after_B (when gradient/update is non-zero)

The exact values depend on initialization and gradients, so equality/inequality should be checked only under a fixture designed to produce a non-zero update. This is stronger than checking a requires_grad flag alone: it verifies actual parameter movement.

Target-module mistakes can create a silent no-op​

A configuration can say “LoRA rank 8” while matching zero real modules because the target-module names are wrong for that architecture. The training loop may still start, but there may be no intended adapter parameters to update. Before training, print or assert:

  • matched target-module names;
  • trainable parameter count;
  • frozen parameter count;
  • one known adapter parameter name.

A non-zero expected trainable count is an important preflight invariant.

Predict

If the LoRA scale is set to zero, what should the adapted layer output?

Build the adapter in the Browser Lab​

This is a learner-authored boundary.

  1. Run the Browser Lab once. The LoRA branch TODO intentionally returns the wrong update and an assertion fails.
  2. Complete the matrix-vector operations for A x and B(Ax).
  3. Apply the configured scale.
  4. Confirm the zero-scale and zero-update invariants.
  5. Change rank in the parameter-count helper and predict the new count before running.

Loading lab…

Quick Check

1. Why can zero-update initialization be useful?
2. What does setting adapter scale to zero test?
3. What provenance is needed for a merged LoRA artifact?

0 of 3 questions answered.

Explain it back​

Trace x → Ax → B(Ax) → scaled adapter update → base + update using your own two-dimensional numbers. State two invariants you would test before trusting the implementation.

Key Takeaways

  • LoRA adds a low-rank correction to the frozen base computation.
  • Intermediate shape and value traces make the implementation debuggable.
  • Zero-update and zero-scale checks provide strong local invariants.
  • Parameter count depends on rank and layer dimensions.
  • Merged artifacts still need base-plus-adapter provenance.

Next Lesson

Complete the mini checkpoint, then study quantization as a separate compression choice that can be combined with adapter training.

References

Lesson actions

Completion is stored locally on this device.

View progress