Skip to main content
L6.1

Decoder-Only Language Models

Goal

By the end of this lesson, you can explain a decoder-only language model as a system that repeatedly predicts the next token from the text available so far, and you can explain why future tokens must be hidden from an earlier prediction.

Start with one prediction​

Suppose the text so far is:

the cat

The text already available is called the prefix. A language model can assign a score to every possible next token in its vocabulary:

sat → 4.2
runs → 2.1
banana → -1.7
...

These raw output scores are called logits. A logit is not yet a probability. It is one score the model gives to one possible output token before a later conversion turns the scores into a probability distribution.

The basic language-model task is therefore:

visible prefix
→ model
→ one logit for every vocabulary token
→ next-token choice

What “decoder-only” means here​

In Level 5 you built a Transformer that lets each token position gather information from allowed positions. A decoder-only language model uses that style of Transformer for repeated next-token prediction.

The model is causal: an earlier prediction may use the prefix that has already appeared, but it may not use future text that would reveal the answer.

For example, from the training text:

the cat sat down

we can create aligned prediction positions:

visible input: [the] [cat] [sat]
next targets: [cat] [sat] [down]

At the position that predicts sat, the model may use the and cat. It must not look at the later sat token as evidence for its own answer.

This rule matters because otherwise training would be like giving a student the answer key while measuring how well they can answer the question.

One sequence contains many training questions​

Take the token sequence:

[the, cat, sat, down]

During training, this single row supplies several questions at once:

prefix "the" → target "cat"
prefix "the cat" → target "sat"
prefix "the cat sat" → target "down"

The implementation can calculate logits for all positions in parallel because the full training row is already known. But each position must behave as if the future were unavailable. Causal masking enforces that information boundary inside attention.

This explains an apparent paradox: training is parallel across positions while generation is sequential across newly created tokens. During training, all target tokens already exist for scoring; during generation, the next token does not exist until the model chooses it.

A good mental test is to ask, “Could this prediction be reproduced at generation time with only the prefix available?” If the answer is no, the training path has leaked information.

Predict

The model is predicting the token after the prefix `the cat`. Which text is it allowed to use?

Read the shapes after the idea is clear​

You will often see a compact shape such as (B, T) or (B, T, V).

In this Level:

  • B = batch size: how many sequences are processed together;
  • T = sequence length: how many token positions are in each sequence;
  • V = vocabulary size: how many possible token choices receive output scores.

So:

input token IDs: (B, T)
logits: (B, T, V)

means that each sequence position gets V raw next-token scores.

For a tiny example with B=2, T=4, and V=10, the input contains 2 sequences of 4 token IDs, and the output contains 10 vocabulary scores at each of those 8 positions.

Inspect causal access in the Lab​

The notebook prints a small table. A 1 means that the row may use that column position; a 0 means the column is in the future and is blocked.

  1. Open the Lab and run the code cell with length = 5 unchanged.
  2. Read the first row. It should contain one 1 followed by four 0s because the first position can use only itself.
  3. Read the last row. It should contain five 1s because the last position has no later position to hide.
  4. Find:
length = 5
  1. Change only 5 to 4.
  2. Before running, predict that the table will become 4 × 4: the first row will still have one allowed position and the last row will have four.
  3. Run the cell again and verify that pattern. The final PASS message should still appear.
  4. Restore length = 5 before moving on.

Loading lab…

The final PASS message checks the access rule. It does not prove that an entire language model is correct; it only gives evidence that this small causal visibility table blocks future positions.

Two properties that must stay true​

As models become larger, it helps to write down properties that should remain true while other settings change. Such a property is often called an invariant.

For this lesson, keep two invariants:

  1. input token IDs have shape (B, T) and output logits have shape (B, T, V);
  2. changing future tokens must not change an earlier position's output, because that earlier position is not allowed to use the future.

The second property is often called prefix invariance: if two examples share the same prefix, an earlier prediction based only on that prefix should not change just because later text differs.

A suspiciously tiny training loss is therefore not automatically good news. First check whether the answer accidentally became visible through the mask or target construction.

Quick Check

1. Why does a causal next-token model block future tokens?
2. In the shape `(B,T,V)`, what does `V` mean?
3. How can training calculate several positions in parallel without breaking causality?

0 of 3 questions answered.

Explain it back​

Explain a decoder-only language model using these four ideas in order:

  1. a visible prefix;
  2. logits for possible next tokens;
  3. a causal rule that hides future answers;
  4. repetition of the next-token step to generate longer text.

That repeated next-token process is also called autoregressive generation. You do not need a separate mental model for the word: it means that each new output is added to the prefix before the next prediction.

Key Takeaways

  • A decoder-only language model repeatedly predicts the next token from the prefix available so far.
  • Logits are raw output scores, one for each vocabulary choice.
  • Causal masking prevents an earlier prediction from using future answers.
  • (B,T,V) means batch × token positions × vocabulary scores.
  • An invariant is a property that should remain true while other settings change.
  • Autoregressive generation means adding each chosen token to the prefix and predicting again.

Next Lesson

Next, in L6.2 — Next-Token Prediction, you will look more closely at how one row of vocabulary logits becomes a next-token training decision.

References

Lesson actions

Completion is stored locally on this device.

View progress