Sampling Text
Goal
Explain generation as a repeated next-token selection loop and use a fixed probability distribution plus seed to distinguish stochastic sampling from model training.
A language model produces one next-token distribution at a time
Training changes model parameters. Sampling does not. Sampling starts after the model has produced scores for the next token.
A decoder generates text by repeating this loop:
current context
↓ model
next-token distribution
↓ choose one token
append that token
↓
new context
↺
For example:
"the moon"
↓ model
P(next token)
↓ choose "is"
"the moon is"
↓ model
P(next token)
↓ choose "quiet"
"the moon is quiet"
The newly selected token becomes part of the next input. That is what autoregressive generation means here: each new choice is conditioned on the prompt plus the tokens already generated.
Predict
Make randomness reproducible before changing policy
The Lab deliberately uses a tiny fixed distribution:
A: 0.6
B: 0.3
C: 0.1
seed: 7
This is smaller than a real vocabulary so you can see the sampling idea without mixing it with temperature, top-k, or top-p yet.
- Click Run unchanged.
- Read the 12 sampled tokens:
A A B A A A A A A A A A. MostlyA, as the weights suggest, but not onlyA. - Click Run again without changing the weights or the seed.
- Confirm that the 12-token sample repeats exactly.
- Change only
rng = random.Random(7)torng = random.Random(8), then run again. - The draws become
A C A B A A C A B A A A: a different sequence, even though the token weights stay[0.6, 0.3, 0.1]. - Press Reset afterward.
Loading lab…
The evidence separates two things:
- the distribution says which outcomes are more or less likely;
- the seed makes a particular sequence of random draws reproducible for an experiment.
A seed does not make A more probable than B. The weights already define that relationship.
Sampling is not greedy decoding
If you always choose the token with the largest score, that is greedy decoding. It is deterministic for fixed logits.
Sampling instead draws from a probability distribution, so several continuations can be possible from the same model state.
Neither method makes the underlying model better trained. They are different ways to choose from what the model already predicted.
Debug the right layer
If generated text looks bad, first separate these questions:
- Model question: Did training and validation produce a useful next-token distribution?
- Sampling question: Given that distribution, are we selecting tokens in the way we intended?
A decoding change cannot repair a model that never learned useful patterns. Conversely, a good checkpoint can still produce different text when the random draws or later decoding controls change.
Quick Check
Key Takeaways
- Generation repeats next-token prediction, selection, append, and predict again.
- Sampling can choose lower-probability tokens; it is not the same as greedy argmax.
- A fixed seed makes stochastic experiments reproducible under fixed conditions.
- Sampling changes the selected continuation, not the trained model weights.
- Diagnose model quality separately from decoding behavior.
Next Lesson
Next, L6.10 — Temperature, Top-k, and Top-p keeps the sampling loop fixed and changes the decoding policy: which next-token choices remain likely or eligible.
References
- Python,
random.choices. - PyTorch,
torch.multinomial.
Completion is stored locally on this device.