Decoding and Determinism
Goal
Separate deterministic selection from sampled generation, explain what a seed can and cannot guarantee, and design repeatable comparisons that record all decoding conditions.
Level 6 introduced temperature, top-k, and top-p. Now treat those settings as part of a workflow contract. If a prompt evaluation cannot be repeated under known conditions, it becomes hard to tell whether a change improved the prompt or merely changed the random sample.
Greedy selection and sampling answer different needs
If the workflow always selects the highest-probability token, each step is greedy. For a fixed model implementation and identical numerical inputs, greedy decoding aims for repeatable token selection. Sampling instead draws from the candidate distribution. That can be useful when diversity matters, but it adds variance to evaluation.
Neither choice is universally correct. A data-extraction workflow may prefer highly repeatable behavior. A creative brainstorming tool may deliberately use sampling.
A seed controls one source of randomness
A pseudo-random generator can produce a repeatable sequence when initialized with the same seed. That gives experiments a useful control:
checkpoint fixed
prompt fixed
decoding settings fixed
seed fixed
But do not turn this into the stronger claim “same seed always guarantees identical output everywhere.” Differences in runtime, hardware kernels, model versions, numerical precision, batching, provider implementation, or hidden server-side changes can affect reproducibility. Record the environment and model identity along with the seed.
Determinism is not correctness
A model can return the same wrong answer every time. Reproducibility makes a behavior easier to measure and debug; it does not make the behavior good. Likewise, sampling variation is not automatically a defect. If the application permits multiple valid phrasings, different outputs may all satisfy the task.
Define success separately from sameness.
Evaluate stochastic behavior with repeated trials
Suppose a prompt passes 9 out of 10 sampled trials. One run would not reveal that reliability rate. For a probabilistic workflow, you may need repeated samples per case, with metrics such as:
- format-pass rate;
- unsupported-claim rate;
- task success rate;
- number of distinct valid outputs.
The number of trials should match the risk and cost of the application. A classroom Lab can use a few seeds. A high-stakes production system needs stronger evaluation design.
Reproducibility needs more than a seed
A random seed controls one source of variation, but it does not freeze the entire inference system. A provider can change model weights, tokenizer behavior, numerical kernels, or serving code; a local runtime can change library versions or hardware math. If an evaluation depends on exact output identity, record the model and runtime identity alongside decoding settings and seed.
Also separate repeatability from correctness. Greedy decoding can return the same unsupported answer every time, while stochastic decoding can sometimes produce a correct one. Reliability work therefore records both the generation configuration and an independent evaluation of the result. The seed helps reproduce an experiment; it is not evidence that the generated content is true.
Predict
Run the Lab
The Lab uses fixed probabilities and a seeded pseudo-random generator.
- Click Run. Compare
seed 7:andseed 7 again:. Both are['A', 'A', 'B', 'A', 'A', 'A', 'A', 'A']. - Compare
seed 11:,['A', 'A', 'C', 'A', 'A', 'A', 'A', 'A']. A different seed gives a different sequence from the same weights. - Read
greedy: A. Greedy selection always picks the highest weight, so it does not use the random generator at all. - Now change the distribution while keeping every seed:
weights = [0.60, 0.30, 0.10]becomesweights = [0.20, 0.30, 0.50]. - Click Run.
greedybecomesC, andseed 7now gives['B', 'A', 'C', 'A', 'C', 'B', 'A', 'C']. The seed is the same, but the outputs are completely different, because the thing being sampled changed. - Explain why “same seed” is not a fair comparison once the distribution has changed. Press Reset afterward.
Loading lab…
Quick Check
Explain it back
Describe an evaluation where greedy decoding is appropriate and another where sampling is appropriate. For the sampled case, list the conditions you would record to make the experiment reproducible.
Key Takeaways
- Greedy selection and sampling serve different workflow goals.
- A seed controls sampling randomness but is not a universal cross-runtime guarantee.
- Deterministic output can still be wrong.
- Probabilistic systems may require repeated trials per evaluation case.
- Record checkpoint, prompt, decoding configuration, seed policy, and environment together.
Next Lesson
Next, distinguish fluent generation from grounded knowledge and design workflows that can abstain when evidence is missing.
References
- Radford et al., Language Models are Unsupervised Multitask Learners.
Completion is stored locally on this device.