Skip to main content
L7.1

How Modern LLMs Behave

Goal

Explain modern LLM behavior as conditional next-token prediction, distinguish model probabilities from sampled outputs, and describe what one fluent response does and does not prove.

You already built a tiny decoder-only language model. A modern LLM uses the same basic idea at much larger scale: given a token prefix, it produces logits for the next token, turns them into a probability distribution, selects a token, appends it, and repeats. That simple loop creates surprisingly rich behavior because the model has learned patterns from a very large amount of text and has many parameters for representing those patterns.

One prompt can support several plausible continuations​

Suppose the current prefix is:

The meeting starts at

A model might assign next-token probability mass to several continuations:

9 .40
10 .25
noon .15
eight .10
other .10

Those probabilities are not a database lookup saying one token is certainly true. They are scores for plausible continuation under the model's learned distribution and the current context. If sampling is enabled, two runs can choose different next tokens even when the prompt and checkpoint are identical.

That does not mean the model “changed its mind” in the human sense. The probability distribution may be the same while the random selection differs.

Training behavior and product behavior are separated by a workflow​

A deployed model is rarely used as “raw logits only.” A system often adds:

  • a task instruction;
  • role/message structure;
  • retrieved documents or other context;
  • output-format requirements;
  • decoding settings;
  • validation and retry logic;
  • evaluation cases;
  • security boundaries.

The model is one component inside that larger workflow.

This matters because many failures are not fixed by changing the model itself. A missing source citation may require better context or validation. A malformed JSON object may require schema checks. A prompt-injection failure may require a stronger trust boundary.

Fluent text is weak evidence by itself​

A model can produce a confident sentence that is wrong. It can also produce a correct sentence for the wrong reason, or because a random sample happened to land on a useful continuation. So separate at least three questions:

  1. fluency: does the text read naturally?
  2. task performance: does it satisfy the requested behavior?
  3. grounding/reliability: is the answer supported by evidence and repeatable enough for the use case?

A single attractive sample mainly answers the first question.

The model has learned patterns, not a live world state​

Unless a workflow supplies fresh information through context or tools, the model's output comes from its learned parameters and the prompt. It does not automatically know that a fact changed yesterday. That is why later Levels introduce retrieval and tools. They give the workflow a way to supply current or private information instead of asking the model to invent it from memory.

Predict

A fixed prompt and checkpoint produce two different sampled continuations. What is the best first explanation?

Run the Lab​

The Lab uses a tiny fixed next-token distribution and seeded sampling.

  1. Click Run unchanged.
  2. Compare the lines. greedy: 9 always picks the most likely candidate. seed 7: and seed 7 again: print the same six sampled tokens, ['9', '9', 'noon', '9', '10', '9'], because they use the same seed. seed 11: prints a different sequence.
  3. Confirm that probabilities: is the same on every line: 9 0.4, 10 0.25, noon 0.2, eight 0.15. Different samples came from one fixed distribution.
  4. Now change only the distribution: probabilities = [0.40, 0.25, 0.20, 0.15] becomes probabilities = [0.10, 0.60, 0.20, 0.10].
  5. Before running, predict the new greedy choice and whether seed 7 will produce more 10s.
  6. Click Run. greedy becomes 10, and seed 7 becomes mostly 10. Changing the seed changes which draws you get; changing the probabilities changes what the model believes. Those are different experiments.
  7. Press Reset afterward.

Loading lab…

Quick Check

1. What does one fluent generated answer prove?
2. Which part of an LLM workflow can often change without retraining the model?
3. Why might a deployed system need retrieval or tools?

0 of 3 questions answered.

Explain it back​

Explain the difference between model probabilities, sampling, and workflow reliability using one example where two outputs differ but the model weights are unchanged.

Key Takeaways

  • Modern LLMs still generate through repeated next-token prediction.
  • Sampling can produce different continuations from one fixed probability distribution.
  • Fluency is not proof of factual correctness or reliability.
  • Prompting, context, decoding, validation, and security belong to the surrounding workflow.
  • Fresh or private information must be supplied explicitly through context or tools.

Next Lesson

Next, treat prompting as interface design: define what the model receives, what it must return, and how success can be checked.

References

Lesson actions

Completion is stored locally on this device.

View progress