본문으로 건너뛰기
L0.12

Debug and Improve a Tiny Predictor

Goal

By the end of this lesson, you can use a repeatable debugging loop to inspect one failure, form a hypothesis, change one thing, compare the result fairly, and decide whether the evidence supports the change.

Debugging is not random tweaking​

When a prediction is wrong, changing numbers until the score improves can work by luck. It does not tell you why the change helped.

Use this loop instead:

  1. Find one failure.
  2. Describe the evidence.
  3. Propose one possible cause.
  4. Make one change that tests that cause.
  5. Evaluate under the same conditions.
  6. Compare with the original setting and a simple baseline.
  7. Explain what the result supports.

Start from one visible failure​

Training examples:

InputLabel
1False
2False
3True
4True
5True

Held-out examples:

InputLabel
2False
3True
4True
6True

The initial predictor uses threshold 4:

predict True when input >= 4

On the held-out set, input 3 is wrong: the predictor says False, but the label is True.

So the initial predictor makes 1 mistake.

A reasonable hypothesis is:

The threshold may be one step too high.

That suggests one exact change: threshold 4 → 3.

With threshold 3, inputs 2, 3, 4, and 6 are all predicted correctly, so we expect the mistake count to fall from 1 to 0.

Compare with a baseline​

The training labels contain more True values than False values. A majority baseline therefore predicts True for every input.

On the held-out set, that baseline misses input 2, so it also makes 1 mistake.

Now we have context:

  • initial threshold 4: 1 mistake;
  • candidate threshold 3: expected 0 mistakes;
  • majority baseline: 1 mistake.

Run the debugging record​

  1. Click Run without changing anything.
  2. Read the printed debug_record dictionary.
  3. Find failure, hypothesis, change, before, after, and baseline.
  4. Confirm that the starter reports before: 1, after: 0, and baseline: 1.

Loading lab…

The evidence supports a careful statement:

On this held-out set, lowering the threshold from 4 to 3 removed the observed mistake and beat the majority baseline.

It does not prove that threshold 3 is best for every future example.

Test a hypothesis that does not help​

Now make one exact edit in the Lab:

candidate_threshold = 3

Change only 3 to 2.

Before clicking Run, predict what will happen to held-out input 2.

With threshold 2, input 2 becomes True, but its label is False. The candidate therefore makes 1 mistake again.

Run the Lab and compare before, after, and baseline.

This result is useful because it rejects the broad idea that any lower threshold must improve the model.

A useful record is:

  • Failure: threshold 4 misses input 3.
  • Hypothesis: a lower threshold may fix the error.
  • Change: threshold 4 → 2.
  • Result: mistake count stays at 1 because input 2 becomes wrong.
  • Next step: test a more specific candidate or ask whether one threshold is expressive enough.

Debug the right layer​

Not every wrong prediction is a model-setting problem. The cause may be:

  • input: a feature value is missing or incorrect;
  • label: the expected answer was recorded incorrectly;
  • feature: important information is absent;
  • model: the predictor is too simple;
  • evaluation: the metric or test cases do not match the real goal;
  • software: the code does not implement the intended rule.

Ask which layer the evidence points toward before changing the model.

Protect final evaluation evidence​

If you repeatedly tune a threshold after reading the same held-out results, those examples start guiding model selection. They are no longer a clean final test.

For this tiny teaching exercise we reuse the held-out examples so you can see the effect clearly. In a real workflow, repeated model selection normally uses training/validation evidence, while a final test stays protected until the end.

Evaluation evidence should not secretly become training information.

A reusable debugging record​

FieldWhat to write
FailureWhich example or behavior is wrong?
EvidenceWhat exactly did you observe?
HypothesisWhat possible cause are you testing?
ChangeWhat one important factor will change?
FixedWhat stays the same?
BeforeOriginal result
AfterResult after the change
BaselineSimple reference result
ConclusionWhat does the evidence support?
Next stepWhat would you test next?

This pattern will stay useful even when later models become much larger.

Before the Level Project: choose your workspace​

The next activity is the Level 0 Project, Data Detective. Local Python is not a hidden Level 0 prerequisite, so you may choose either path:

Core browser path​

Stay in the browser. Use the evidence you just produced and complete the Data Detective Core Evidence Report. This path demonstrates the Level 0 exit skills without requiring repository or terminal setup.

Builder local path​

Implement the same reasoning in the provided Python starter and run the validator. If words such as repository root, terminal, or validator are new, read the Project Workbench first. It teaches that environment explicitly before later projects depend on it.

Both paths use the same rubric. The difference is the implementation environment, not the conceptual standard.

Quick Check

1. What should normally come before changing a model setting during debugging?
2. Why evaluate before and after under the same conditions?
3. A candidate change does not improve the result. What should you do?

0 of 3 questions answered.

Key Takeaways

  • Debugging starts from a specific failure and evidence.
  • A hypothesis should suggest a focused change.
  • Keep comparisons fair by holding evaluation conditions fixed.
  • Use a baseline so “better” has context.
  • Failed hypotheses are useful when recorded.
  • Do not assume every error is a model-setting problem.
  • Local coding is a Builder skill, not a hidden requirement for Level 0 conceptual mastery.

Next Lesson

Complete Data Detective using either the Core browser path or the Builder local path. After the project, Level 1 introduces more formal machine-learning models and metrics.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Data Detective: Build and Explain a Tiny Predictor

View progress