Skip to main content
L1.2

Datasets as Tables

Goal

By the end of this lesson, you can read a dataset as rows of examples and columns of variables, identify candidate features and a target, and notice data-quality problems before modeling.

Start by asking what one row means​

Consider this tiny student dataset:

study_hoursattendancepassed
270%No
588%Yes
482%Yes

If each row represents one student record, then the first row says that one example has 2 study hours, 70% attendance, and a No result for passed.

This gives us the basic table vocabulary:

  • a row usually represents one example;
  • a column represents one measured property;
  • a feature is a column used as model input;
  • the target or label is the answer we want to predict.

If our task is to predict passed, then study_hours and attendance are candidate features and passed is the target.

The word candidate matters. A column can exist in a table without being a sensible model input.

A numerically valid column can still be wrong​

Suppose attendance is stored as fractions such as 0.70, 0.88, and 0.82. Those mean 70%, 88%, and 82% if the column contract says 0.0 to 1.0.

Now imagine one row contains 82 because someone stored that student's attendance as a percent instead of a fraction. The computer can still read the number. The meaning is inconsistent.

Other quiet data problems include:

  • 0 meaning “missing” in one system and a real zero in another;
  • duplicate rows for the same person;
  • an age of -3 caused by a data-entry error;
  • a column collected after the prediction moment;
  • two columns that use different units for the same quantity.

This is why understanding the table comes before fitting a model.

Keep the answer out of the inputs​

If passed is the target, putting passed inside the feature matrix would let the model read the answer it is supposed to predict.

That can produce an impressive score for the wrong reason.

A useful question for every feature is:

Would I know this value at the moment I need to make the prediction?

If the answer is no, the column may create leakage even if it is present in the dataset.

Shape is a compact description of the table​

In scikit-learn, feature data is commonly stored in X and targets in y.

If there are six examples and two features, then:

  • X has shape (6, 2) — six rows, two feature columns;
  • y has shape (6,) — one target value for each example.

Shape is not just Python bookkeeping. It tells you whether the number of examples and features matches the task you think you built.

Inspect the browser Lab​

The Lab uses a six-row student table. attendance is stored as a fraction, so 0.70 means 70%.

  1. Click Run without editing anything.
  2. Find features:, X shape:, and y shape:. Confirm the starter reports two features, X shape: (6, 2), and y shape: (6,).
  3. Read the first row in rows and translate it into ordinary language: 2.0 study hours, 0.70 attendance (70%), and passed = 0.
  4. After the final existing row, add exactly this seventh row:
{"study_hours": 5.5, "attendance": 0.84, "passed": 1},
  1. Before running, predict the shapes. There is still the same pair of features, but now there are seven examples.
  2. Click Run. Confirm X shape: (7, 2) and y shape: (7,).
  3. Remove the added row to restore the original six-row starter.

Loading lab…

After the guided pass, add a different valid row of your own. Keep the same three fields and explain why only the number of examples—not the feature count—changes.

A useful data dictionary​

Before modeling a real table, write down at least:

ColumnMeaningUnit/categoryAvailable when?
study_hourstime spent studying before the examhoursbefore prediction
attendanceattendance before the examfraction 0.0–1.0before prediction
passedexam outcome0/1after the exam

This simple habit catches many mistakes that code alone will not catch.

Quick Check

1. What does one row usually represent?
2. What is a feature?
3. Why write down units and meanings?

0 of 3 questions answered.

Key Takeaways

  • Rows are examples and columns are variables.
  • Features are model inputs; the target is the answer to predict.
  • Shape summarizes examples by features.
  • Data meaning, units, timing, missingness, and duplication matter before modeling.
  • A target or future-only field must not quietly become an input feature.

Next Lesson

Next, you will split examples into different roles so the final evaluation does not secretly help the model learn.

References

Lesson actions

Completion is stored locally on this device.

View progress