Step 4

Getting the file into a usable shape

Nine of the columns you are allowed to use hold words, and every model in this course refuses them. You turn all nine into numbers, hand the model a table with 47 columns instead of 7, and then find out how hard it is to tell whether that helped.

About 4 hours 50 minutes13 partsBest over two sittingsThe plumbing step
Read this first3 min

What this step gives you

First, the thing Step 3 asked you to bring

Step 3 sent you to data/bank-names.txt, to the heading related with the last contact of the current campaign, and asked what you made of the columns underneath it. Read your lines back now. If you cannot find them, the next paragraph gives you enough to carry on.

Four columns sit under that heading: contact, day, month and duration. Only day has been in your seven, so if that is all you found, you read the file properly. The second column worth arguing about is campaign, which is filed under other attributes and gives itself away in its own description rather than by its heading.

Here is what I decided, so you can argue with it. day is the day of the month the call happened, and if you are choosing who to ring on the fourteenth then the day is something you know, so it stays. campaign counts the calls made to this person in this campaign including the one being made, so on Monday morning the honest value is one less than the file says. It is off by one, always in the same direction, and a tree cutting at 2.5 does not care much. It stays, with that written down beside it.

Neither is like duration. Both are worth a sentence in your README, because the person who reads your work after you will ask.

Every model in this course from here to the end takes a table of numbers and nothing else. Not words. Not blanks. Not a column where somebody typed -1 to mean "never happened".

A few models outside this course will take words directly, if you set them up to. They are the exception, they are not on this ladder, and they still need everything else in this step.

This step is where you learn to hand a model a table it will accept. It is the least glamorous step in the ladder, and it is where most real mistakes get made.

The two things that happen in Parts 5 and 6

Your seven-column table has been scoring 0.893 on people it has never seen since Step 3, catching 22 of the 138 who would say yes.

In Part 5 you hand the same model 40 more columns, every job title and every month and everything else the file knows. The score is 0.893. The catches are 22. The wasted calls are 5. Not close to Step 3. The same, to the last digit.

In Part 6 you let it ask one more question and the score goes to 0.8983, catching 28 for the same 5 wasted calls. Six more customers for nothing, and the tree gets there by asking about October, which nobody told it about and which you found by hand in Step 1.

Then, in the same part, you do what Step 3 taught you and check whether that win is real. Run the same comparison over 30 different splits and the wider table wins 11 times out of 30, and loses on average. The split you were handed was the sixth best of the thirty.

That is the step. Not a win, and not a failure: a result you cannot read off one number, and the first honest look at how small a difference has to be before you stop being able to see it.

What you actually do, in order

  1. Give five job titles a number each, on paper, and find the lie you just told.
  2. Count how many different answers each word column holds.
  3. Turn the three yes-or-no columns into 0 and 1.
  4. Turn the six others into 38 tick boxes.
  5. Train on all 47 columns and watch nothing whatsoever happen.
  6. Let it ask one more question, watch the score rise, then find out how much to believe it.
  7. Send twenty new people through it and watch the whole thing break.
  8. Meet the -1 from Step 0 again, and decide what to do with it.
  9. Scale the columns, and find out what your test rows told your scaler.
How the time actually goes

Parts 0 and 1 are twenty minutes on paper, away from the screen. You do not open the notebook until Part 2. If you have less than an hour tonight, do Parts 0 and 1 and stop there. They stand on their own.

Part 5 ends on a shock and Part 5's explanation is the cure, so it wants twenty minutes in one go. If you have less than that left, stop at the end of Part 4.

The good places to stop are the end of Part 2, the end of Part 4, the end of Part 6 and the end of Part 9. Part 10 is a sitting of its own. In three-quarter-hour evenings this step is five or six of them, and that is normal for the longest step in the course.

Part 0Unplug15 minAway from the screen

Give five jobs a number

Paper and pen. Read the table below, copy the five job titles down the left of your page, then shut the laptop.

Here are five of the twelve job titles in the file.

JobYour number
student
blue-collar
management
housemaid
retired

The model needs numbers. So give each job a number, 1 to 5, in whatever order feels natural, writing them beside the titles on your paper. It takes ten seconds and there is no wrong answer.

Now read back what you just wrote

Whatever order you chose, you have told the machine four things you did not mean.

  1. That these five things have an order at all.
  2. That the gap between number 1 and number 2 is the same size as the gap between 4 and 5.
  3. That number 3 sits between number 2 and number 4, in some real sense.
  4. That number 4 is, in some way, twice number 2.

A tree hears the first one and only the first one. It asks whether a number is below a cut, so with your numbering it can ask "is this job below 3.5", which lumps your first three jobs together and splits off the other two. That grouping is an accident of the order you happened to write them in. Somebody else numbering the same five jobs gets a different model, and neither of you meant to say anything about order at all.

The other three lies are wasted on a tree, which never adds or averages anything. They are not wasted on the models in Steps 9 and 10, which do both. Numbering a word column is the kind of mistake that lies quiet under one model and goes off under the next.

Do this one, it is the whole part

Rub out your numbers. Draw five columns instead, one for each job, and write the five job titles across the top.

Now take three people: a student, a housemaid, and a manager. Give each of them a row, and put a 1 in the column that matches their job and a 0 in the other four.

You have three rows of five numbers. No order. No gaps. Nothing between anything.

That is one-hot encoding, and you have just done all of it. One column per possible answer, one 1 per row, zeroes everywhere else. The name is silly and comes from electronics. The idea is a row of tick boxes.

What it costs

Count the columns you just used. Five jobs became five columns. The real file has twelve job titles, so twelve columns. Six word columns with several answers each will become 38.

That is fine here. It is not always fine. Write down what you think happens to a column holding five thousand different customer names, and keep it. Step 18 is where that problem gets solved properly.

One question to answer before you open the laptop

Three of the nine word columns hold only yes and no. Do those need five columns, or two, or one?

Write your answer and why. Part 3 is that question.

Part 1Words5 min

Words to know

Six words. All six turn up in job adverts for this work.

Encoding
Turning words into numbers so a model will take them. The whole of Parts 3 and 4.
One-hot
The tick-box way of doing it. One column per answer, a 1 in the one that applies.
Imputing
Filling a gap with something. The something is a decision, and it is yours.
Scaling
Putting every column on comparable footing, so that a column measured in thousands cannot shout down a column measured in ones. There are several kinds. This step uses the simplest: squeeze everything onto 0 to 1.
Fit
Look at the data and work out what to remember. A min-max scaler fits by finding the smallest and largest value; other scalers remember other things.
Transform
Apply what was remembered. Fit once, on the training rows. Transform everything, including rows that arrive next year.
The one sentence to keep

Preparing your data is not something that happens before the model. It is part of the model, and everything Step 3 said about the split applies to it.

Part 2Run15 min

Nine columns of words

Open band-04.ipynb. Every cell on this page is already in it, in this order. You run them; you do not type them. The only typing you do in this step is in Part 10, on your own table.

If you are coming back on a new evening

Run every cell from the top before you start. Nothing is remembered overnight, and the cells later in this step need names that were made earlier in it.

If you see NameError: name 'ready' is not defined, that is all this is. Nothing is broken.

If your numbers are different from mine

Check that every random_state=0 on the page is in your cell too. Without it the rows are shuffled differently every run, so your numbers change every time you press Shift and Enter and no two people in the world can compare notes.

If the numbers still differ, run every cell from the top in a fresh kernel: Kernel, then Restart Kernel and Run All Cells.

Step 2 ended on a refusal. You handed the tree a column of job titles and got ValueError: could not convert string to float: 'unemployed'. Since then you have worked with seven columns: the six that were already numbers, and said_yes_before, which you built by hand in Step 2 because you needed one word column badly enough to do it yourself.

Here is what you have been leaving out.

import pandas as pd
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split

df = pd.read_csv("data/bank.csv", sep=";")
y = (df["y"] == "yes").astype(int)

words = ["job", "marital", "education", "default",
         "housing", "loan", "contact", "month", "poutcome"]

for name in words:
    print(name, df[name].nunique())

.nunique() counts how many different answers a column holds. The loop prints the name and the count for each of the nine, one line each. You met value_counts() in Step 0, which shows every answer and how often it appears; this is the short version, when all you want is how many kinds there are.

You should see
job 12
marital 3
education 4
default 2
housing 2
loan 2
contact 3
month 12
poutcome 4

Nine columns. Three of them hold two answers. Six hold more. That split is the plan for the next two parts, and it is the answer to the question you wrote at the end of Part 0.

If you get FileNotFoundError

Jupyter is not standing in the course folder. Close it, open a terminal, type cd and a space and the path to the folder that holds data, press Enter, then type jupyter lab and press Enter. Parts 6 and 7 of the setup page walk through it slowly.

Part 3Run15 min

Two answers become 0 and 1

A column holding only yes and no does not need tick boxes. Two boxes would be one column of information written twice, because the second is always whatever the first is not.

One column is enough, and you have done this before. In Step 2 you wrote (df["poutcome"] == "success").astype(int) to make said_yes_before. Same move, three times.

ready = df[["age", "balance", "day", "campaign", "pdays", "previous"]].copy()

for name in ["default", "housing", "loan"]:
    ready[name] = (df[name] == "yes").astype(int)

print(ready.shape)
print(ready["housing"].sum())

The loop does the same thing to three columns rather than copying the line out three times. ready starts as the six columns that were already numbers, and gains three more.

Note what is not in ready: said_yes_before is gone. You built that by hand in Step 2 because you needed one word column badly enough to do it manually. In Part 4 the machine builds it for you, along with 37 others, and it will be called poutcome_success.

You should see

(4521, 9), then 2559

Nine columns, and 2,559 of the 4,521 people have a housing loan. The notes say housing: has housing loan? and nothing more, so a housing loan is what you call it, not a mortgage.

Why 1 for yes and 0 for no, and not the other way round

For a tree it makes no difference at all. It will find the same cut either way, because the two groups are the same two groups.

It makes a difference to you at two in the morning. Pick the convention that 1 means the thing the column is named after, and never break it. housing is 1 for a person with a housing loan. If half your columns say yes with a 1 and the other half say yes with a 0, every sentence you write about the model will be wrong half the time.

Part 4Run20 min

Many answers become boxes

Now the six columns with more than two answers. This is Part 0, done by machine.

boxes = pd.get_dummies(df[["job", "marital", "education",
                           "contact", "month", "poutcome"]], dtype=int)

print(boxes.shape)
print(list(boxes.columns[:4]))

pd.get_dummies does exactly what you did on paper. It reads a column, finds every different answer in it, and makes one new column per answer, holding 1 for the rows that match and 0 for the rest.

The [:4] on the last line means "the first four". A colon inside square brackets asks for a run of things rather than one thing, and with nothing before the colon it starts at the beginning. There are 38 names and four is enough to see the pattern.

dtype=int asks for 1 and 0 rather than True and False. Both work. Ones and zeroes are easier to read when you print the table, and you are going to print it.

The word "dummies" is a statistics term for these columns and it means nothing useful. Read it as "tick boxes" every time you see it.

You should see

(4521, 38), then ['job_admin.', 'job_blue-collar', 'job_entrepreneur', 'job_housemaid']

Six columns became 38, named for the column they came from and the answer they stand for.

Check that it did what you think it did, rather than trusting the shape.

print(boxes["job_student"].sum())
print((df["job"] == "student").sum())

Adding up a column of 1s and 0s counts the 1s. The second line counts the students the Step 1 way, by comparing a column to a value and adding up the Trues. That shape, a column of True and False you use to pick people out, is called a mask. Step 1 had you write them all evening without giving them a name; the name is worth having now, because Part 6 stacks three of them.

The two numbers have to match, or something is wrong.

You should see

84, then 84

The habit, not the number

Nobody made you run that second cell. Get used to running it anyway.

Every time you reshape a table you should check one thing you can count two ways. Nothing about this step turns red when it goes wrong. It just quietly gives you a different table from the one you think you have.

Join the two halves together.

ready = pd.concat([ready, boxes], axis=1)

print(ready.shape)

pd.concat you met in Step 3, where you stacked the file on top of itself. That used the default, axis=0, which means downwards, adding rows. axis=1 means sideways, adding columns. Same tool, turned ninety degrees.

You should see

(4521, 47)

Six numbers, three yes-or-no columns, 38 boxes. Every one of them a number, and not a word left anywhere.

Part 5PredictRun20 min

Forty more columns, and nothing happens

Predict first. Write one number.

Since Step 3 your best honest score has been 0.893, catching 22 of the 138 and wasting 5 calls. That came from seven columns.

The model is about to get 47. Every job title, every month, every education level, everything the file knows and you have been throwing away.

Same depth 2 as Step 3, so that the only thing that changes is the width of the table. What does it score? Write the number down before you run it.

X_train, X_test, y_train, y_test = train_test_split(
    ready, y, test_size=0.25, random_state=0)

flat = DecisionTreeClassifier(max_depth=2, random_state=0)
flat.fit(X_train, y_train)

print(round(flat.score(X_test, y_test), 4))

flat_guess = flat.predict(X_test)
print(((flat_guess == 1) & (y_test == 1)).sum())
print(((flat_guess == 1) & (y_test == 0)).sum())

The same split as Step 3, with the same random_state=0, so the same people are held back. The same depth. The only change in the world is that the table is 40 columns wider.

You should see

0.893, then 22, then 5

Go back and look at Step 3, Part 4. Those are the same three numbers. Not similar. The same.

Sit with this before you read the explanation

You did four parts of work. You turned nine columns of words into 41 columns of numbers, correctly, and checked them. The model was handed all of it.

It changed nothing at all. Write down what you think happened, in one line, before you read on.

Here is the part that surprises people, and it is not what you probably wrote down. Print the tree and one of your new columns is in it: with 47 columns to choose from, the second question is month_oct, where the seven-column tree asked about age. Your work did win a place.

It bought nothing, because both branches under that question still say no. The tree is different and the answers are not, which is the same thing Part 4 of Step 2 showed you when a split changed nothing.

A depth 2 tree gets two questions and stops. Your new columns are not useless. They are just not strong enough to flip an answer inside two questions, and a tree with two questions never gets to the third.

One number for scale, since this is the step about not trusting one split: run the two tables against each other over thirty splits and they give identical predictions on sixteen of them. On the other fourteen they differ by a handful of people, in both directions.

The habit this step exists to give you

This is worth more to you than a win would have been. Most of the work in a real project is like this: correct, necessary, and worth nothing on its own.

The columns you just built pay off in Part 6, in Step 5, and in every model from Step 8 on, all of which can hold more than two ideas at once. Work that pays later still has to be done right now, and nobody claps.

Part 6ModifyInvestigate35 min

One more question, and how much to believe it

Give it a third question and see whether the new columns get used.

tree = DecisionTreeClassifier(max_depth=3, random_state=0)
tree.fit(X_train, y_train)

print(round(tree.score(X_test, y_test), 4))

guess = tree.predict(X_test)
print(((guess == 1) & (y_test == 1)).sum())
print(((guess == 1) & (y_test == 0)).sum())
You should see

0.8983, then 28, then 5

0.893 became 0.8983. Twenty-two caught became 28. Wasted calls stayed at 5.

Six more customers for nothing. Write it in your README, and then keep reading, because this part is not over.

First, find out what it did with them

print(export_text(tree, feature_names=list(ready.columns)))
You should see
|--- poutcome_success <= 0.50
|   |--- month_oct <= 0.50
|   |   |--- age <= 60.50
|   |   |   |--- class: 0
|   |   |--- age >  60.50
|   |   |   |--- class: 0
|   |--- month_oct >  0.50
|   |   |--- day <= 16.50
|   |   |   |--- class: 0
|   |   |--- day >  16.50
|   |   |   |--- class: 1
|--- poutcome_success >  0.50
|   |--- balance <= 8053.00
|   |   |--- education_tertiary <= 0.50
|   |   |   |--- class: 1
|   |   |--- education_tertiary >  0.50
|   |   |   |--- class: 1
|   |--- balance >  8053.00
|   |   |--- class: 0
Read the second line again

month_oct.

In Step 1 you sat with a table of twelve months and worked out by hand that March, September, October and December were different from the rest. You wrote a wider rule around them and it scored 0.8821, which was worse than saying no to everybody, and you kept the line anyway because the months themselves were a real finding badly used.

Nobody told this tree anything about months. It got twelve columns named after month names, in alphabetical order, meaning nothing. It went and found one of your four.

It found something you did not give it, which is the same thing that happened in Step 2 when it found your Step 1 rule. A model is not cleverer than you. It is faster, and it never gets bored, and if the thing is there it will find it.

Read the rule properly before you believe your own summary of it. It is not "October". It is October, after the sixteenth. That is one branch of one tree, and you are about to find out how much weight it will take.

Two questions that change the answer less than they look

Look at age <= 60.50. Both branches under it end in class: 0. Look at education_tertiary. Both branches end in class: 1.

The decision is the same on both sides, so as far as the yes-or-no answer goes, those two questions could be deleted. They are not decoration though, and it matters that you know why: the tree picked them because they were the best splits it had left, and each one changes how sure the model is. On the university branch it moves from about half the people saying yes to about four in five.

Today you are reading a yes or a no, so you cannot see that. In Step 9 you start reading the chance instead, and those two lines come back to life.

Now check it the way Step 3 taught you

Step 3's Silver task asked you to run the same split with five different random_state values and look at the spread. Here is that task, made compulsory, on a result you have a reason to want to be true.

Predict first, and be honest

You are about to run the same comparison over 30 different splits: the wide table with three questions against the Step 3 table with two.

How many of the 30 will the wide table win? Write the number down. Nobody will see it but you.

seven = df[["age", "balance", "day", "campaign", "pdays", "previous"]].copy()
seven["said_yes_before"] = (df["poutcome"] == "success").astype(int)

wide_scores = []
narrow_scores = []
wins = 0

for seed in range(30):
    a, b, c, d = train_test_split(ready, y, test_size=0.25, random_state=seed)
    wide = DecisionTreeClassifier(max_depth=3, random_state=0)
    wide.fit(a, c)
    wide_score = wide.score(b, d)

    a, b, c, d = train_test_split(seven, y, test_size=0.25, random_state=seed)
    narrow = DecisionTreeClassifier(max_depth=2, random_state=0)
    narrow.fit(a, c)
    narrow_score = narrow.score(b, d)

    wide_scores.append(wide_score)
    narrow_scores.append(narrow_score)

    if wide_score > narrow_score:
        wins = wins + 1

print(round(min(wide_scores), 4), round(max(wide_scores), 4))
print(round(sum(wide_scores) / 30, 4))
print(round(sum(narrow_scores) / 30, 4))
print(wins)

Longer than usual, and every line of it is something you have done before. Four things are new and none is hard.

The wins counter starts at 0, and wins = wins + 1 means take what is in wins, add one, and put it back. It only happens on the passes where the wide table came out ahead.

You should see

0.8753 0.9045, then 0.8902, then 0.8915, then 11

Read those four numbers in order. They are the point of this step.

The wide table scores anywhere from 0.8753 to 0.9045 depending on nothing but which quarter of the people you happened to hold back. That is a spread of about three points, and the win you were celebrating was half a point.

On average the wide table gets 0.8902 and the Step 3 table gets 0.8915. The wide one is slightly worse.

It won 11 of the 30. If it were a coin you would expect 15.

The split you were handed, random_state=0, was the sixth best of the thirty for the wide table. You did not choose it. It was chosen in Step 3, for other reasons, long before any of this. And it happened to be the one where the new columns look good.

What this comparison does not separate

Two things differ between the two models: the wide one has 47 columns and three questions, the narrow one has 7 and two. So this tells you about the pair as a whole, not about columns or depth on their own.

You already know one half of it, because Part 5 held depth still and the columns bought nothing. Separating the rest properly is Step 6's job.

What that does and does not mean

It does not mean the work was wasted. Every model from Step 8 onwards can use forty columns properly, and not one of them can use a column of words. You did the work that makes them possible.

It does not mean the extra columns are worthless. It means this measurement cannot tell, because the difference you are chasing is smaller than the noise in the way you are measuring.

What it means is that "0.893 became 0.8983" was never a sentence you were entitled to write. You would have written it. I did write it, in the first version of this page, and somebody who ran 30 splits took it out.

The habit this step exists to give you

Before you believe a difference between two numbers, find out how much each number moves on its own when nothing important has changed.

One split gives you one number and no idea how much it wobbles. Thirty splits cost you four seconds and tell you whether you are looking at a result or at weather.

Step 14 is where this becomes a proper tool with a name. You do not need the name to do it, and you have just done it.

And find the six people

One more thing before you leave it. The whole difference in Part 6 was six extra customers caught. Go and look at them.

october_leaf = ((X_test["poutcome_success"] == 0) &
                (X_test["month_oct"] == 1) &
                (X_test["day"] > 16.5))

print(october_leaf.sum())
print((october_leaf & (y_test == 1)).sum())

Three masks with & between them, which is the shape from Part 4 stacked three deep. Each set of brackets is one question from the tree, read straight off the printed rules above, and together they pick out exactly the people who land in the leaf that says yes about October.

The first line counts them. The second adds one more condition, that they really said yes, and counts those.

You should see

6, then 6

Six people in the test half reach that leaf. All six said yes.

Your entire improvement is six people, and every one of them went the right way. On the training side that leaf holds 34 people, 20 of whom said yes, which is a real signal and not nothing. But six out of six is the kind of luck that does not repeat, and now you know exactly where the half point came from.

Step 6 is where you learn to ask which columns are pulling their weight, properly, instead of squinting at one tree.

Part 7ModifyInvestigate25 min

Twenty people arrive on Monday

You have a model that works. Now use it, which is the point of having one.

Twenty people from the test half arrive as they would in real life: raw, in the shape the file has them, not the shape the model wants.

newcomers = df.loc[X_test.index[:20]]

new_boxes = pd.get_dummies(newcomers[["job", "marital", "education",
                                      "contact", "month", "poutcome"]], dtype=int)

print(new_boxes.shape)

df.loc[X_test.index[:20]] takes the first twenty row numbers from the test half and pulls those rows out of the original file. They are people the model has never seen, in their original words.

You should see

(20, 23)

Stop. Twenty-three.

The model was trained on 38 box columns. These twenty people produced 23.

Work out why before you read on. It is not a bug in pandas.

Twenty people do not have twelve different jobs between them. They do not cover all twelve months. get_dummies makes one column per answer it can see, and twenty people cannot show it everything 4,521 people could.

So the table has the wrong shape, and the columns that do exist are in the wrong order. Hand that to the model and see what it says.

print(len(new_boxes.columns), len(X_train.columns))

try:
    tree.predict(new_boxes)
except ValueError as error:
    print("ValueError:", error)

len(...) counts things. You met it in Step 1 as len(df), counting rows; len(new_boxes.columns) counts the names in the list of columns instead.

The try and except shape is the one from Step 2: attempt this, and if it goes wrong, catch the complaint and print it as ordinary text so the rest of the notebook still runs.

You should see

23 47, then a complaint several lines long that begins:

ValueError: The feature names should match those that were passed during fit.

and then lists the columns it expected and cannot find. Read the first line. The list underneath is the detail, and it is worth a glance because it names the plain columns like age too, not only the boxes.

This is one of the standard ways a model that worked in a notebook fails the first time somebody tries to use it, and I have watched it happen. Nothing about the model is wrong. The preparation was done twice, by two different people or by the same person on two different days, and the second one did not know what the first had decided.

The fix is a thing that remembers

get_dummies looks at whatever you hand it and decides the columns on the spot. You need something that decides once, on the training rows, writes the decision down, and applies the same decision forever after.

from sklearn.preprocessing import OneHotEncoder

many = ["job", "marital", "education", "contact", "month", "poutcome"]
plain = ["age", "balance", "day", "campaign", "pdays", "previous"]

raw_train = df.loc[X_train.index]

encoder = OneHotEncoder(sparse_output=False, handle_unknown="ignore")
encoder.set_output(transform="pandas")
encoder.fit(raw_train[many])

print(len(encoder.get_feature_names_out()))

df.loc[X_train.index] is the same move as the one you used on the newcomers: take the row numbers of the training half and pull those rows, in their original words, out of the file.

Four new things, and they are the shape of every tool in the rest of this course.

get_feature_names_out() is the list of column names it decided on, and len counts them.

Notice which rows it was fitted on: raw_train only. Part 9 is about why.

You should see

38

The same 38 columns, decided once and written down inside the encoder.

Now one function that puts any set of raw rows into the model's shape.

def prepare(raw):
    out = raw[plain].copy()
    for name in ["default", "housing", "loan"]:
        out[name] = (raw[name] == "yes").astype(int)
    return pd.concat([out, encoder.transform(raw[many])], axis=1)

print(prepare(newcomers).shape)

A def gives a block of work a name so you can use it again without copying it out. Everything indented under the first line is the work, and return hands the answer back to whoever asked.

raw is a stand-in. It means "whatever table somebody hands this thing", and inside the block that is what raw refers to. Below, prepare(newcomers) hands it the twenty newcomers, so on that run every raw in the block means newcomers. Tomorrow you hand it a different table and every raw means that one.

plain, many and encoder are not handed in. The block reaches out and uses them where they sit. That is fine and normal, and it is also the reason a function like this stops working if you rebuild the encoder later without rerunning it.

encoder.transform(...) is the other half of fit. Fit remembered; transform applies. Because of the set_output line it hands back a table with the 38 names already on it, which pd.concat then joins sideways to the nine you built by hand.

This is the first function in the course and it will not be the last. The alternative is doing these five lines in three places and getting one of them wrong.

You should see

(20, 47)

Twenty people, 47 columns, in the same order the model was trained on. Now it works.

print(tree.predict(prepare(newcomers)).sum())
You should see

0

It says no to all twenty. Roughly 12 people in 100 say yes, and this model only says yes to about 3 in 100, so twenty people producing no yeses at all is ordinary. It is not the error you just fixed coming back.

Write this in your README

One line: what get_dummies does that OneHotEncoder does not, and when you would still use get_dummies.

There is a real answer to the second half. It is a good tool for looking at a table. It is a bad tool for feeding a model you intend to use twice.

Part 8Investigate25 min

The number that is not a number

In Step 0 you found that pdays has an average of 39.77, and that the average was a lie, because 3,705 of the 4,521 rows hold -1, which does not mean minus one day. It means this person was never contacted before.

That column is sitting in your model right now. Time to deal with it.

print((df["pdays"] == -1).sum())
print(round(df["pdays"].mean(), 2))
print(round(df.loc[df["pdays"] != -1, "pdays"].mean(), 2))
You should see

3705, then 39.77, then 224.87

The honest average, over the 816 people who really were contacted before, is 224.87 days. Not 39.77. The 39.77 was 3,705 rows of "never" being counted as "one day ago, in the wrong direction".

Say what you mean, then fill the gap

The standard tool for a gap is an imputer. Watch what it does here, because this is the trap.

import numpy as np
from sklearn.impute import SimpleImputer

gaps = df[["pdays"]].replace(-1, np.nan)
print(gaps["pdays"].isna().sum())

filler = SimpleImputer(strategy="mean")
filler.fit(gaps)

print(round(filler.statistics_[0], 2))

A second library appears here. numpy is what pandas is built on top of, and it is where the plain number-crunching lives. You will not use much of it directly. as np gives it a short name, the same bargain as as pd on the first line of every notebook you have opened.

.replace(-1, np.nan) turns the sentinel into a real blank. Sentinel is the working name for what Step 0 called a stand-in: a value somebody wrote to mean "no value here". Same thing, and from here on the course uses the shorter word. np.nan is the blank pandas understands: the one .isna() counts, and the one that was not there in Step 0 when you ran df.isna().sum().sum() and got zero.

The double brackets in df[["pdays"]] ask for a table with one column in it, rather than the column on its own. Scikit-learn tools always want a table, because they are built to take many columns at once. Single brackets give you a column; double brackets give you a table. It is worth saying out loud once, because it will bite you.

SimpleImputer(strategy="mean") fills every blank with the average of the ones that are not blank. .statistics_ is what it decided to use, one number per column, so [0] takes the first and only one. The trailing underscore is scikit-learn's mark for "this was learned from data, it was not something you set".

You should see

3705, then 224.87

Read what that would do

It is about to write 224.87 into 3,705 rows. That says: every one of these people was last contacted about 225 days ago.

Not one of them was contacted at all. You would be inventing a phone call for 3,705 people, and the model would believe you, and the number would look perfectly reasonable to everybody who read it afterwards.

The tool is not broken. It did exactly what it was asked. Nobody asked it whether the blanks meant "we do not know" or "it never happened", because it has no way to ask.

Blank almost never means one thing. Learn to sort it into three, because the right fix is different for each.

fixed = ready.copy()
fixed["contacted_before"] = (df["pdays"] != -1).astype(int)
fixed["pdays"] = df["pdays"].where(df["pdays"] != -1, 0)

F_train, F_test, y_train, y_test = train_test_split(
    fixed, y, test_size=0.25, random_state=0)

fixed_tree = DecisionTreeClassifier(max_depth=3, random_state=0)
fixed_tree.fit(F_train, y_train)

print(round(fixed_tree.score(F_test, y_test), 4))

.where(condition, other) keeps the value where the condition is true and puts other everywhere else. So the real gaps in days survive, and every -1 becomes 0.

Two things changed there, not one. The sentinel became 0, and a new column arrived, so fixed has 48 columns rather than 47.

You should see

0.8983

Exactly the same score. Sit with that.

You just did the right thing and were paid nothing for it. Twice over: the sentinel is gone and a new column arrived, and the tree is identical.

Half of that is Step 2's lesson about money in cents. A tree only asks whether a number is below a cut. The smallest real pdays is 1, so swapping -1 for 0 keeps everybody in exactly the same order and no cut can tell. The other half is simpler: contacted_before never won a place in the top three questions, so it is sitting in the table doing nothing yet.

Be careful with the first half of that argument, because it has a condition. If any real value had been 0, the sentinel would have landed on top of it and two different meanings would share a number. Check the smallest real value before you pick a filler.

Who does care? Not much, on this file, today. I measured it: the model in Step 10, which judges people by how far apart they are, moves a little, because -1 against 224 is a distance it takes seriously. The model in Step 9 barely moves at all on this column. What does not survive is the reading: a person who runs .mean() on a column of sentinels and puts 39.77 in a slide is wrong by a factor of five, and no model will warn them.

So the honest reason to do this is not the score. It is that -1 means "never" and the file does not say so anywhere, and every person and every model that meets your table later will take it at face value. You are writing down what you know, in a form the next reader cannot misread.

Part 9PredictModify25 min

What the test rows told your scaler

One column left to fix, and it is the one that carries the whole point of this step.

balance runs from -3313 to 71188. previous runs from 0 to 25. To a tree that is fine, because it never compares one column against another.

To some models it is not fine at all. Anything that measures how far apart two people are adds the columns up, so a difference of 5,000 euros drowns a difference of three phone calls, for no better reason than the units somebody chose. That is Step 10. Anything that is told to keep its numbers small will squash the column with the big units hardest, and that is the model in Step 9. Anything that walks downhill towards an answer walks badly across a lopsided landscape, and that is how the Step 9 model is trained.

Trees, forests and everything built out of them do not care, which covers Steps 2, 11 and 12. It is worth knowing which of your models care rather than scaling out of habit, because a habit you cannot explain is a habit you cannot defend in a review.

The fix is scaling. Squeeze every column onto the same range, usually 0 to 1, by asking where each value sits between the smallest and the largest.

Which raises the question this whole step has been walking towards. The smallest and largest of what?

from sklearn.preprocessing import MinMaxScaler

print(X_train["balance"].min(), X_train["balance"].max())
print(ready["balance"].min(), ready["balance"].max())

The first line looks at the training half only. The second looks at the whole file.

You should see

-2082 71188, then -3313 71188

The poorest person in the file, at -3313, is not in the training half. They are one of the 1,131 people you promised not to look at.

Predict first. This is the marked question of the step.

You are about to scale one training person's balance twice: once with a scaler fitted on the training rows only, and once with a scaler fitted on the whole file.

Will the two numbers be the same? Write yes or no, and why.

honest = MinMaxScaler()
honest.fit(X_train[["balance"]])

leaky = MinMaxScaler()
leaky.fit(ready[["balance"]])

three = X_train[["balance"]].head(3).copy()
three["fitted_on_training"] = honest.transform(three[["balance"]])
three["fitted_on_everything"] = leaky.transform(three[["balance"]])

print(three.round(4))

Three real people from the training half, their balances scaled twice. .head(3) is from Step 0 and .copy() is from Step 2, where it stopped you damaging a table you still needed. .round(4) on a whole table rounds every number in it, so the two new columns fit on one line.

You should see
      balance  fitted_on_training  fitted_on_everything
4384        4              0.0285                0.0445
2560     1071              0.0430                0.0588
1470     4103              0.0844                0.0995

Read that slowly. Take the first row: one person, with 4 euros in the bank, in the training half.

Their number is 0.0285 or 0.0445 depending on nothing whatever about them. It depends on whether a stranger in the test half, the one with -3313, was allowed in the room when the scaler decided where the bottom of the range was. Every row in the table shifts, and every one of them is a training row.

This is the whole lesson of the second half

Fitting on everything and splitting afterwards feels harmless. Nobody looked at the answers. No y was involved anywhere.

But the training rows have been quietly told something about the test rows, and every score you produce afterwards is measured on people who already leaked into the preparation. Your test set is a little less of a stranger than you think it is.

Fit on the training rows. Transform everything. Every single time, for every scaler, every imputer, every encoder, forever.

How much did it cost, here, today?

Nothing measurable, and you deserve to be told that plainly.

I ran it. Fitting the scaler on everything rather than on the training rows changes the tree not at all, and changes the Step 9 and Step 10 models by less than a thousandth, in no reliable direction. A minimum and a maximum are about the weakest thing a preparation step can learn, because they say nothing whatever about who said yes.

So why the fuss? Because the shape of the mistake is the thing, not this instance of it. The same mistake, made with a preparation step that does learn something about the answer, is not a thousandth. Replacing a column of job titles with the yes-rate of each job, fitted over the whole file, hands the test rows' answers straight to the training rows. Choosing which columns to keep by looking at all the rows does the same. Both are ordinary things that ordinary people do, and both are this mistake with the volume turned up.

There is a worse property than being wrong, and this has it: you cannot say which way it is wrong. An error you can sign, you can correct for. An error like this leaves you with a number you cannot defend in either direction.

You are learning the habit now, on a model where getting it wrong is free, so that it is already a habit when it stops being free. Step 15 is where you meet the tool that makes it impossible to get wrong, and it will make no sense at all unless you have done it by hand first.

Part 10Make60 minA sitting of its own

Your own table, in a shape a model takes

No code here and no answer at the end. Your own table, the one you signed up in Step 0.

Use new names

mine, not df. my_ready, not ready. If you reuse the names above, the cells you already ran will still work and print numbers about the wrong table.

The brief

Get every column of your own table into numbers, split it, and train the same depth 3 tree on all of it.

It is done when

If one of your columns has hundreds of different answers

Boxes will give you hundreds of columns, most of them almost all zeroes. A tree will usually shrug; the models in Steps 10 and 13 will not, and either way you have paid a lot of columns for very little. Leave it for now.

Leave the column out today, and write down how many different answers it has. Step 18 is where that gets handled properly, and the number you write down now is the one that makes that step make sense.

If your score got worse after adding all the columns

That happens and it is a real result. More columns give a tree more ways to find something that is only true of your training rows, which is Step 3's lesson arriving in a new place.

Write down both scores and keep going. Step 6 is where you learn which columns were worth their place.

Working alone

Read the answers in one of your word columns out loud to yourself tomorrow, slowly, off the value_counts() output.

You are listening for "those two mean the same thing". Real tables are full of N/A and n/a and Not applicable sitting in one column as three separate answers, and every one of them becomes its own box, and nothing anywhere will tell you. If the data is about work you do, you are the person best placed to hear it.

If you know somebody who works with this data

Read them the same list. Two people hear different pairs, and neither of you hears all of them.

Part 11Ship15 min

Write it up

Add to the README you started in Step 0.

Then say it to somebody

Find a person who does not code. Explain why a column of job titles cannot just be numbered 1 to 12.

Use the five jobs from Part 0 and a pen. If they say "so it thinks a manager is worth three housemaids", they have it, and so do you.

Part 12Check15 min

Check yourself

Six questions. Answer them in writing before you open the answers.

One counts more than the other five. You can get five right and still not pass this step.

  1. You number twelve job titles 1 to 12. Name two things you have told the model that are not true.
  2. housing holds yes and no. Why does it get one column while marital gets three?
  3. Your tree scored 0.8983 with pdays full of -1, and 0.8983 after you fixed it properly. Give the reason, and name one model that would not have shrugged.
  4. This is not the Step 3 question. Read it twice: the colleague is doing something different, and the answer is different.

    A colleague scales every column of her file so they all sit between 0 and 1. Then she splits into training and test rows, trains, and scores 0.88 on the test rows. Her rows are all different people, with no duplicates, and the split itself is done correctly.

    What did the test rows tell her scaler? What is her 0.88 worth, and what should she do instead?

    Answer the first question with something specific, not with the word leakage. Then say, in one line, what you would need to know before you could say how much her 0.88 is off by.

  5. You train on 4,521 people and get 38 box columns. Twenty new people arrive and produce 23. What went wrong, and which tool fixes it?
  6. An imputer offers to fill your blanks with the mean. Give one case where that is right and one where it is wrong, and say what makes the difference.
Open the answers, once all six are written down
  1. Any two of: that the jobs have an order; that the gaps between them are equal; that one sits between two others; that job 12 is in some way twelve times job 1. All four are invented by the numbering and none of them are in the data. Full marks if you also said that a tree only hears the first of the four, and the other three are waiting for Steps 9 and 10.
  2. Because with two answers the second column is the first one upside down. It carries nothing the first does not, and one column of information written twice is one column too many. Three answers cannot be squeezed into one column without inventing an order, so they get three.
  3. A tree only asks whether a number is below a cut, and swapping -1 for 0 does not change which people are below which, because no real pdays is below 1. Nothing moved. The model in Step 10 does move, a little, because it measures how far apart people are and -1 against 224 is a distance it takes seriously. The reader who runs .mean() on the column moves furthest of all: 39.77 against 224.87.
  4. The test rows handed her scaler their own smallest and largest values, so the range it adopted came partly from rows she had promised not to look at. Every training number was then converted using that range. Her training data depends on her test people.

    What her 0.88 is worth is the harder half, and the honest answer is: nobody can say, including her. A minimum and a maximum carry nothing about who said yes, so the error is small, and it is not reliably in either direction. That is worse than a known bias, not better. An error you can sign is an error you can correct for.

    To find out how much it cost her she would have to do it both ways and compare, over enough splits to see past the noise, which is Part 6 of this step.

    She should split first, fit the scaler on the training rows only, then transform both halves with it. Mark yourself right only if you said what the test rows actually handed over, which is the range. Saying "it leaks" is half.

  5. Twenty people do not contain every job and every month, and get_dummies makes columns only for the answers it can see. The fix is an encoder fitted once on the training rows, which remembers all 38 columns and produces the same 38 for anybody, including a person whose job it has never seen.
  6. Right when the value existed and was lost, like a weight nobody wrote down. Wrong when the thing never happened, like pdays, because then the mean invents an event. The difference is what the blank means, and the file cannot tell you: you have to ask whoever collected it.
If you got the marked one wrong

Go back to Part 9 and run the two scalers again, and this time write down where the number -3313 came from before you read anything else.

Every step from here on prepares data before it models. If you do not fully believe that preparation is part of the model, you will do this by accident, and you will be left holding a number you cannot defend in either direction.

StretchOptionalHarder

If you want more

Bronze

Part 2 lists the nine word columns by name, written out by hand. Get pandas to find them for you instead: df.select_dtypes(include=["object", "string"]) hands back only the columns holding text.

Both names are in there for a reason. Older pandas calls a text column object and newer pandas calls it string, so asking for both is how you write one line that works on a colleague's laptop as well as yours. Check that what it finds matches your nine, and note what it does with y.

Silver

Train the depth 3 tree at every width: the six plain columns, then plus the three yes-or-no columns, then plus the boxes. Three scores.

Write down which of the three additions actually paid, and then find the honest way to say the result. One of the three is doing all the work, and a sentence that says "adding the word columns raised the score" is true and misleading at the same time.

Gold

Build the whole thing properly, in order: split the raw file first, fit the encoder on the training rows only, prepare both halves through prepare, and train.

You will get 0.8983 again, and that is the interesting part. Write the paragraph you would put in a code review explaining why the leaky version and the correct version give the same number here, and why that is not an argument for the leaky version.

Before you move on

All seven, or it is a not yet.

The thing you built.
Your own table, every column a number, split, and scored with a depth 3 tree. Before and after in your README.
The decisions.
Every word column sorted into two-answer, boxes, or left out, with a reason. Every gap and sentinel sorted into the three kinds from Part 8.
The thing that remembers.
An encoder fitted on your training rows only, and a prepare function you have run twice.
Check yourself.
Five of six right, including the marked one about the scaler.
The build log.
One thing that broke or surprised you, and one prediction you wrote down that turned out wrong. Parts 5, 6 and 9 all asked you to predict.
It still runs.
Kernel, then Restart Kernel and Run All Cells. One cell prints a ValueError as ordinary black text on purpose. Nothing should turn red.
Ninety seconds, out loud.
Record yourself explaining why job titles cannot be numbered 1 to 12. Every step ends this way.

Keep your 47-column table and the month_oct line. Step 5 starts from the fact that the tree found October on its own, and asks what it could never have found, however deep you let it go and however many columns you gave it.