Nine of the columns you are allowed to use hold words, and every model in this course refuses them. You turn all nine into numbers, hand the model a table with 47 columns instead of 7, and then find out how hard it is to tell whether that helped.
Step 3 sent you to data/bank-names.txt, to the heading related with the last contact of the current campaign, and asked what you made of the columns underneath it. Read your lines back now. If you cannot find them, the next paragraph gives you enough to carry on.
Four columns sit under that heading: contact, day, month and duration. Only day has been in your seven, so if that is all you found, you read the file properly. The second column worth arguing about is campaign, which is filed under other attributes and gives itself away in its own description rather than by its heading.
Here is what I decided, so you can argue with it. day is the day of the month the call happened, and if you are choosing who to ring on the fourteenth then the day is something you know, so it stays. campaign counts the calls made to this person in this campaign including the one being made, so on Monday morning the honest value is one less than the file says. It is off by one, always in the same direction, and a tree cutting at 2.5 does not care much. It stays, with that written down beside it.
Neither is like duration. Both are worth a sentence in your README, because the person who reads your work after you will ask.
Every model in this course from here to the end takes a table of numbers and nothing else. Not words. Not blanks. Not a column where somebody typed -1 to mean "never happened".
A few models outside this course will take words directly, if you set them up to. They are the exception, they are not on this ladder, and they still need everything else in this step.
This step is where you learn to hand a model a table it will accept. It is the least glamorous step in the ladder, and it is where most real mistakes get made.
Your seven-column table has been scoring 0.893 on people it has never seen since Step 3, catching 22 of the 138 who would say yes.
In Part 5 you hand the same model 40 more columns, every job title and every month and everything else the file knows. The score is 0.893. The catches are 22. The wasted calls are 5. Not close to Step 3. The same, to the last digit.
In Part 6 you let it ask one more question and the score goes to 0.8983, catching 28 for the same 5 wasted calls. Six more customers for nothing, and the tree gets there by asking about October, which nobody told it about and which you found by hand in Step 1.
Then, in the same part, you do what Step 3 taught you and check whether that win is real. Run the same comparison over 30 different splits and the wider table wins 11 times out of 30, and loses on average. The split you were handed was the sixth best of the thirty.
That is the step. Not a win, and not a failure: a result you cannot read off one number, and the first honest look at how small a difference has to be before you stop being able to see it.
-1 from Step 0 again, and decide what to do with it.Parts 0 and 1 are twenty minutes on paper, away from the screen. You do not open the notebook until Part 2. If you have less than an hour tonight, do Parts 0 and 1 and stop there. They stand on their own.
Part 5 ends on a shock and Part 5's explanation is the cure, so it wants twenty minutes in one go. If you have less than that left, stop at the end of Part 4.
The good places to stop are the end of Part 2, the end of Part 4, the end of Part 6 and the end of Part 9. Part 10 is a sitting of its own. In three-quarter-hour evenings this step is five or six of them, and that is normal for the longest step in the course.
Paper and pen. Read the table below, copy the five job titles down the left of your page, then shut the laptop.
Here are five of the twelve job titles in the file.
| Job | Your number |
|---|---|
| student | |
| blue-collar | |
| management | |
| housemaid | |
| retired |
The model needs numbers. So give each job a number, 1 to 5, in whatever order feels natural, writing them beside the titles on your paper. It takes ten seconds and there is no wrong answer.
Whatever order you chose, you have told the machine four things you did not mean.
A tree hears the first one and only the first one. It asks whether a number is below a cut, so with your numbering it can ask "is this job below 3.5", which lumps your first three jobs together and splits off the other two. That grouping is an accident of the order you happened to write them in. Somebody else numbering the same five jobs gets a different model, and neither of you meant to say anything about order at all.
The other three lies are wasted on a tree, which never adds or averages anything. They are not wasted on the models in Steps 9 and 10, which do both. Numbering a word column is the kind of mistake that lies quiet under one model and goes off under the next.
Rub out your numbers. Draw five columns instead, one for each job, and write the five job titles across the top.
Now take three people: a student, a housemaid, and a manager. Give each of them a row, and put a 1 in the column that matches their job and a 0 in the other four.
You have three rows of five numbers. No order. No gaps. Nothing between anything.
That is one-hot encoding, and you have just done all of it. One column per possible answer, one 1 per row, zeroes everywhere else. The name is silly and comes from electronics. The idea is a row of tick boxes.
Count the columns you just used. Five jobs became five columns. The real file has twelve job titles, so twelve columns. Six word columns with several answers each will become 38.
That is fine here. It is not always fine. Write down what you think happens to a column holding five thousand different customer names, and keep it. Step 18 is where that problem gets solved properly.
Three of the nine word columns hold only yes and no. Do those need five columns, or two, or one?
Write your answer and why. Part 3 is that question.
Six words. All six turn up in job adverts for this work.
Preparing your data is not something that happens before the model. It is part of the model, and everything Step 3 said about the split applies to it.
Open band-04.ipynb. Every cell on this page is already in it, in this order. You run them; you do not type them. The only typing you do in this step is in Part 10, on your own table.
Run every cell from the top before you start. Nothing is remembered overnight, and the cells later in this step need names that were made earlier in it.
If you see NameError: name 'ready' is not defined, that is all this is. Nothing is broken.
Check that every random_state=0 on the page is in your cell too. Without it the rows are shuffled differently every run, so your numbers change every time you press Shift and Enter and no two people in the world can compare notes.
If the numbers still differ, run every cell from the top in a fresh kernel: Kernel, then Restart Kernel and Run All Cells.
Step 2 ended on a refusal. You handed the tree a column of job titles and got ValueError: could not convert string to float: 'unemployed'. Since then you have worked with seven columns: the six that were already numbers, and said_yes_before, which you built by hand in Step 2 because you needed one word column badly enough to do it yourself.
Here is what you have been leaving out.
import pandas as pd
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split
df = pd.read_csv("data/bank.csv", sep=";")
y = (df["y"] == "yes").astype(int)
words = ["job", "marital", "education", "default",
"housing", "loan", "contact", "month", "poutcome"]
for name in words:
print(name, df[name].nunique())
.nunique() counts how many different answers a column holds. The loop prints the name and the count for each of the nine, one line each. You met value_counts() in Step 0, which shows every answer and how often it appears; this is the short version, when all you want is how many kinds there are.
job 12
marital 3
education 4
default 2
housing 2
loan 2
contact 3
month 12
poutcome 4
Nine columns. Three of them hold two answers. Six hold more. That split is the plan for the next two parts, and it is the answer to the question you wrote at the end of Part 0.
Jupyter is not standing in the course folder. Close it, open a terminal, type cd and a space and the path to the folder that holds data, press Enter, then type jupyter lab and press Enter. Parts 6 and 7 of the setup page walk through it slowly.
A column holding only yes and no does not need tick boxes. Two boxes would be one column of information written twice, because the second is always whatever the first is not.
One column is enough, and you have done this before. In Step 2 you wrote (df["poutcome"] == "success").astype(int) to make said_yes_before. Same move, three times.
ready = df[["age", "balance", "day", "campaign", "pdays", "previous"]].copy()
for name in ["default", "housing", "loan"]:
ready[name] = (df[name] == "yes").astype(int)
print(ready.shape)
print(ready["housing"].sum())
The loop does the same thing to three columns rather than copying the line out three times. ready starts as the six columns that were already numbers, and gains three more.
Note what is not in ready: said_yes_before is gone. You built that by hand in Step 2 because you needed one word column badly enough to do it manually. In Part 4 the machine builds it for you, along with 37 others, and it will be called poutcome_success.
(4521, 9), then 2559
Nine columns, and 2,559 of the 4,521 people have a housing loan. The notes say housing: has housing loan? and nothing more, so a housing loan is what you call it, not a mortgage.
For a tree it makes no difference at all. It will find the same cut either way, because the two groups are the same two groups.
It makes a difference to you at two in the morning. Pick the convention that 1 means the thing the column is named after, and never break it. housing is 1 for a person with a housing loan. If half your columns say yes with a 1 and the other half say yes with a 0, every sentence you write about the model will be wrong half the time.
Now the six columns with more than two answers. This is Part 0, done by machine.
boxes = pd.get_dummies(df[["job", "marital", "education",
"contact", "month", "poutcome"]], dtype=int)
print(boxes.shape)
print(list(boxes.columns[:4]))
pd.get_dummies does exactly what you did on paper. It reads a column, finds every different answer in it, and makes one new column per answer, holding 1 for the rows that match and 0 for the rest.
The [:4] on the last line means "the first four". A colon inside square brackets asks for a run of things rather than one thing, and with nothing before the colon it starts at the beginning. There are 38 names and four is enough to see the pattern.
dtype=int asks for 1 and 0 rather than True and False. Both work. Ones and zeroes are easier to read when you print the table, and you are going to print it.
The word "dummies" is a statistics term for these columns and it means nothing useful. Read it as "tick boxes" every time you see it.
(4521, 38), then ['job_admin.', 'job_blue-collar', 'job_entrepreneur', 'job_housemaid']
Six columns became 38, named for the column they came from and the answer they stand for.
Check that it did what you think it did, rather than trusting the shape.
print(boxes["job_student"].sum())
print((df["job"] == "student").sum())
Adding up a column of 1s and 0s counts the 1s. The second line counts the students the Step 1 way, by comparing a column to a value and adding up the Trues. That shape, a column of True and False you use to pick people out, is called a mask. Step 1 had you write them all evening without giving them a name; the name is worth having now, because Part 6 stacks three of them.
The two numbers have to match, or something is wrong.
84, then 84
Nobody made you run that second cell. Get used to running it anyway.
Every time you reshape a table you should check one thing you can count two ways. Nothing about this step turns red when it goes wrong. It just quietly gives you a different table from the one you think you have.
Join the two halves together.
ready = pd.concat([ready, boxes], axis=1)
print(ready.shape)
pd.concat you met in Step 3, where you stacked the file on top of itself. That used the default, axis=0, which means downwards, adding rows. axis=1 means sideways, adding columns. Same tool, turned ninety degrees.
(4521, 47)
Six numbers, three yes-or-no columns, 38 boxes. Every one of them a number, and not a word left anywhere.
Since Step 3 your best honest score has been 0.893, catching 22 of the 138 and wasting 5 calls. That came from seven columns.
The model is about to get 47. Every job title, every month, every education level, everything the file knows and you have been throwing away.
Same depth 2 as Step 3, so that the only thing that changes is the width of the table. What does it score? Write the number down before you run it.
X_train, X_test, y_train, y_test = train_test_split(
ready, y, test_size=0.25, random_state=0)
flat = DecisionTreeClassifier(max_depth=2, random_state=0)
flat.fit(X_train, y_train)
print(round(flat.score(X_test, y_test), 4))
flat_guess = flat.predict(X_test)
print(((flat_guess == 1) & (y_test == 1)).sum())
print(((flat_guess == 1) & (y_test == 0)).sum())
The same split as Step 3, with the same random_state=0, so the same people are held back. The same depth. The only change in the world is that the table is 40 columns wider.
0.893, then 22, then 5
Go back and look at Step 3, Part 4. Those are the same three numbers. Not similar. The same.
You did four parts of work. You turned nine columns of words into 41 columns of numbers, correctly, and checked them. The model was handed all of it.
It changed nothing at all. Write down what you think happened, in one line, before you read on.
Here is the part that surprises people, and it is not what you probably wrote down. Print the tree and one of your new columns is in it: with 47 columns to choose from, the second question is month_oct, where the seven-column tree asked about age. Your work did win a place.
It bought nothing, because both branches under that question still say no. The tree is different and the answers are not, which is the same thing Part 4 of Step 2 showed you when a split changed nothing.
A depth 2 tree gets two questions and stops. Your new columns are not useless. They are just not strong enough to flip an answer inside two questions, and a tree with two questions never gets to the third.
One number for scale, since this is the step about not trusting one split: run the two tables against each other over thirty splits and they give identical predictions on sixteen of them. On the other fourteen they differ by a handful of people, in both directions.
This is worth more to you than a win would have been. Most of the work in a real project is like this: correct, necessary, and worth nothing on its own.
The columns you just built pay off in Part 6, in Step 5, and in every model from Step 8 on, all of which can hold more than two ideas at once. Work that pays later still has to be done right now, and nobody claps.
Give it a third question and see whether the new columns get used.
tree = DecisionTreeClassifier(max_depth=3, random_state=0)
tree.fit(X_train, y_train)
print(round(tree.score(X_test, y_test), 4))
guess = tree.predict(X_test)
print(((guess == 1) & (y_test == 1)).sum())
print(((guess == 1) & (y_test == 0)).sum())
0.8983, then 28, then 5
0.893 became 0.8983. Twenty-two caught became 28. Wasted calls stayed at 5.
Six more customers for nothing. Write it in your README, and then keep reading, because this part is not over.
print(export_text(tree, feature_names=list(ready.columns)))
|--- poutcome_success <= 0.50
| |--- month_oct <= 0.50
| | |--- age <= 60.50
| | | |--- class: 0
| | |--- age > 60.50
| | | |--- class: 0
| |--- month_oct > 0.50
| | |--- day <= 16.50
| | | |--- class: 0
| | |--- day > 16.50
| | | |--- class: 1
|--- poutcome_success > 0.50
| |--- balance <= 8053.00
| | |--- education_tertiary <= 0.50
| | | |--- class: 1
| | |--- education_tertiary > 0.50
| | | |--- class: 1
| |--- balance > 8053.00
| | |--- class: 0
month_oct.
In Step 1 you sat with a table of twelve months and worked out by hand that March, September, October and December were different from the rest. You wrote a wider rule around them and it scored 0.8821, which was worse than saying no to everybody, and you kept the line anyway because the months themselves were a real finding badly used.
Nobody told this tree anything about months. It got twelve columns named after month names, in alphabetical order, meaning nothing. It went and found one of your four.
It found something you did not give it, which is the same thing that happened in Step 2 when it found your Step 1 rule. A model is not cleverer than you. It is faster, and it never gets bored, and if the thing is there it will find it.
Read the rule properly before you believe your own summary of it. It is not "October". It is October, after the sixteenth. That is one branch of one tree, and you are about to find out how much weight it will take.
Look at age <= 60.50. Both branches under it end in class: 0. Look at education_tertiary. Both branches end in class: 1.
The decision is the same on both sides, so as far as the yes-or-no answer goes, those two questions could be deleted. They are not decoration though, and it matters that you know why: the tree picked them because they were the best splits it had left, and each one changes how sure the model is. On the university branch it moves from about half the people saying yes to about four in five.
Today you are reading a yes or a no, so you cannot see that. In Step 9 you start reading the chance instead, and those two lines come back to life.
Step 3's Silver task asked you to run the same split with five different random_state values and look at the spread. Here is that task, made compulsory, on a result you have a reason to want to be true.
You are about to run the same comparison over 30 different splits: the wide table with three questions against the Step 3 table with two.
How many of the 30 will the wide table win? Write the number down. Nobody will see it but you.
seven = df[["age", "balance", "day", "campaign", "pdays", "previous"]].copy()
seven["said_yes_before"] = (df["poutcome"] == "success").astype(int)
wide_scores = []
narrow_scores = []
wins = 0
for seed in range(30):
a, b, c, d = train_test_split(ready, y, test_size=0.25, random_state=seed)
wide = DecisionTreeClassifier(max_depth=3, random_state=0)
wide.fit(a, c)
wide_score = wide.score(b, d)
a, b, c, d = train_test_split(seven, y, test_size=0.25, random_state=seed)
narrow = DecisionTreeClassifier(max_depth=2, random_state=0)
narrow.fit(a, c)
narrow_score = narrow.score(b, d)
wide_scores.append(wide_score)
narrow_scores.append(narrow_score)
if wide_score > narrow_score:
wins = wins + 1
print(round(min(wide_scores), 4), round(max(wide_scores), 4))
print(round(sum(wide_scores) / 30, 4))
print(round(sum(narrow_scores) / 30, 4))
print(wins)
Longer than usual, and every line of it is something you have done before. Four things are new and none is hard.
wide_scores = [] makes an empty list, which is a place to put things in order. Nothing is in it yet.wide_scores.append(wide_score) puts one number on the end of it. Thirty passes, thirty numbers, in the order they were made. Every loop you have written so far printed as it went and kept nothing; this one keeps everything and prints at the end.range(30) hands out the numbers 0 to 29, one per pass, and each becomes a random_state. So each pass holds a different quarter of the people back and everything else is held still.min, max and sum you have used as df["age"].min(), on the end of a column. On a plain list of numbers you write them the other way round, with the list inside the brackets. Same job.The wins counter starts at 0, and wins = wins + 1 means take what is in wins, add one, and put it back. It only happens on the passes where the wide table came out ahead.
0.8753 0.9045, then 0.8902, then 0.8915, then 11
The wide table scores anywhere from 0.8753 to 0.9045 depending on nothing but which quarter of the people you happened to hold back. That is a spread of about three points, and the win you were celebrating was half a point.
On average the wide table gets 0.8902 and the Step 3 table gets 0.8915. The wide one is slightly worse.
It won 11 of the 30. If it were a coin you would expect 15.
The split you were handed, random_state=0, was the sixth best of the thirty for the wide table. You did not choose it. It was chosen in Step 3, for other reasons, long before any of this. And it happened to be the one where the new columns look good.
Two things differ between the two models: the wide one has 47 columns and three questions, the narrow one has 7 and two. So this tells you about the pair as a whole, not about columns or depth on their own.
You already know one half of it, because Part 5 held depth still and the columns bought nothing. Separating the rest properly is Step 6's job.
It does not mean the work was wasted. Every model from Step 8 onwards can use forty columns properly, and not one of them can use a column of words. You did the work that makes them possible.
It does not mean the extra columns are worthless. It means this measurement cannot tell, because the difference you are chasing is smaller than the noise in the way you are measuring.
What it means is that "0.893 became 0.8983" was never a sentence you were entitled to write. You would have written it. I did write it, in the first version of this page, and somebody who ran 30 splits took it out.
Before you believe a difference between two numbers, find out how much each number moves on its own when nothing important has changed.
One split gives you one number and no idea how much it wobbles. Thirty splits cost you four seconds and tell you whether you are looking at a result or at weather.
Step 14 is where this becomes a proper tool with a name. You do not need the name to do it, and you have just done it.
One more thing before you leave it. The whole difference in Part 6 was six extra customers caught. Go and look at them.
october_leaf = ((X_test["poutcome_success"] == 0) &
(X_test["month_oct"] == 1) &
(X_test["day"] > 16.5))
print(october_leaf.sum())
print((october_leaf & (y_test == 1)).sum())
Three masks with & between them, which is the shape from Part 4 stacked three deep. Each set of brackets is one question from the tree, read straight off the printed rules above, and together they pick out exactly the people who land in the leaf that says yes about October.
The first line counts them. The second adds one more condition, that they really said yes, and counts those.
6, then 6
Six people in the test half reach that leaf. All six said yes.
Your entire improvement is six people, and every one of them went the right way. On the training side that leaf holds 34 people, 20 of whom said yes, which is a real signal and not nothing. But six out of six is the kind of luck that does not repeat, and now you know exactly where the half point came from.
Step 6 is where you learn to ask which columns are pulling their weight, properly, instead of squinting at one tree.
You have a model that works. Now use it, which is the point of having one.
Twenty people from the test half arrive as they would in real life: raw, in the shape the file has them, not the shape the model wants.
newcomers = df.loc[X_test.index[:20]]
new_boxes = pd.get_dummies(newcomers[["job", "marital", "education",
"contact", "month", "poutcome"]], dtype=int)
print(new_boxes.shape)
df.loc[X_test.index[:20]] takes the first twenty row numbers from the test half and pulls those rows out of the original file. They are people the model has never seen, in their original words.
(20, 23)
The model was trained on 38 box columns. These twenty people produced 23.
Work out why before you read on. It is not a bug in pandas.
Twenty people do not have twelve different jobs between them. They do not cover all twelve months. get_dummies makes one column per answer it can see, and twenty people cannot show it everything 4,521 people could.
So the table has the wrong shape, and the columns that do exist are in the wrong order. Hand that to the model and see what it says.
print(len(new_boxes.columns), len(X_train.columns))
try:
tree.predict(new_boxes)
except ValueError as error:
print("ValueError:", error)
len(...) counts things. You met it in Step 1 as len(df), counting rows; len(new_boxes.columns) counts the names in the list of columns instead.
The try and except shape is the one from Step 2: attempt this, and if it goes wrong, catch the complaint and print it as ordinary text so the rest of the notebook still runs.
23 47, then a complaint several lines long that begins:
ValueError: The feature names should match those that were passed during fit.
and then lists the columns it expected and cannot find. Read the first line. The list underneath is the detail, and it is worth a glance because it names the plain columns like age too, not only the boxes.
This is one of the standard ways a model that worked in a notebook fails the first time somebody tries to use it, and I have watched it happen. Nothing about the model is wrong. The preparation was done twice, by two different people or by the same person on two different days, and the second one did not know what the first had decided.
get_dummies looks at whatever you hand it and decides the columns on the spot. You need something that decides once, on the training rows, writes the decision down, and applies the same decision forever after.
from sklearn.preprocessing import OneHotEncoder
many = ["job", "marital", "education", "contact", "month", "poutcome"]
plain = ["age", "balance", "day", "campaign", "pdays", "previous"]
raw_train = df.loc[X_train.index]
encoder = OneHotEncoder(sparse_output=False, handle_unknown="ignore")
encoder.set_output(transform="pandas")
encoder.fit(raw_train[many])
print(len(encoder.get_feature_names_out()))
df.loc[X_train.index] is the same move as the one you used on the newcomers: take the row numbers of the training half and pull those rows, in their original words, out of the file.
Four new things, and they are the shape of every tool in the rest of this course.
encoder.fit(...) looks at the training rows and remembers every answer it saw. It is the encoder learning, and there is nothing to catch on the left of an equals sign.handle_unknown="ignore" decides what happens when a job title turns up that was not in the training rows. Ignore means all twelve boxes stay 0 for that person. The alternative is to stop with an error. Ignoring is usually right, and you should know you chose it.sparse_output=False asks for an ordinary grid of numbers. Without it you get a compressed thing that saves memory on very wide tables and prints as gibberish.set_output(transform="pandas") asks it to hand back a proper pandas table with column names on, rather than that bare grid. Without this line you get numbers with no names, and putting the names back is a fiddly job you would then have to do by hand every time.get_feature_names_out() is the list of column names it decided on, and len counts them.
Notice which rows it was fitted on: raw_train only. Part 9 is about why.
38
The same 38 columns, decided once and written down inside the encoder.
Now one function that puts any set of raw rows into the model's shape.
def prepare(raw):
out = raw[plain].copy()
for name in ["default", "housing", "loan"]:
out[name] = (raw[name] == "yes").astype(int)
return pd.concat([out, encoder.transform(raw[many])], axis=1)
print(prepare(newcomers).shape)
A def gives a block of work a name so you can use it again without copying it out. Everything indented under the first line is the work, and return hands the answer back to whoever asked.
raw is a stand-in. It means "whatever table somebody hands this thing", and inside the block that is what raw refers to. Below, prepare(newcomers) hands it the twenty newcomers, so on that run every raw in the block means newcomers. Tomorrow you hand it a different table and every raw means that one.
plain, many and encoder are not handed in. The block reaches out and uses them where they sit. That is fine and normal, and it is also the reason a function like this stops working if you rebuild the encoder later without rerunning it.
encoder.transform(...) is the other half of fit. Fit remembered; transform applies. Because of the set_output line it hands back a table with the 38 names already on it, which pd.concat then joins sideways to the nine you built by hand.
This is the first function in the course and it will not be the last. The alternative is doing these five lines in three places and getting one of them wrong.
(20, 47)
Twenty people, 47 columns, in the same order the model was trained on. Now it works.
print(tree.predict(prepare(newcomers)).sum())
0
It says no to all twenty. Roughly 12 people in 100 say yes, and this model only says yes to about 3 in 100, so twenty people producing no yeses at all is ordinary. It is not the error you just fixed coming back.
One line: what get_dummies does that OneHotEncoder does not, and when you would still use get_dummies.
There is a real answer to the second half. It is a good tool for looking at a table. It is a bad tool for feeding a model you intend to use twice.
In Step 0 you found that pdays has an average of 39.77, and that the average was a lie, because 3,705 of the 4,521 rows hold -1, which does not mean minus one day. It means this person was never contacted before.
That column is sitting in your model right now. Time to deal with it.
print((df["pdays"] == -1).sum())
print(round(df["pdays"].mean(), 2))
print(round(df.loc[df["pdays"] != -1, "pdays"].mean(), 2))
3705, then 39.77, then 224.87
The honest average, over the 816 people who really were contacted before, is 224.87 days. Not 39.77. The 39.77 was 3,705 rows of "never" being counted as "one day ago, in the wrong direction".
The standard tool for a gap is an imputer. Watch what it does here, because this is the trap.
import numpy as np
from sklearn.impute import SimpleImputer
gaps = df[["pdays"]].replace(-1, np.nan)
print(gaps["pdays"].isna().sum())
filler = SimpleImputer(strategy="mean")
filler.fit(gaps)
print(round(filler.statistics_[0], 2))
A second library appears here. numpy is what pandas is built on top of, and it is where the plain number-crunching lives. You will not use much of it directly. as np gives it a short name, the same bargain as as pd on the first line of every notebook you have opened.
.replace(-1, np.nan) turns the sentinel into a real blank. Sentinel is the working name for what Step 0 called a stand-in: a value somebody wrote to mean "no value here". Same thing, and from here on the course uses the shorter word. np.nan is the blank pandas understands: the one .isna() counts, and the one that was not there in Step 0 when you ran df.isna().sum().sum() and got zero.
The double brackets in df[["pdays"]] ask for a table with one column in it, rather than the column on its own. Scikit-learn tools always want a table, because they are built to take many columns at once. Single brackets give you a column; double brackets give you a table. It is worth saying out loud once, because it will bite you.
SimpleImputer(strategy="mean") fills every blank with the average of the ones that are not blank. .statistics_ is what it decided to use, one number per column, so [0] takes the first and only one. The trailing underscore is scikit-learn's mark for "this was learned from data, it was not something you set".
3705, then 224.87
It is about to write 224.87 into 3,705 rows. That says: every one of these people was last contacted about 225 days ago.
Not one of them was contacted at all. You would be inventing a phone call for 3,705 people, and the model would believe you, and the number would look perfectly reasonable to everybody who read it afterwards.
The tool is not broken. It did exactly what it was asked. Nobody asked it whether the blanks meant "we do not know" or "it never happened", because it has no way to ask.
Blank almost never means one thing. Learn to sort it into three, because the right fix is different for each.
fixed = ready.copy()
fixed["contacted_before"] = (df["pdays"] != -1).astype(int)
fixed["pdays"] = df["pdays"].where(df["pdays"] != -1, 0)
F_train, F_test, y_train, y_test = train_test_split(
fixed, y, test_size=0.25, random_state=0)
fixed_tree = DecisionTreeClassifier(max_depth=3, random_state=0)
fixed_tree.fit(F_train, y_train)
print(round(fixed_tree.score(F_test, y_test), 4))
.where(condition, other) keeps the value where the condition is true and puts other everywhere else. So the real gaps in days survive, and every -1 becomes 0.
Two things changed there, not one. The sentinel became 0, and a new column arrived, so fixed has 48 columns rather than 47.
0.8983
You just did the right thing and were paid nothing for it. Twice over: the sentinel is gone and a new column arrived, and the tree is identical.
Half of that is Step 2's lesson about money in cents. A tree only asks whether a number is below a cut. The smallest real pdays is 1, so swapping -1 for 0 keeps everybody in exactly the same order and no cut can tell. The other half is simpler: contacted_before never won a place in the top three questions, so it is sitting in the table doing nothing yet.
Be careful with the first half of that argument, because it has a condition. If any real value had been 0, the sentinel would have landed on top of it and two different meanings would share a number. Check the smallest real value before you pick a filler.
Who does care? Not much, on this file, today. I measured it: the model in Step 10, which judges people by how far apart they are, moves a little, because -1 against 224 is a distance it takes seriously. The model in Step 9 barely moves at all on this column. What does not survive is the reading: a person who runs .mean() on a column of sentinels and puts 39.77 in a slide is wrong by a factor of five, and no model will warn them.
So the honest reason to do this is not the score. It is that -1 means "never" and the file does not say so anywhere, and every person and every model that meets your table later will take it at face value. You are writing down what you know, in a form the next reader cannot misread.
One column left to fix, and it is the one that carries the whole point of this step.
balance runs from -3313 to 71188. previous runs from 0 to 25. To a tree that is fine, because it never compares one column against another.
To some models it is not fine at all. Anything that measures how far apart two people are adds the columns up, so a difference of 5,000 euros drowns a difference of three phone calls, for no better reason than the units somebody chose. That is Step 10. Anything that is told to keep its numbers small will squash the column with the big units hardest, and that is the model in Step 9. Anything that walks downhill towards an answer walks badly across a lopsided landscape, and that is how the Step 9 model is trained.
Trees, forests and everything built out of them do not care, which covers Steps 2, 11 and 12. It is worth knowing which of your models care rather than scaling out of habit, because a habit you cannot explain is a habit you cannot defend in a review.
The fix is scaling. Squeeze every column onto the same range, usually 0 to 1, by asking where each value sits between the smallest and the largest.
Which raises the question this whole step has been walking towards. The smallest and largest of what?
from sklearn.preprocessing import MinMaxScaler
print(X_train["balance"].min(), X_train["balance"].max())
print(ready["balance"].min(), ready["balance"].max())
The first line looks at the training half only. The second looks at the whole file.
-2082 71188, then -3313 71188
The poorest person in the file, at -3313, is not in the training half. They are one of the 1,131 people you promised not to look at.
You are about to scale one training person's balance twice: once with a scaler fitted on the training rows only, and once with a scaler fitted on the whole file.
Will the two numbers be the same? Write yes or no, and why.
honest = MinMaxScaler()
honest.fit(X_train[["balance"]])
leaky = MinMaxScaler()
leaky.fit(ready[["balance"]])
three = X_train[["balance"]].head(3).copy()
three["fitted_on_training"] = honest.transform(three[["balance"]])
three["fitted_on_everything"] = leaky.transform(three[["balance"]])
print(three.round(4))
Three real people from the training half, their balances scaled twice. .head(3) is from Step 0 and .copy() is from Step 2, where it stopped you damaging a table you still needed. .round(4) on a whole table rounds every number in it, so the two new columns fit on one line.
balance fitted_on_training fitted_on_everything
4384 4 0.0285 0.0445
2560 1071 0.0430 0.0588
1470 4103 0.0844 0.0995
Read that slowly. Take the first row: one person, with 4 euros in the bank, in the training half.
Their number is 0.0285 or 0.0445 depending on nothing whatever about them. It depends on whether a stranger in the test half, the one with -3313, was allowed in the room when the scaler decided where the bottom of the range was. Every row in the table shifts, and every one of them is a training row.
Fitting on everything and splitting afterwards feels harmless. Nobody looked at the answers. No y was involved anywhere.
But the training rows have been quietly told something about the test rows, and every score you produce afterwards is measured on people who already leaked into the preparation. Your test set is a little less of a stranger than you think it is.
Fit on the training rows. Transform everything. Every single time, for every scaler, every imputer, every encoder, forever.
Nothing measurable, and you deserve to be told that plainly.
I ran it. Fitting the scaler on everything rather than on the training rows changes the tree not at all, and changes the Step 9 and Step 10 models by less than a thousandth, in no reliable direction. A minimum and a maximum are about the weakest thing a preparation step can learn, because they say nothing whatever about who said yes.
So why the fuss? Because the shape of the mistake is the thing, not this instance of it. The same mistake, made with a preparation step that does learn something about the answer, is not a thousandth. Replacing a column of job titles with the yes-rate of each job, fitted over the whole file, hands the test rows' answers straight to the training rows. Choosing which columns to keep by looking at all the rows does the same. Both are ordinary things that ordinary people do, and both are this mistake with the volume turned up.
There is a worse property than being wrong, and this has it: you cannot say which way it is wrong. An error you can sign, you can correct for. An error like this leaves you with a number you cannot defend in either direction.
You are learning the habit now, on a model where getting it wrong is free, so that it is already a habit when it stops being free. Step 15 is where you meet the tool that makes it impossible to get wrong, and it will make no sense at all unless you have done it by hand first.
No code here and no answer at the end. Your own table, the one you signed up in Step 0.
mine, not df. my_ready, not ready. If you reuse the names above, the cells you already ran will still work and print numbers about the wrong table.
Get every column of your own table into numbers, split it, and train the same depth 3 tree on all of it.
.nunique() for each, and decided for every one whether it is a two-answer column, a boxes column, or something else.prepare function, or something like it, that takes raw rows and hands back model-shaped rows, and you have run it on five rows of your own to prove it works twice.value_counts() on every word column and min() and max() on every number column, the Step 0 way. Anything odd goes on the list.Boxes will give you hundreds of columns, most of them almost all zeroes. A tree will usually shrug; the models in Steps 10 and 13 will not, and either way you have paid a lot of columns for very little. Leave it for now.
Leave the column out today, and write down how many different answers it has. Step 18 is where that gets handled properly, and the number you write down now is the one that makes that step make sense.
That happens and it is a real result. More columns give a tree more ways to find something that is only true of your training rows, which is Step 3's lesson arriving in a new place.
Write down both scores and keep going. Step 6 is where you learn which columns were worth their place.
Read the answers in one of your word columns out loud to yourself tomorrow, slowly, off the value_counts() output.
You are listening for "those two mean the same thing". Real tables are full of N/A and n/a and Not applicable sitting in one column as three separate answers, and every one of them becomes its own box, and nothing anywhere will tell you. If the data is about work you do, you are the person best placed to hear it.
Read them the same list. Two people hear different pairs, and neither of you hears all of them.
Add to the README you started in Step 0.
day and campaign from Step 3, and whether you still agree with them.month_oct: what you found by hand in Step 1, what the tree found on its own, and how many test people were sitting in that leaf.Find a person who does not code. Explain why a column of job titles cannot just be numbered 1 to 12.
Use the five jobs from Part 0 and a pen. If they say "so it thinks a manager is worth three housemaids", they have it, and so do you.
Six questions. Answer them in writing before you open the answers.
One counts more than the other five. You can get five right and still not pass this step.
housing holds yes and no. Why does it get one column while marital gets three?pdays full of -1, and 0.8983 after you fixed it properly. Give the reason, and name one model that would not have shrugged.This is not the Step 3 question. Read it twice: the colleague is doing something different, and the answer is different.
A colleague scales every column of her file so they all sit between 0 and 1. Then she splits into training and test rows, trains, and scores 0.88 on the test rows. Her rows are all different people, with no duplicates, and the split itself is done correctly.
What did the test rows tell her scaler? What is her 0.88 worth, and what should she do instead?
Answer the first question with something specific, not with the word leakage. Then say, in one line, what you would need to know before you could say how much her 0.88 is off by.
-1 for 0 does not change which people are below which, because no real pdays is below 1. Nothing moved. The model in Step 10 does move, a little, because it measures how far apart people are and -1 against 224 is a distance it takes seriously. The reader who runs .mean() on the column moves furthest of all: 39.77 against 224.87.The test rows handed her scaler their own smallest and largest values, so the range it adopted came partly from rows she had promised not to look at. Every training number was then converted using that range. Her training data depends on her test people.
What her 0.88 is worth is the harder half, and the honest answer is: nobody can say, including her. A minimum and a maximum carry nothing about who said yes, so the error is small, and it is not reliably in either direction. That is worse than a known bias, not better. An error you can sign is an error you can correct for.
To find out how much it cost her she would have to do it both ways and compare, over enough splits to see past the noise, which is Part 6 of this step.
She should split first, fit the scaler on the training rows only, then transform both halves with it. Mark yourself right only if you said what the test rows actually handed over, which is the range. Saying "it leaks" is half.
get_dummies makes columns only for the answers it can see. The fix is an encoder fitted once on the training rows, which remembers all 38 columns and produces the same 38 for anybody, including a person whose job it has never seen.pdays, because then the mean invents an event. The difference is what the blank means, and the file cannot tell you: you have to ask whoever collected it.Go back to Part 9 and run the two scalers again, and this time write down where the number -3313 came from before you read anything else.
Every step from here on prepares data before it models. If you do not fully believe that preparation is part of the model, you will do this by accident, and you will be left holding a number you cannot defend in either direction.
Part 2 lists the nine word columns by name, written out by hand. Get pandas to find them for you instead: df.select_dtypes(include=["object", "string"]) hands back only the columns holding text.
Both names are in there for a reason. Older pandas calls a text column object and newer pandas calls it string, so asking for both is how you write one line that works on a colleague's laptop as well as yours. Check that what it finds matches your nine, and note what it does with y.
Train the depth 3 tree at every width: the six plain columns, then plus the three yes-or-no columns, then plus the boxes. Three scores.
Write down which of the three additions actually paid, and then find the honest way to say the result. One of the three is doing all the work, and a sentence that says "adding the word columns raised the score" is true and misleading at the same time.
Build the whole thing properly, in order: split the raw file first, fit the encoder on the training rows only, prepare both halves through prepare, and train.
You will get 0.8983 again, and that is the interesting part. Write the paragraph you would put in a code review explaining why the leaky version and the correct version give the same number here, and why that is not an argument for the leaky version.
All seven, or it is a not yet.
prepare function you have run twice.ValueError as ordinary black text on purpose. Nothing should turn red.Keep your 47-column table and the month_oct line. Step 5 starts from the fact that the tree found October on its own, and asks what it could never have found, however deep you let it go and however many columns you gave it.