Twelve shops on a grid, and no line straight up or straight across separates them. That is a picture of something a tree cannot say cheaply, and of the one column that fixes it. Then four sensible columns for the bank file, measured properly, and not one of them pays.
Halfway through Step 4 a tree found October on its own, out of twelve month columns nobody explained to it. Then the rest of that step was about how little you could safely conclude from it.
This step starts from the other side of that. There are things it can only reach by building a staircase, however deep you let it go and however many columns you hand it, and they are not exotic. The commonest one is the gap between two numbers.
A model can only use what it can say cheaply. A tree says "is this number below this cut", one question at a time, so a relationship between two columns costs it many questions and it still only gets close. Hand the relationship over as a column and it costs one.
In the first half you meet a small table where the answer is exactly that: a gap. You watch a tree fail on it, then hand it one column you built yourself, and watch a single question get every person right.
In the second half you go back to the bank file and try the same trick four times, properly measured, and it does not work once. Then you work out why, which is worth more than the trick.
The two halves together are the actual skill: knowing which missing column is worth building, and being willing to find out that the answer is none of them.
Parts 0 and 1 are half an hour on paper, and you do not open the notebook until Part 2. Part 0 takes longer than it looks, because plotting twelve points by hand is slower than reading about it.
If you have an hour and a half, Parts 0 to 4 are a whole story and a good place to stop. If you have twenty minutes, this is not the night for Step 5.
If three quarters of an hour is what you always have, here is the whole step mapped to it: Parts 0 and 1 the first night, Parts 2 and 3 the second, Part 4 the third, Parts 5 and 6 the fourth, Part 7 on its own the fifth. Parts 5 and 6 are one idea in two halves and must not be split. Part 5 ends by naming the column that did the damage; Part 6 is the reason.
Paper and pen. You want a page with room for a square about the size of your hand.
A chain of shops. For each shop you know the month it opened, counted from the month the company started, and the month a particular sale happened, counted the same way. Some sales were big and some were not.
Here are twelve real rows out of the table you will open in Part 2.
| Opened, month | Sale, month | Big sale |
|---|---|---|
| 5 | 48 | yes |
| 8 | 17 | no |
| 14 | 83 | yes |
| 21 | 69 | yes |
| 25 | 82 | yes |
| 30 | 73 | yes |
| 35 | 40 | no |
| 40 | 48 | no |
| 44 | 53 | no |
| 58 | 114 | yes |
| 61 | 64 | no |
| 64 | 69 | no |
Draw two lines for the edges of a square, one along the bottom and one up the left.
Along the bottom is opened. Mark 0 at the left, then 20, 40 and 60, evenly spaced. Up the left is sale. Mark 0 at the bottom, then 40, 80 and 120, evenly spaced. Squared paper makes this easier, and the back of an envelope is fine.
Now put a mark for each of the twelve shops: across to its opened month, up to its sale month. Use a small circle for the big sales and a cross for the rest. Twelve marks in all. Being a few millimetres out will not matter.
Draw one straight line, either straight up and down or straight across, with circles on one side of it and crosses on the other.
Spend three minutes on it. Try up and down. Try across. Count how many of the twelve your best line gets right, and write the number down.
You cannot do it with one line. The best line across gets 10 of the 12, and so does the best line up and down. Never better, either way.
Now break the rule. Draw one line at any angle you like.
A line running diagonally, up and to the right, puts all six circles above it and all six crosses below. Every one of the twelve, and with room to spare: you can be quite sloppy about the angle and still get all of them.
Here is the same thing drawn properly, so you can check yours. The left panel is the best straight-across line there is, with the two shops it gets wrong. The right panel is one line at an angle.
Two straight lines, one across and one up and down, get eleven of the twelve. Three get all twelve, which is worth knowing, because that is exactly what a tree with three questions would do. It is not that a tree cannot separate these twelve. It is what it costs, and what happens to the thirteenth shop it has never seen.
Look at the twelve rows again. Add a fourth column on your paper and, for each shop, subtract the opened month from the sale month. Twelve subtractions, done on the page rather than in your head.
Then write one sentence saying what the circles have in common that the crosses do not.
Write your sentence down before Part 2. If you get it now, the rest of the first half will feel obvious, and that is exactly the right way for it to feel.
Every question a tree asks is one number against one cut. On a grid, that is a line straight up and down, or straight across. Never at an angle.
More questions means more lines, so a tree builds a staircase alongside a diagonal. With enough questions and enough shops the staircase gets close. It never becomes the diagonal, and every extra step costs a question and covers fewer shops.
Before you build one, ask: could the model already say this?
Splitting age into five bands is something a tree can already do, so it gains nothing. Subtracting one date from another is not, so it might gain everything.
Open band-05.ipynb. Every cell on this page is already in it, in this order. You run them; you do not type them. The typing is in Part 7, on your own table.
Run every cell from the top before you start. Nothing is remembered overnight, and the cells later in this step need names that were made earlier in it.
If you see NameError: name 'extras' is not defined, that is all this is. Nothing is broken.
Check that every random_state=0 on the page is in your cell too. Without it the rows are shuffled differently every run, so your numbers change every time you press Shift and Enter and no two people in the world can compare notes.
If the numbers still differ, run every cell from the top in a fresh kernel: Kernel, then Restart Kernel and Run All Cells.
If it is a while since you started JupyterLab: open a terminal, move into the course folder with cd, then type jupyter lab. Parts 6 and 7 of the setup page walk through it slowly.
The rest of this course runs on one real file. This part uses a made-up one, and you should know that before you draw any conclusions from it.
It holds 600 shops and exactly one relationship, so that you can see that relationship on its own. The bank file has no gap like this in it, for a reason that is the whole of Part 6.
It is built by tools/make_stores.py, out of real numbers from the bank file rather than out of a random generator, so that everybody gets the same 600 shops.
import pandas as pd
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split
stores = pd.read_csv("data/stores.csv")
print(stores.shape)
print(stores["big_sale"].mean().round(4))
A plain read_csv with no sep=";" this time, because this file uses commas like most of them do. The bank file was the odd one.
(600, 3), then 0.6267
600 shops, three columns, and 62.67 percent of the sales were big ones. So the lazy rule here is say yes to everybody, rather than say no as it was for the bank.
Hold that number loosely. In a moment you will keep a quarter of the shops back, and the lazy rule scores slightly differently on that quarter, the same way the bar in Step 3 was 0.878 on the test half and 0.8848 across the whole file.
print(stores.head(5))
opened_month sold_month big_sale
0 34 101 1
1 48 109 1
2 6 45 1
3 15 83 1
4 46 110 1
The same shape as your paper grid. Two months and an answer. The first five all happen to be big sales, which is what a 63 percent yes rate looks like when you only glance at five rows, and is a good reason never to judge a table by its head.
You are about to split these 600 shops the usual way and train the usual trees on the two columns as they stand.
What does a tree with two questions score on the shops it has never seen? What does a tree with no limit at all score?
You have an advantage here that you did not have in Step 3: you already know, from your paper grid, what the answer depends on. Use it. A good prediction now is worth more than a good score later.
X = stores[["opened_month", "sold_month"]]
y = stores["big_sale"]
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=0)
flat = DecisionTreeClassifier(max_depth=2, random_state=0)
flat.fit(X_train, y_train)
print(round(flat.score(X_test, y_test), 4))
0.7933
Better than saying yes to everybody, which scores 0.5667 on this particular test half. That is the 62.67 percent from two cells ago, measured on the 150 shops that were held back rather than on all 600. Better than lazy, then, and not good.
deeper = DecisionTreeClassifier(max_depth=3, random_state=0)
deeper.fit(X_train, y_train)
print(round(deeper.score(X_test, y_test), 4))
loose = DecisionTreeClassifier(random_state=0)
loose.fit(X_train, y_train)
print(round(loose.score(X_train, y_train), 4))
print(round(loose.score(X_test, y_test), 4))
print(loose.get_n_leaves())
0.8733, then 1.0, then 0.9467, then 20
Read the last three together, because you have seen this shape before. Perfect on the shops it learned from, 0.9467 on the ones it has not, and twenty endings to do it. That is Step 3's memorising tree, in a new place, and this time you can see exactly what it is memorising: a staircase, climbing beside a diagonal it cannot draw.
Now the thing you worked out on paper.
stores["months_open"] = stores["sold_month"] - stores["opened_month"]
print(stores["months_open"].min(), stores["months_open"].max())
print(stores.head(3))
One line. Subtract one column from another and pandas does it for all 600 rows at once, the same way multiplying a column by a hundred did in Step 2.
0 71, then a table with a fourth column on the end holding 67, 61 and 39.
Now a tree with one question, on all three columns, including the new one.
One question. Write down what it scores on the shops it has never seen.
X2 = stores[["opened_month", "sold_month", "months_open"]]
X2_train, X2_test, y_train, y_test = train_test_split(
X2, y, test_size=0.25, random_state=0)
one_question = DecisionTreeClassifier(max_depth=1, random_state=0)
one_question.fit(X2_train, y_train)
print(round(one_question.score(X2_train, y_train), 4))
print(round(one_question.score(X2_test, y_test), 4))
1.0, then 1.0
Every shop it learned from, and every shop it never saw. One question.
In Step 3 a perfect score was a warning. Here it is not, and there are two separate reasons, and you should have both.
This table was built so that a big sale is exactly a gap of more than 24. That is in tools/make_stores.py, and you can read it. There is no noise in it at all, so a model that can say the rule gets every row, and nothing you ever do on real data will be this clean.
That is on purpose. You are being shown a mechanism, not a result. Part 5 is the same idea meeting a real file, and it goes very differently.
The second reason is that nothing leaked, and here is the test, which is narrower than you might guess.
A column is safe when each row is worked out from that row alone, out of things you would have for a new person on the day, using arithmetic you can read.
All three parts matter. Row by row is the one people miss: the moment a column is worked out across the whole table, the other rows have had a say in it, and some of those rows are your test set. The marked question at the end of this step is exactly that mistake.
months_open passes on all three. Every shop's gap is its own two numbers subtracted, and no other shop is consulted.
print(export_text(one_question, feature_names=list(X2.columns)))
|--- months_open <= 24.50
| |--- class: 0
|--- months_open > 24.50
| |--- class: 1
Two years. A shop that has been open more than two years makes big sales, and one that has not does not.
The tree threw away both of the columns it was given and used only the one you built. It was never going to find that rule in one question, and it found it immediately once somebody wrote it down.
You did not add any information. Both numbers were already there, in every row, in front of the model, from the beginning.
What you added was a way of saying it.
Look at what the unlimited tree was doing instead. This is the first four questions of the twenty-ending tree from Part 2.
print(export_text(loose, feature_names=list(X.columns), max_depth=2))
max_depth=2 here is a limit on the printing, not on the tree. The tree is unchanged; you are just asking to see the top of it.
|--- sold_month <= 51.50
| |--- opened_month <= 17.50
| | |--- sold_month <= 36.00
and six questions in all at this printing depth, with truncated branch of depth lines where it has stopped showing you.
Read the first three cuts as lines on your paper grid. One across at 51.5, one up and down at 17.5, then another across at 36. Across, up and down, across, and on it goes: twenty boxes in all, each a small rectangle with a single answer in it.
That is a staircase. It follows the diagonal, roughly, and every extra step costs a question and covers fewer shops. Sorted by size, the twenty boxes hold 1, 1, 2, 2, 4, 5 shops and upwards. Two of them hold a single shop each. Step 2 put it this way: a rule built around one person is not a rule, it is a note about that person.
It is not that a tree cannot get there. Given enough questions and enough rows it will build a staircase against almost any boundary, and Part 2 watched it climb to 0.9467 doing exactly that.
What is true is narrower and more useful:
So the column does not tell the model anything new. It makes something cheap that was expensive, and cheap is what you are buying.
A tree is unusually bad at gaps because it looks at one column at a time. The model in Step 9 is not: it adds the columns up with weights it works out for itself, and a gap is a subtraction, which is an addition with one weight turned negative. On this shop table that model gets every shop right from the two raw columns, without your help.
Gold at the bottom asks you to run it. The lesson is not that trees are poor. It is that "the model cannot express this" is a question about a particular model, and you have to know which one you are holding.
You cannot find months_open by staring at a table. You find it by knowing that shops take time to establish, which is something a shopkeeper knows and a spreadsheet does not.
Every step has asked you to show your work to somebody who does not code. This is the step where that stops being a courtesy and starts being the method.
Ask which kind you are about to build before you write any code. It will save you most of the time people waste on this.
Back to the real file, and to the 47-column table you built in Step 4.
Here are four columns a sensible person would build for it. Each one is a real idea, and one of them is yours.
peak_month: one for March, September, October and December, zero otherwise. This is your Step 1 rule, the four months you found by hand, written as a column.contacts_total: this campaign's calls plus the previous campaign's. How much this person has been bothered altogether.balance_per_year: money in the bank divided by age. A rough measure of how well somebody is doing for their age, which a bank might care about.contacted_before: one if this person was ever called in an earlier campaign, zero otherwise. The flag from Step 4's Part 8.One is a sum and one is a ratio, which Part 4 says can pay. Two are things one question could arguably already say. Before you run anything, write down which is which, and which single one you would bet on.
bank = pd.read_csv("data/bank.csv", sep=";")
y_bank = (bank["y"] == "yes").astype(int)
plain = ["age", "balance", "day", "campaign", "pdays", "previous"]
many = ["job", "marital", "education", "contact", "month", "poutcome"]
ready = bank[plain].copy()
for name in ["default", "housing", "loan"]:
ready[name] = (bank[name] == "yes").astype(int)
ready = pd.concat([ready, pd.get_dummies(bank[many], dtype=int)], axis=1)
print(ready.shape)
Step 4's Part 4, in one cell. Every piece of it is from that step: .copy(), the loop turning three yes-or-no columns into ones and zeros, get_dummies making 38 boxes, and concat joining them on sideways.
You should be. Step 4's Part 7 showed you that get_dummies breaks the moment you send new rows through a trained model, and had you write down when you would still use it.
This is that case. Nothing here is ever asked about a row from outside this table: every score comes from splitting these same 4,521 people. For looking at a table and comparing models on it, get_dummies is fine and shorter. The encoder is for the day you use the model.
If that made you stop and check, you did the right thing, and it is exactly what Step 4 was for.
(4521, 47)
extras = ready.copy()
extras["peak_month"] = bank["month"].isin(["mar", "sep", "oct", "dec"]).astype(int)
extras["contacts_total"] = bank["campaign"] + bank["previous"]
extras["balance_per_year"] = (bank["balance"] / bank["age"]).round(2)
extras["contacted_before"] = (bank["pdays"] != -1).astype(int)
print(extras.shape)
print(extras["peak_month"].sum())
.isin([...]) is from Step 1, where you used it to write the wider rule. The division and the addition work a column at a time, like the subtraction in Part 3.
(4521, 51), then 201
Only 201 of the 4,521 calls were made in one of your four months. That is worth noticing now rather than in Part 6: the column you were most confident about is switched on for one call in twenty-two.
You are about to run Step 4's thirty splits on the 47-column table and the 51-column one, side by side.
What does each average? And how many of the thirty does the wider one win?
base_scores = []
extra_scores = []
wins = 0
for seed in range(30):
a, b, c, d = train_test_split(ready, y_bank, test_size=0.25, random_state=seed)
plain_tree = DecisionTreeClassifier(max_depth=3, random_state=0)
plain_tree.fit(a, c)
base_score = plain_tree.score(b, d)
a, b, c, d = train_test_split(extras, y_bank, test_size=0.25, random_state=seed)
extra_tree = DecisionTreeClassifier(max_depth=3, random_state=0)
extra_tree.fit(a, c)
extra_score = extra_tree.score(b, d)
base_scores.append(base_score)
extra_scores.append(extra_score)
if extra_score > base_score:
wins = wins + 1
print(round(sum(base_scores) / 30, 4))
print(round(sum(extra_scores) / 30, 4))
print(wins)
The same loop as Step 4's Part 6, with one thing changed: both models get three questions, so the only difference between them is the four new columns.
The four short names come out of train_test_split in its fixed order, and it is worth saying them out loud once more: a is the training columns, b the test columns, c the training answers, d the test answers. Short names here because they are used twice and thrown away.
0.8902, then 0.8882, then 5
Four thought-out columns, built from a real finding and two honest ratios, and the average went down. They won 5 of the 30.
In Step 4 the one split you were handed happened to flatter the new columns. Not this time. I ran random_state=0 on its own, and the wider table scores 0.8912, catching 23 of the people who would say yes for 8 wasted calls, against Step 4's 0.8983 catching 28 for 5. You do not need to run it; the cell above already told you the important thing.
Worse on the average and worse on the split. There is nothing to argue about, which is a comfortable place to be and does not happen often.
I also ran them one at a time, because four columns failing together could hide one that helps. The averages, against the 0.8902 to beat: peak_month 0.8882, contacts_total 0.8902, balance_per_year 0.8901, contacted_before 0.8902. Bronze at the bottom asks you to reproduce that.
One of the four is doing all the damage, and it is the one you were most sure of.
This is the part that is worth the step, and it starts with a question you can put to the model directly rather than guessing at: did it even use the column?
new_names = ["peak_month", "contacts_total",
"balance_per_year", "contacted_before"]
used = pd.Series(0, index=new_names)
for seed in range(30):
a, b, c, d = train_test_split(extras, y_bank, test_size=0.25, random_state=seed)
tree = DecisionTreeClassifier(max_depth=3, random_state=0)
tree.fit(a, c)
weight = pd.Series(tree.feature_importances_, index=extras.columns)
used = used + (weight[new_names] > 0)
print(used.to_string())
Three new things, all small.
pd.Series(0, index=new_names) makes four counters with names on them, all starting at zero. You met pd.Series in Step 1, where you made a column of False the same way.tree.feature_importances_ is one number per column, saying how much work that column did. The trailing underscore means the tree worked it out, like statistics_ in Step 4. Step 6 is about what the size of these numbers means. Today you only care whether one is zero, because zero means the tree never asked about that column at all.weight[new_names] > 0 gives four True or False answers, and adding those to the counters turns each True into a 1, the way .astype(int) has all course.peak_month 29
contacts_total 1
balance_per_year 5
contacted_before 0
Not one of them is what I expected when I wrote this page, and the first line is the opposite of it. Take them in turn.
This is the one I got wrong when I first wrote this page, and it is more interesting than what I expected.
I assumed the tree would leave it alone, because it already had twelve month columns and had picked October out of them. It does the opposite. It reaches for peak_month in 29 of the 30 splits, and it reaches high, because lumping four good months together makes a bigger, cleaner-looking split than any single month can.
Print the tree at random_state=0 and you can watch the damage. In Step 4 the second question was month_oct, and under it sat the leaf that caught six people in October after the sixteenth. With peak_month available, the tree asks that instead, and then both leaves underneath say no. The leaf that was catching people is gone, and the catches drop from 28 to 23.
So the coarse column won the argument and then had nothing useful to say. That is worth remembering: a tree takes the biggest split it can see right now, and the biggest split now is not always the best three questions later. Nobody warns you about this, and it is why Step 12 exists.
One correction to something you may believe from Step 1, while you are here. October is not special. In October 37 of 80 calls ended in a yes. In December it was 9 of 20 and in March 21 of 49, which are the same kind of rate. October wins because it holds 80 calls to December's 20, so the tree can build a bigger, steadier leaf out of it. Four months are genuinely good, and the tree picks the one it can lean on hardest.
These two are the kind Part 4 said to look for: a sum and a ratio, neither of which one question can say.
The tree barely touched them: contacts_total once out of thirty and balance_per_year five times, and the average did not move.
The reason is simpler than the shape. Whether somebody opens a savings account has very little to do with their total number of phone calls, or with their money divided by their age. The arithmetic was legal. The idea was wrong.
A feature is a guess about how the world works, written as a column. Getting the shape right does not make the guess right, and no amount of care with the shape will rescue a guess that is not true.
Zero out of thirty. That bottom line is what a column with nothing whatever to add looks like, and Step 4 half told you already: it said contacted_before "never won a place in the top three questions". Now you have watched it lose thirty times.
It is the same 816 people, three times. pdays is -1 for exactly the other 3,705, so a cut anywhere between -1 and 1 makes the same two groups. poutcome_unknown, one of your 38 boxes, marks exactly the same 3,705, not almost. And previous is 0 for exactly those people too, so previous <= 0.5 is a third way to say it.
Three columns already said it in one question each. You built a fourth.
The gap that would matter for this data is time. How long since this campaign began, how long between the last campaign and this one, whether this call came before or after the neighbouring bank went under.
You cannot build any of it. The file has a day and a month and no year, and you found that out in Step 0, in Part 6, when you went looking for a year column and there wasn't one.
That is what a missing column actually looks like in real work. Not a gap you can fill with arithmetic. A question the table cannot answer, which no amount of cleverness will fix, and which somebody has to be told.
Look at pdays again. Days since this person was last contacted. That is not something the bank's computer recorded on the day. It is a subtraction between two dates that somebody did before they gave you the file, which makes it the one hand-built gap in the whole table.
Hold it to the same standard as your four.
without_pdays = ready.drop(columns=["pdays"])
scores = []
for seed in range(30):
a, b, c, d = train_test_split(without_pdays, y_bank, test_size=0.25, random_state=seed)
tree = DecisionTreeClassifier(max_depth=3, random_state=0)
tree.fit(a, c)
scores.append(tree.score(b, d))
print(round(sum(scores) / 30, 4))
.drop(columns=["pdays"]) hands back a copy of the table with that column missing. It leaves ready alone, the way .copy() does.
0.891
Against 0.8902 with it. Very slightly better without.
So somebody did the job you are learning, on the right column, for the right reason, and on this file it did not pay either. That is the honest shape of this work, and it is why the measuring matters more than the ideas.
Three lines. The two averages and the win count. Which of the four columns you would defend in a review even though it lost. And one sentence naming a column this table would need that cannot be built from what is in it.
No code here and no answer at the end. Your own table, the one you signed up in Step 0 and got into shape in Step 4.
mine, not bank. my_extras, not extras. If you reuse the names above, the cells you already ran will still work and print numbers about the wrong table.
Build the columns your table needs, decide which of them are worth anything, and prove it with thirty splits rather than one.
Start there. The gap between two dates is the feature your model is least able to find for itself, so it is the first one worth testing. Days between, months between, and whether one came before the other. Whether it pays is a separate question, and Part 6 is what that looks like when the answer is no.
Step 17 is where dates get their own step, and everything you write down now feeds it.
Be suspicious in the direction Step 3 taught you. Ask the Part 6 question of it: would you have this, for a new row, at the moment you have to answer?
A feature built out of columns you will not have on the day is still leakage, however honest the arithmetic that made it.
If the table is about work you do, you are the person this step keeps telling you to go and find. Sit down with your three ideas before you build them and answer two questions in writing: what would I look at, and what do I wish somebody had recorded?
The second question is the one that pays, and it is the fifth item on the list above. Domain knowledge is not a thing other people have. It is what you know about this data that the table does not say.
Ask them both questions before you build anything, and write their answer down word for word rather than in your own words. The wording is where the column hides.
Add to the README you started in Step 0.
Find a person who does not code and draw them the grid from Part 0, with the dots and the diagonal.
Then tell them that the computer is only allowed to draw lines straight up and straight across. If they say "so you had to do the diagonal for it", you have explained the whole step in twenty seconds, and so have they.
Six questions. Answer them in writing before you open the answers.
One counts more than the other five. You can get five right and still not pass this step.
Two of the answers below offer half marks. Half is not a pass on the marked one: if you got half there, read the answer, then come back tomorrow and write it out again from memory before you go on. Half on any of the other five is a pass.
age with five age bands and the score does not move at all. A colleague says the bands must be wrong. What do you say?contacts_total is a sum of two columns, which is exactly the shape Part 4 said to look for, and it did nothing. Why not?A colleague has a table of hospital visits, one row per visit. It holds admitted_on and discharged_on as dates, the ward, the patient's age, and twenty other columns. She is predicting whether a patient will be readmitted within a month. Her tree scores badly and she asks you which of three new columns to build.
For each of the three, say whether it can help her and why. One of them has almost nothing to offer a tree, and one of them is not safe to build the way she has described it.
Length of stay: build this one. It is the gap between two columns, which costs a tree a staircase and costs you one subtraction, and it is the thing a nurse would tell you matters. It is the shop table with a different hat on. It is also safe: worked out row by row, from two things she has the moment a patient leaves.
Age bands: very little to offer. A tree can already cut age anywhere it likes, so the bands say nothing it could not say. The one case where they help is if her tree is short of questions and a band saves it one, and that is a reason to give the tree another question rather than to build the column. Full marks if you said that; half if you said they cannot help at all.
The ward readmission rate: not like that. It is built from the answers of other rows, and worked out across the whole table, so rows in her test set have helped build a column her training rows learn from. That is Step 4's scaler with a new face, and it is worse than the scaler, because a scaler learns a range and this learns the answer.
How badly does it flatter her? Nobody can tell her without measuring, and how much depends on how many visits each ward has: a ward with 40 visits leaks a lot into each row, a ward with 4,000 leaks almost nothing, and the same code gives both. If she wants the column, it is worked out on the training rows only, applied to the rest, and measured the Part 5 way over thirty splits before she believes a word of it.
Mark yourself right only if you said the third one is built out of answers and that it is worked out across the whole table. Either fault alone is half.
Go back to Part 4 and read the two kinds of new column again, then read Step 4's Part 9 about the scaler.
Those three columns are the three mistakes people actually make with features, and you will be handed all three within a month of starting work.
Run the four bank columns one at a time rather than all together: four loops of thirty splits, four averages against the 0.8902 to beat.
You should get 0.8882, 0.8902, 0.8901 and 0.8902. Then answer this: if one of them had come out at 0.8915, what would you have needed to do before believing it?
Take months_open away again and find out what the tree buys with extra rows instead. Split the shops the usual way first, so that you have 450 to train on and 150 held back. Then train an unlimited tree on the first 100 of the training rows, then 200, then 300, then all 450, scoring every one of them on the same 150.
Take the training rows out of the training half, not out of the whole table. Slicing the first 100 rows off the full 600 puts some of your held-back shops into the training set, which is Step 3's duplicate trap in a new coat, and your last score would be a memory test.
Write down the four scores and the four leaf counts. It climbs and then flattens out somewhere below 1.0, and the flattening is the point: one column got there with one question and 450 rows. Say in one line what a feature buys you, in the currency of rows.
Go back to the shop table, the two raw columns, no months_open, and train a different model on them:
from sklearn.linear_model import LogisticRegression
straight = LogisticRegression()
straight.fit(X_train, y_train)
print(round(straight.score(X_test, y_test), 4))
print(straight.coef_.round(3))
You have not met this model and Step 9 is where it belongs, so run it, do not study it. It scores 1.0 on shops it has never seen, from the two columns a tree needed twenty boxes for, and the two numbers it prints are close to equal and opposite.
Write the paragraph explaining why. What does one weight of about -1.7 and one of about +1.7 do to a pair of columns, and what does that have to do with the column you built by hand in Part 3? Then say what it costs: name one thing the tree could do on this table that this model cannot.
All six, or it is a not yet.
Keep the sentence you wrote in Part 0, and the shop grid. Step 6 asks which of your columns are pulling their weight, and it starts by handing you a way to ask the model directly instead of guessing.