Your model will hand you a ranked list of its columns if you ask. You add a column that is nothing but the row number, and the list does not put it anywhere near last. Then you learn the two questions that catch that, and the one thing no ranking can ever tell you.
Step 5 ended with you asking the model a question directly for the first time: did you even use this column? You used feature_importances_ and only looked at whether the number was zero.
This step is about the size of those numbers, and it is mostly about how badly you can be fooled by them.
An importance table tells you what a model leaned on. It does not tell you what matters, it does not tell you what would help, and it cannot tell you that a column is cheating.
You add one column to your table. It holds the row number: 0 for the first person, 1 for the second, and so on to 4,520. It is not about anybody. It cannot possibly mean anything.
You train a tree six questions deep and ask for the ranking. Your row number lands high in it, above columns you worked hard for in Step 4.
Then in Part 4 you ask a different question, of the people the model never saw, and get a completely different answer about the same column. Both are correct. They are answers to different questions, and almost nobody who quotes the first one knows that.
The page does not tell you either position in advance. You are asked to write down a guess first, three times, and being wrong is the point.
duration back in and watch both rankings fall for it at once.Parts 0 and 1 are twenty-five minutes on paper. You do not open the notebook until Part 2.
Parts 3 and 4 are one idea in two halves and want an hour together. If you have less than an hour left, stop at the end of Part 2 and start fresh with Part 3 another day, rather than stopping between them. The other good places to stop are the end of Part 5, and before Part 7, which is a sitting of its own.
Paper and pen. Laptop shut. Copy the five clues onto your paper first, so the table below is something you can work on.
A coat shop wants to know which visitors will buy a coat, so that it can decide who to send an offer to on Friday morning. Here are five things it could know about a visitor.
| Clue | What you would pay |
|---|---|
| Whether it rained that week | |
| How much they spent in the shop last winter | |
| How many minutes they spent in the shop on the day | |
| Their coat size | |
| Whether they came in with somebody else |
You have one hundred pounds to spend on clues. Split it between the five, in whatever way you like, giving most to the one you would least want to lose. Write the five numbers in the second column. They must add up to a hundred.
This takes longer than you expect and the arguing with yourself is the exercise. Do it before you read on.
It is Friday morning. Nobody has come into the shop yet, and you are choosing who to send the offer to.
Go down your five clues and ask of each one: would I have this, for a person I am about to choose, right now?
Cross out the ones you would not. Then add up the money you put on the ones you crossed out.
One of the five cannot survive that question. How many minutes somebody spent in the shop is a fact about a visit that has not happened. On Friday morning it does not exist for anybody you care about.
It is also, if you are honest, probably one of the two you paid most for, because it is the best clue on the list. Somebody who spends forty minutes in a coat shop is buying a coat.
You have met this before. In Step 3 it was called duration, the length of a phone call that had not happened yet.
Write one line: what was different about the way you found it this time, when you were pricing clues rather than reading a score?
Look at "whether it rained that week". It is the same for every visitor on the same day, so it can never tell two of Friday's visitors apart. A clue that does not vary between the people you are choosing between is worth nothing for choosing between them, however much it explains about the shop's takings overall.
And look at "their coat size" next to "how much they spent last winter". If everybody who spent a lot last winter also happens to be recorded with a coat size, and the people who spent nothing are not, then coat size is quietly saying the same thing as last winter’s spending, and taking it away might cost you nothing at all.
Keep your list of five with the money written on it. Part 6 comes back to it.
feature_importances_. Worked out from the training rows, while the tree was being built.There are three ways a column can be at the top of an importance table: it is genuinely useful, it is noise the model was able to carve up, or it is cheating.
The table looks identical in all three cases.
Open band-06.ipynb. Every cell on this page is already in it, in this order. You run them; you do not type them. The typing is in Part 7, on your own table.
Open a terminal, move into the course folder with cd, then type jupyter lab. It does not stay running between days.
Run every cell from the top before you start. Nothing is remembered overnight, and the cells in Parts 5 and 6 need names that were made back in this one.
If you see NameError: name 'ready' is not defined, that is all this is. Nothing is broken.
Check that every random_state=0 on the page is in your cell too. Without it the rows are shuffled differently every run, so your numbers change every time you press Shift and Enter and no two people in the world can compare notes.
If the numbers still differ, run every cell from the top in a fresh kernel: Kernel, then Restart Kernel and Run All Cells.
The first cell is Step 4's table and Step 3's split, in one go. Nothing here is new except the last import, which you will need in Part 4.
import pandas as pd
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.inspection import permutation_importance
df = pd.read_csv("data/bank.csv", sep=";")
y = (df["y"] == "yes").astype(int)
plain = ["age", "balance", "day", "campaign", "pdays", "previous"]
many = ["job", "marital", "education", "contact", "month", "poutcome"]
ready = df[plain].copy()
for name in ["default", "housing", "loan"]:
ready[name] = (df[name] == "yes").astype(int)
ready = pd.concat([ready, pd.get_dummies(df[many], dtype=int)], axis=1)
X_train, X_test, y_train, y_test = train_test_split(
ready, y, test_size=0.25, random_state=0)
print(ready.shape)
(4521, 47)
Now the tree from Step 4, and its ranking.
small = DecisionTreeClassifier(max_depth=3, random_state=0)
small.fit(X_train, y_train)
weight = pd.Series(small.feature_importances_, index=X_train.columns)
print(round(weight.sum(), 4))
print((weight > 0).sum())
print(weight.sort_values(ascending=False).head(6).round(4).to_string())
index=X_train.columns writes the column names down the side of the list, so weight["age"] reads off one column's number by name. .sort_values(ascending=False) puts the biggest first. .head(6) keeps the top six, the way it kept the top rows of a table in Step 0. .to_string() prints the names beside the numbers without the extra line pandas adds about the kind of thing it is.
1.0
6
poutcome_success 0.5974
month_oct 0.1542
age 0.1104
day 0.0547
balance 0.0460
education_tertiary 0.0373
Three things in that, and the first two are the ones people miss.
poutcome_success takes 60 percent of the total, which is the same column your Step 1 rule was made of, arriving again with a number on it.Now the same thing with no limit on the depth.
big = DecisionTreeClassifier(random_state=0)
big.fit(X_train, y_train)
big_weight = pd.Series(big.feature_importances_, index=X_train.columns)
print((big_weight > 0).sum())
print(round(big.score(X_test, y_test), 4))
47, then 0.8073
Every single column now has an importance above zero. All 47 of them matter. Or so the table says.
And the model they all matter to scores 0.8073 on people it has never seen, which is well below the 0.878 you get by saying no to everybody. Step 3's memorising tree, with a ranking attached.
The number of columns with a non-zero importance went from 6 to 47 because the model got deeper, not because the columns got better.
Importance is a fact about your model. It is not a fact about your data, and it is certainly not a fact about the world.
Time to find out how much that ranking is worth.
You are going to add one column holding nothing but the row number. The first person gets 0, the second gets 1, and so on to 4,520. The rows of this file are a random tenth of a bigger file, so the order means nothing whatever: it is not time, it is not the order the calls were made, it is not anything.
You will train a tree six questions deep on 48 columns, one of which is that row number, and rank them.
Where does the row number come? Write a position from 1 to 48 before you run it. Most people write 48.
noisy = ready.copy()
noisy["row_number"] = range(len(noisy))
N_train, N_test, y_train, y_test = train_test_split(
noisy, y, test_size=0.25, random_state=0)
deep = DecisionTreeClassifier(max_depth=6, random_state=0)
deep.fit(N_train, y_train)
deep_weight = pd.Series(deep.feature_importances_, index=N_train.columns)
print(deep_weight.sort_values(ascending=False).head(5).round(4).to_string())
range(len(noisy)) counts from 0 up to one less than the number of rows, which is exactly the row numbers. Six questions rather than three, because a small tree does not have room to be fooled, and you will see that in a moment.
poutcome_success 0.3168
age 0.1086
row_number 0.0946
pdays 0.0933
month_oct 0.0914
print(round(deep_weight["row_number"], 4))
print(round(deep.score(N_test, y_test), 4))
0.0946, then 0.8806
Third out of forty-eight. Ahead of pdays. Ahead of the October column that you found by hand in Step 1 and watched the model rediscover in Step 4.
Third is this split, not a law. Over thirty different splits it lands between second and eleventh, and it never once falls out of the top quarter of the table. The position moves; the fact that a column made of nothing sits near the top does not.
The column is the row number. There is nothing in it. If you put that ranking in front of a manager, one of its top lines would be a lie, and you would have no idea.
Built-in importance is worked out while the tree is being built, on the training rows, and it counts how much each question tidied up the heap underneath it.
Now think about what the row number offers a tree. It has 3,390 different values in the training half, so there are 3,389 places to cut it. Most other columns offer far fewer: month_oct offers exactly one cut, because it is a one or a zero. The exception proves the rule, and it is worth knowing: balance has 1,927 different values, and it floats up this table too.
Give a tree three thousand places to cut and some of them will separate the training rows a little, by luck. The tree takes the best of those, and it looks like work done, and it is counted as work done.
Built-in importance favours columns with many different values, whether or not those values mean anything.
So a customer number, an account reference, a postcode, a timestamp, a price to two decimal places: all of them will float up your ranking for no better reason than that they are finely divided. If the top of your table is an id column, one whose only job is to name a row rather than say anything about the person in it, that is not a discovery. It is this.
One more thing before you leave it. Give the same 48 columns to a tree with three questions instead of six.
shallow = DecisionTreeClassifier(max_depth=3, random_state=0)
shallow.fit(N_train, y_train)
shallow_weight = pd.Series(shallow.feature_importances_, index=N_train.columns)
print(round(shallow_weight["row_number"], 4))
print((shallow_weight > 0).sum())
0.0, then 6
Exactly zero, tied at the bottom with 41 other columns. The row number was offered to this tree too, with the same three thousand places to cut, and it never got used.
This tree had no room to be fooled. It only had three questions and it spent them on the columns that were worth it.
Do not make that a promise, though. Run the same three questions on thirty different splits and the row number picks up a real number on seven of them, once as high as fourth. Three questions is thin protection, not armour. Depth does not only cost you honesty in the score, which is Step 3. It costs you honesty in the ranking too, and shallowness only makes it less likely.
There is a second way to ask which columns matter, and it is slower, better, and almost never the one people use.
Take the model you already trained. Take the 1,131 people it has never seen. Now pick one column and shuffle it, so that everybody has somebody else's value for that column and the column still looks perfectly normal. Score the model again.
If the score falls, the model was using that column for something real. If the score does not move, it was not.
Same model, same 48 columns. Shuffle each column in turn, on the people the model has never seen.
Where does row_number come now? Write the position down before you run it.
shuffled = permutation_importance(deep, N_test, y_test,
n_repeats=10, random_state=0)
honest = pd.Series(shuffled.importances_mean, index=N_test.columns)
print(honest.sort_values(ascending=False).head(5).round(4).to_string())
permutation_importance does the shuffling for you, once per column. You hand it the trained model and the rows to test it on, and it hands back a result you read the numbers off. n_repeats=10 shuffles each column ten times and averages, because one shuffle is one go, the way one split was one go in Step 4. importances_mean is the average drop in score for each column.
poutcome_success 0.0134
month_oct 0.0068
day 0.0035
housing 0.0020
education_tertiary 0.0015
print(round(honest["row_number"], 4))
order = honest.sort_values(ascending=False)
print(list(order.index).index("row_number") + 1)
The second line finds where in the sorted list that name sits and counts from 1, because you predicted a position and the ranking prints numbers.
-0.0019, then 47
Below zero means the model got better when that column was scrambled. Which makes sense: everything the tree learned from the row number was fitted to the training people and was noise for everybody else, so wrecking it helped.
Two rankings. Same model, same columns, same afternoon. One puts the row number third and one puts it forty-seventh.
Both are correct answers to their own question.
The first asks: how much did this column shape the model I built? The row number shaped it a lot. That is true, and it is why the model is worse than it looks.
The second asks: how much does this model lose when this column is scrambled, on people I have not met? Nothing at all. Also true, and it is much closer to the question you were actually asking.
Closer, not the same. It is what scrambling costs the model you have. It is not what you would lose by dropping the column and training again, and Part 5 is about the gap between those two.
Notice how much smaller the second set of numbers is. The best column in the table costs 0.0134 when it is scrambled, which is just over one point of score. Built-in importance is a share of one and always adds to 1; permutation importance is a real amount of score, and real amounts are usually small. Never compare a number from one table with a number from the other.
Permutation importance re-scores the model once per column per repeat. That is 48 columns times 10 repeats, and it took a moment. On a big model with a thousand columns you would feel it.
And it has a real weakness, which the next part is about: shuffle one of two columns that say the same thing, and the model just reads the other one, so both look worthless.
Step 5 found something you can use here. Four columns in this file say the same thing about the same 816 people: pdays, previous, poutcome_unknown, and the contacted_before flag you built. They are redundant, in the Part 1 sense.
Put the flag back and ask the ranking about all four.
dup = ready.copy()
dup["contacted_before"] = (df["pdays"] != -1).astype(int)
D_train, D_test, y_train, y_test = train_test_split(
dup, y, test_size=0.25, random_state=0)
dup_tree = DecisionTreeClassifier(max_depth=6, random_state=0)
dup_tree.fit(D_train, y_train)
dup_weight = pd.Series(dup_tree.feature_importances_, index=D_train.columns)
print(dup_weight[["pdays", "previous",
"poutcome_unknown", "contacted_before"]].round(4).to_string())
Handing a list of names inside the square brackets picks out those rows of the ranking, in that order.
pdays 0.1148
previous 0.0000
poutcome_unknown 0.0000
contacted_before 0.0000
One column takes everything and three take nothing.
There is nothing to choose between them. The tree met pdays first, split on it, and by the time it looked at the other three there was nothing left for them to explain. A ranking cannot show you a tie, so it shows you a landslide.
Watch what happens when the winner leaves.
gone = dup.drop(columns=["pdays"])
G_train, G_test, y_train, y_test = train_test_split(
gone, y, test_size=0.25, random_state=0)
without = DecisionTreeClassifier(max_depth=6, random_state=0)
without.fit(G_train, y_train)
without_weight = pd.Series(without.feature_importances_, index=G_train.columns)
print(without_weight[["previous", "poutcome_unknown",
"contacted_before"]].round(4).to_string())
print(round(dup_tree.score(D_test, y_test), 4))
print(round(without.score(G_test, y_test), 4))
previous 0.0183
poutcome_unknown 0.0000
contacted_before 0.0000
then 0.8859, then 0.8842
previous was worth zero and is now worth something. It did not change. The redundant column that beat it left.
And the score barely moved, because the information never went anywhere. You deleted the column that took 11 percent of the credit and lost less than two thousandths of a point.
That was one split, and you know from Step 4 what one split is worth. Here is the same question asked thirty times.
with_it, without_it = [], []
for seed in range(30):
a, b, c, d = train_test_split(dup, y, test_size=0.25, random_state=seed)
tree = DecisionTreeClassifier(max_depth=6, random_state=0).fit(a, c)
with_it.append(tree.score(b, d))
a, b, c, d = train_test_split(gone, y, test_size=0.25, random_state=seed)
tree = DecisionTreeClassifier(max_depth=6, random_state=0).fit(a, c)
without_it.append(tree.score(b, d))
print(round(sum(with_it) / 30, 4))
print(round(sum(without_it) / 30, 4))
Both splits use the same seed each time round, so the two models are always being asked about the same people. That is the only fair way to compare them.
0.8844, then 0.8845
Averaged over thirty different sets of people, the table without pdays is very slightly ahead. Not meaningfully ahead: a ten-thousandth of a point is nothing. But the column that took 11 percent of the ranking is worth, as far as anyone can measure, exactly nothing at all.
This is the loop you will need in Part 7, and it is the only tool on this page that answers the question people think a ranking answers.
Deleting a zero. "It scored zero, so I dropped it" is how people delete the backup for a column they are about to lose. If your top column is one day recorded differently, or stops arriving, the zeros are what you have left.
Believing a landslide. A column at the top with three redundant columns underneath it is not four times as important as anything. It won a race by arriving first.
The fix is the loop you just ran: take the column out, run thirty splits, and see what actually happens to the score. It is slower than reading a ranking, and it is the only one of the three that answers the question you meant.
One warning about it, so you do not swing too far the other way. A column can be redundant, cost nothing to delete, and still be related to the answer. The score not moving tells you the model does not need it. It does not tell you the column has nothing to do with anything.
One last one, and it is the reason this step cannot be the last word.
Put duration back in. You met it in Step 3: the length of the phone call, which does not exist until the call has happened.
Where does duration come in the built-in ranking, and where in the permutation ranking?
You know it is leakage. The question is whether either ranking knows.
leaky = ready.copy()
leaky["duration"] = df["duration"]
L_train, L_test, y_train, y_test = train_test_split(
leaky, y, test_size=0.25, random_state=0)
cheat = DecisionTreeClassifier(max_depth=6, random_state=0)
cheat.fit(L_train, y_train)
print(round(cheat.score(L_test, y_test), 4))
built_in = pd.Series(cheat.feature_importances_, index=L_train.columns)
print(built_in.sort_values(ascending=False).head(3).round(4).to_string())
0.9036, then
duration 0.4587
poutcome_success 0.1816
age 0.0740
shuffled_leaky = permutation_importance(cheat, L_test, y_test,
n_repeats=10, random_state=0)
honest_leaky = pd.Series(shuffled_leaky.importances_mean, index=L_test.columns)
print(honest_leaky.sort_values(ascending=False).head(3).round(4).to_string())
duration 0.0580
poutcome_success 0.0213
contact_cellular 0.0118
Top of both. By a distance, in both. Nearly half of the built-in table, and getting on for three times the next column in the permutation table.
The score is 0.9036, as high as anything this course has produced. Step 3 got a number just like it, in the same way. Every check you have agrees this is the best column you own.
Permutation importance did not catch it, and it never will, because duration genuinely does predict the answer on rows the model has never seen. Every one of those rows is a call that already happened.
A ranking measures how well a column predicts. Leakage is a column that predicts well and will not be there. No amount of measuring the first thing will ever tell you about the second.
The only thing that catches it is the Part 0 question, asked by a person: would I have this, for somebody I am about to choose, at the moment I have to choose?
Which gives you what a ranking is actually for. The ranking is not the check. The ranking is the shortlist. Every column near the top of it gets that question asked out loud, and the higher it is, the harder you ask.
So you can answer the question after all, when somebody asks which columns matter most in your model. Not with the built-in table, and not with a shrug. You say which of the two rankings you are quoting, what its numbers mean, and which of the columns near the top you checked by hand and how. That answer is three sentences long and you now have all three.
No code here and no answer at the end. Your own table, the one you signed up in Step 0 and have carried through five steps. Give it a sitting of its own; it is the longest part in the step.
mine, not df. my_weight, not weight. If you reuse the names above, the cells you already ran will still work and print a ranking of the wrong table.
Produce one importance table for your own model that you would be willing to show somebody, and delete one column with the reason written next to it.
Go and look at where it comes in the built-in ranking before you drop it. This is the single most common version of Part 3 in real work, and seeing it happen on your own data is worth more than reading about it here.
Then drop it. An id is a name, not a fact about anybody.
Drop n_repeats to 3 and say in your README that you did. Fewer repeats means a noisier answer, and the honest thing is to say which you took rather than to quote it as if it were the ten.
You are the person who knows your subject. Put the README away, and tomorrow read your top five out loud, in order, before you look at the numbers again. You are listening for your own surprise: "that one is at the top?" is the beginning of finding either a leak or something real that nobody had noticed.
Read the same five to them, in order, without saying what the numbers are. Their surprise is worth more than yours, because they do not know what the model did and you do.
Add to the README you started in Step 0.
Find a person who does not code. Tell them you added a column of row numbers to a model and it came third out of forty-eight.
If they laugh, they have understood it. That reaction is the correct one, and most people who read importance tables for a living have never had it.
If they do not laugh, tell them the column was 0, 1, 2, 3 counting down the page and nothing else, and try again.
Six questions. Answer them in writing before you open the answers.
One counts more than the other five. You can get five right and still not pass this step.
A colleague is predicting which customers will cancel their subscription. She shows you the built-in importance table from a tree eight questions deep. Three lines matter.
customer_id.days_since_last_login, which her system records as of the day she pulled the data.plan_type. She says this proves the plan somebody is on has nothing to do with whether they cancel.Take the three in turn. Say what is happening in each case and what you would do about it. One of them is not a discovery at all, one might be a leak and you cannot tell from the table, and one conclusion does not follow.
duration first. Both were right. Why is that a problem, and what catches it?customer_id is not a discovery, it is Part 3. An id has as many different values as there are rows, so it offers the tree thousands of places to cut, and some of them separate the training rows by luck. Built-in importance counts that as work done. Drop the column, and if she wants proof first, run permutation importance and read the size of the number rather than its place in the list: it will be a few ten-thousandths either side of zero, which is nothing, even on a run where that still leaves it partway up a table of near-zeros.
days_since_last_login might be a leak and the table cannot tell her. If it is counted as of the day she pulled the data, then for somebody who cancelled three months ago it has been counting since they left, which is a fact about the cancellation. It would predict beautifully on held-back rows and be useless on a customer who is still here. The table will not catch it; the question will: at the moment she has to decide, for a customer still subscribed, what value would that column hold?
The zero proves nothing about plan type. It proves the tree did not use that column, and one obvious reason is that something else in her table already says it: the price paid, the seat count, the date they joined. Take the plan out and run thirty splits. If the score does not move, what you have learned is that the model does not need the column, which is not what she said. To answer what she actually said, stop modelling and count: cancellation rate by plan, one line per plan, the way you did months in Step 1.
Mark yourself right only if you got all three, and only if the middle one is a maybe in your answer rather than a verdict. Certainty about the second one is the failure this question is looking for.
duration is genuinely predictive on every row in the file, including the held-back ones, because every one of those calls already happened. Nothing you can compute will tell a leak apart from a genuinely brilliant column, because the two look identical in every table. What computation can do is hand you the shortlist: duration on its own predicts far better than any other column here, and a column that far ahead of the field is always worth an interrogation. The interrogation is a person's job, and the question is whether you would have that column for a person you are about to choose, at the moment you choose.Go back to Part 6 and read the box about what a ranking measures, then re-read your own Part 0 list with the money on it.
Those three lines are the three importance-table conversations you will actually have, and the middle one is where this gets uncomfortable.
Then write it out again tomorrow, from memory, and go on. That counts.
Run Part 3's noisy table at depth 1, 2, 3, 6 and with no limit, and write down where row_number comes each time and what its importance is.
It starts at nothing at depths 1 and 2 and ends at the top of the table. Write one line on what that curve means for the advice "use a deeper model, it will find more".
Do Part 4's permutation importance five times, with random_state set to 0, 1, 2, 3 and 4, and look at the top five each time.
The order will not be identical. Write down which columns hold their place and which move, then answer this: what would you have to do before quoting a ranking as if the order were a fact?
Take the six columns from Part 2's small tree and train a model on those six alone. Score it over thirty splits and compare it with the 47-column model at the same depth. Then do the same with the six columns from the bottom of the permutation ranking.
The top six will beat the full table and the bottom six will not, which is the opposite of a cautionary tale. So write the honest version instead: say what that result does and does not license you to conclude, and then answer this one question. What would have happened to it if row_number had been one of the six at the top?
All six, or it is a not yet.
Keep your two rankings. Step 7 asks the question underneath all of them: what is a score actually measuring, who pays when it is wrong, and why accuracy has been the wrong number since Step 1.