Step 6

Which columns are pulling their weight

Your model will hand you a ranked list of its columns if you ask. You add a column that is nothing but the row number, and the list does not put it anywhere near last. Then you learn the two questions that catch that, and the one thing no ranking can ever tell you.

About 4 hours10 partsBest over three sittingsThe one about not trusting a ranking
Read this first3 min

What this step gives you

Step 5 ended with you asking the model a question directly for the first time: did you even use this column? You used feature_importances_ and only looked at whether the number was zero.

This step is about the size of those numbers, and it is mostly about how badly you can be fooled by them.

The one sentence

An importance table tells you what a model leaned on. It does not tell you what matters, it does not tell you what would help, and it cannot tell you that a column is cheating.

The thing that happens in Part 3

You add one column to your table. It holds the row number: 0 for the first person, 1 for the second, and so on to 4,520. It is not about anybody. It cannot possibly mean anything.

You train a tree six questions deep and ask for the ranking. Your row number lands high in it, above columns you worked hard for in Step 4.

Then in Part 4 you ask a different question, of the people the model never saw, and get a completely different answer about the same column. Both are correct. They are answers to different questions, and almost nobody who quotes the first one knows that.

The page does not tell you either position in advance. You are asked to write down a guess first, three times, and being wrong is the point.

What you actually do, in order

  1. Rank five clues about a shopper by what you would pay for each, on paper, and find the one you would only know afterwards.
  2. Ask a small tree which columns it leaned on, and find that 41 of the 47 got no look at all.
  3. Add the row number and see how high it comes.
  4. Ask the test rows the same question and get a different answer.
  5. Find three columns with an importance of exactly zero that are not useless.
  6. Put duration back in and watch both rankings fall for it at once.
How the time actually goes

Parts 0 and 1 are twenty-five minutes on paper. You do not open the notebook until Part 2.

Parts 3 and 4 are one idea in two halves and want an hour together. If you have less than an hour left, stop at the end of Part 2 and start fresh with Part 3 another day, rather than stopping between them. The other good places to stop are the end of Part 5, and before Part 7, which is a sitting of its own.

Part 0Unplug20 minAway from the screen

Which clue would you give up?

Paper and pen. Laptop shut. Copy the five clues onto your paper first, so the table below is something you can work on.

A coat shop wants to know which visitors will buy a coat, so that it can decide who to send an offer to on Friday morning. Here are five things it could know about a visitor.

ClueWhat you would pay
Whether it rained that week
How much they spent in the shop last winter
How many minutes they spent in the shop on the day
Their coat size
Whether they came in with somebody else

First, rank them

You have one hundred pounds to spend on clues. Split it between the five, in whatever way you like, giving most to the one you would least want to lose. Write the five numbers in the second column. They must add up to a hundred.

This takes longer than you expect and the arguing with yourself is the exercise. Do it before you read on.

Now the question that matters

It is Friday morning. Nobody has come into the shop yet, and you are choosing who to send the offer to.

Go down your five clues and ask of each one: would I have this, for a person I am about to choose, right now?

Cross out the ones you would not. Then add up the money you put on the ones you crossed out.

One of the five cannot survive that question. How many minutes somebody spent in the shop is a fact about a visit that has not happened. On Friday morning it does not exist for anybody you care about.

It is also, if you are honest, probably one of the two you paid most for, because it is the best clue on the list. Somebody who spends forty minutes in a coat shop is buying a coat.

Write this down and keep it

You have met this before. In Step 3 it was called duration, the length of a phone call that had not happened yet.

Write one line: what was different about the way you found it this time, when you were pricing clues rather than reading a score?

The other four are not equal either

Look at "whether it rained that week". It is the same for every visitor on the same day, so it can never tell two of Friday's visitors apart. A clue that does not vary between the people you are choosing between is worth nothing for choosing between them, however much it explains about the shop's takings overall.

And look at "their coat size" next to "how much they spent last winter". If everybody who spent a lot last winter also happens to be recorded with a coat size, and the people who spent nothing are not, then coat size is quietly saying the same thing as last winter’s spending, and taking it away might cost you nothing at all.

Keep your list of five with the money written on it. Part 6 comes back to it.

Part 1Words5 min

Words to know

Importance
A number per column, saying how much of the model's work that column did. Every model that offers one means something slightly different by it, and none of them mean "how much this matters in the world".
Built-in importance
The kind a tree hands you for free in feature_importances_. Worked out from the training rows, while the tree was being built.
Permutation importance
The other kind, and the one you can trust further. Shuffle one column so it says nothing, score the model again, and see how much worse it got. Worked out on rows the model never saw.
Redundant
A column that says something another column already said. Two redundant columns will share the credit, or one will take it all and the other will show as zero. Step 3 used the word twin for two copies of the same person; this is two columns saying the same thing, which is not the same problem.
Weight
Used loosely in this step’s title and its variable names for how much a column did. Step 9 uses the same word for a number a different kind of model multiplies a column by. They are not the same thing.
Ranking
A list in order. Easy to read, easy to put in a slide, and the reason this step exists.
The shape of the whole step

There are three ways a column can be at the top of an importance table: it is genuinely useful, it is noise the model was able to carve up, or it is cheating.

The table looks identical in all three cases.

Part 2Run20 min

Ask the model what it leaned on

Open band-06.ipynb. Every cell on this page is already in it, in this order. You run them; you do not type them. The typing is in Part 7, on your own table.

If JupyterLab is not open in your browser

Open a terminal, move into the course folder with cd, then type jupyter lab. It does not stay running between days.

If you are coming back on a new evening

Run every cell from the top before you start. Nothing is remembered overnight, and the cells in Parts 5 and 6 need names that were made back in this one.

If you see NameError: name 'ready' is not defined, that is all this is. Nothing is broken.

If your numbers are different from mine

Check that every random_state=0 on the page is in your cell too. Without it the rows are shuffled differently every run, so your numbers change every time you press Shift and Enter and no two people in the world can compare notes.

If the numbers still differ, run every cell from the top in a fresh kernel: Kernel, then Restart Kernel and Run All Cells.

The first cell is Step 4's table and Step 3's split, in one go. Nothing here is new except the last import, which you will need in Part 4.

import pandas as pd
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.inspection import permutation_importance

df = pd.read_csv("data/bank.csv", sep=";")
y = (df["y"] == "yes").astype(int)

plain = ["age", "balance", "day", "campaign", "pdays", "previous"]
many = ["job", "marital", "education", "contact", "month", "poutcome"]

ready = df[plain].copy()
for name in ["default", "housing", "loan"]:
    ready[name] = (df[name] == "yes").astype(int)

ready = pd.concat([ready, pd.get_dummies(df[many], dtype=int)], axis=1)

X_train, X_test, y_train, y_test = train_test_split(
    ready, y, test_size=0.25, random_state=0)

print(ready.shape)
You should see

(4521, 47)

Now the tree from Step 4, and its ranking.

small = DecisionTreeClassifier(max_depth=3, random_state=0)
small.fit(X_train, y_train)

weight = pd.Series(small.feature_importances_, index=X_train.columns)

print(round(weight.sum(), 4))
print((weight > 0).sum())
print(weight.sort_values(ascending=False).head(6).round(4).to_string())

index=X_train.columns writes the column names down the side of the list, so weight["age"] reads off one column's number by name. .sort_values(ascending=False) puts the biggest first. .head(6) keeps the top six, the way it kept the top rows of a table in Step 0. .to_string() prints the names beside the numbers without the extra line pandas adds about the kind of thing it is.

You should see
1.0
6
poutcome_success      0.5974
month_oct             0.1542
age                   0.1104
day                   0.0547
balance               0.0460
education_tertiary    0.0373

Three things in that, and the first two are the ones people miss.

Now the same thing with no limit on the depth.

big = DecisionTreeClassifier(random_state=0)
big.fit(X_train, y_train)

big_weight = pd.Series(big.feature_importances_, index=X_train.columns)

print((big_weight > 0).sum())
print(round(big.score(X_test, y_test), 4))
You should see

47, then 0.8073

Every single column now has an importance above zero. All 47 of them matter. Or so the table says.

And the model they all matter to scores 0.8073 on people it has never seen, which is well below the 0.878 you get by saying no to everybody. Step 3's memorising tree, with a ranking attached.

Hold those two together

The number of columns with a non-zero importance went from 6 to 47 because the model got deeper, not because the columns got better.

Importance is a fact about your model. It is not a fact about your data, and it is certainly not a fact about the world.

Part 3PredictModify30 min

A column that cannot mean anything

Time to find out how much that ranking is worth.

You are going to add one column holding nothing but the row number. The first person gets 0, the second gets 1, and so on to 4,520. The rows of this file are a random tenth of a bigger file, so the order means nothing whatever: it is not time, it is not the order the calls were made, it is not anything.

Predict first. Write one number.

You will train a tree six questions deep on 48 columns, one of which is that row number, and rank them.

Where does the row number come? Write a position from 1 to 48 before you run it. Most people write 48.

noisy = ready.copy()
noisy["row_number"] = range(len(noisy))

N_train, N_test, y_train, y_test = train_test_split(
    noisy, y, test_size=0.25, random_state=0)

deep = DecisionTreeClassifier(max_depth=6, random_state=0)
deep.fit(N_train, y_train)

deep_weight = pd.Series(deep.feature_importances_, index=N_train.columns)

print(deep_weight.sort_values(ascending=False).head(5).round(4).to_string())

range(len(noisy)) counts from 0 up to one less than the number of rows, which is exactly the row numbers. Six questions rather than three, because a small tree does not have room to be fooled, and you will see that in a moment.

You should see
poutcome_success    0.3168
age                 0.1086
row_number          0.0946
pdays               0.0933
month_oct           0.0914
print(round(deep_weight["row_number"], 4))
print(round(deep.score(N_test, y_test), 4))
You should see

0.0946, then 0.8806

Third out of forty-eight. Ahead of pdays. Ahead of the October column that you found by hand in Step 1 and watched the model rediscover in Step 4.

Third is this split, not a law. Over thirty different splits it lands between second and eleventh, and it never once falls out of the top quarter of the table. The position moves; the fact that a column made of nothing sits near the top does not.

The column is the row number. There is nothing in it. If you put that ranking in front of a manager, one of its top lines would be a lie, and you would have no idea.

Why it happened, and it is not a bug

Built-in importance is worked out while the tree is being built, on the training rows, and it counts how much each question tidied up the heap underneath it.

Now think about what the row number offers a tree. It has 3,390 different values in the training half, so there are 3,389 places to cut it. Most other columns offer far fewer: month_oct offers exactly one cut, because it is a one or a zero. The exception proves the rule, and it is worth knowing: balance has 1,927 different values, and it floats up this table too.

Give a tree three thousand places to cut and some of them will separate the training rows a little, by luck. The tree takes the best of those, and it looks like work done, and it is counted as work done.

The rule this gives you, and it is worth more than the demonstration

Built-in importance favours columns with many different values, whether or not those values mean anything.

So a customer number, an account reference, a postcode, a timestamp, a price to two decimal places: all of them will float up your ranking for no better reason than that they are finely divided. If the top of your table is an id column, one whose only job is to name a row rather than say anything about the person in it, that is not a discovery. It is this.

One more thing before you leave it. Give the same 48 columns to a tree with three questions instead of six.

shallow = DecisionTreeClassifier(max_depth=3, random_state=0)
shallow.fit(N_train, y_train)

shallow_weight = pd.Series(shallow.feature_importances_, index=N_train.columns)

print(round(shallow_weight["row_number"], 4))
print((shallow_weight > 0).sum())
You should see

0.0, then 6

Exactly zero, tied at the bottom with 41 other columns. The row number was offered to this tree too, with the same three thousand places to cut, and it never got used.

This tree had no room to be fooled. It only had three questions and it spent them on the columns that were worth it.

Do not make that a promise, though. Run the same three questions on thirty different splits and the row number picks up a real number on seven of them, once as high as fourth. Three questions is thin protection, not armour. Depth does not only cost you honesty in the score, which is Step 3. It costs you honesty in the ranking too, and shallowness only makes it less likely.

Part 4Run25 min

Ask the test rows instead

There is a second way to ask which columns matter, and it is slower, better, and almost never the one people use.

Take the model you already trained. Take the 1,131 people it has never seen. Now pick one column and shuffle it, so that everybody has somebody else's value for that column and the column still looks perfectly normal. Score the model again.

If the score falls, the model was using that column for something real. If the score does not move, it was not.

Predict first, and this one is the point of the step

Same model, same 48 columns. Shuffle each column in turn, on the people the model has never seen.

Where does row_number come now? Write the position down before you run it.

shuffled = permutation_importance(deep, N_test, y_test,
                                  n_repeats=10, random_state=0)

honest = pd.Series(shuffled.importances_mean, index=N_test.columns)

print(honest.sort_values(ascending=False).head(5).round(4).to_string())

permutation_importance does the shuffling for you, once per column. You hand it the trained model and the rows to test it on, and it hands back a result you read the numbers off. n_repeats=10 shuffles each column ten times and averages, because one shuffle is one go, the way one split was one go in Step 4. importances_mean is the average drop in score for each column.

You should see
poutcome_success      0.0134
month_oct             0.0068
day                   0.0035
housing               0.0020
education_tertiary    0.0015
print(round(honest["row_number"], 4))

order = honest.sort_values(ascending=False)
print(list(order.index).index("row_number") + 1)

The second line finds where in the sorted list that name sits and counts from 1, because you predicted a position and the ranking prints numbers.

You should see

-0.0019, then 47

Below zero means the model got better when that column was scrambled. Which makes sense: everything the tree learned from the row number was fitted to the training people and was noise for everybody else, so wrecking it helped.

Two rankings. Same model, same columns, same afternoon. One puts the row number third and one puts it forty-seventh.

Which is right, and why the answer is not "the second one"

Both are correct answers to their own question.

The first asks: how much did this column shape the model I built? The row number shaped it a lot. That is true, and it is why the model is worse than it looks.

The second asks: how much does this model lose when this column is scrambled, on people I have not met? Nothing at all. Also true, and it is much closer to the question you were actually asking.

Closer, not the same. It is what scrambling costs the model you have. It is not what you would lose by dropping the column and training again, and Part 5 is about the gap between those two.

Notice how much smaller the second set of numbers is. The best column in the table costs 0.0134 when it is scrambled, which is just over one point of score. Built-in importance is a share of one and always adds to 1; permutation importance is a real amount of score, and real amounts are usually small. Never compare a number from one table with a number from the other.

The cost, and the catch

Permutation importance re-scores the model once per column per repeat. That is 48 columns times 10 repeats, and it took a moment. On a big model with a thousand columns you would feel it.

And it has a real weakness, which the next part is about: shuffle one of two columns that say the same thing, and the model just reads the other one, so both look worthless.

Part 5Investigate25 min

Zero does not mean useless

Step 5 found something you can use here. Four columns in this file say the same thing about the same 816 people: pdays, previous, poutcome_unknown, and the contacted_before flag you built. They are redundant, in the Part 1 sense.

Put the flag back and ask the ranking about all four.

dup = ready.copy()
dup["contacted_before"] = (df["pdays"] != -1).astype(int)

D_train, D_test, y_train, y_test = train_test_split(
    dup, y, test_size=0.25, random_state=0)

dup_tree = DecisionTreeClassifier(max_depth=6, random_state=0)
dup_tree.fit(D_train, y_train)

dup_weight = pd.Series(dup_tree.feature_importances_, index=D_train.columns)

print(dup_weight[["pdays", "previous",
                   "poutcome_unknown", "contacted_before"]].round(4).to_string())

Handing a list of names inside the square brackets picks out those rows of the ranking, in that order.

You should see
pdays               0.1148
previous            0.0000
poutcome_unknown    0.0000
contacted_before    0.0000

One column takes everything and three take nothing.

There is nothing to choose between them. The tree met pdays first, split on it, and by the time it looked at the other three there was nothing left for them to explain. A ranking cannot show you a tie, so it shows you a landslide.

Watch what happens when the winner leaves.

gone = dup.drop(columns=["pdays"])

G_train, G_test, y_train, y_test = train_test_split(
    gone, y, test_size=0.25, random_state=0)

without = DecisionTreeClassifier(max_depth=6, random_state=0)
without.fit(G_train, y_train)

without_weight = pd.Series(without.feature_importances_, index=G_train.columns)

print(without_weight[["previous", "poutcome_unknown",
                      "contacted_before"]].round(4).to_string())
print(round(dup_tree.score(D_test, y_test), 4))
print(round(without.score(G_test, y_test), 4))
You should see
previous            0.0183
poutcome_unknown    0.0000
contacted_before    0.0000

then 0.8859, then 0.8842

previous was worth zero and is now worth something. It did not change. The redundant column that beat it left.

And the score barely moved, because the information never went anywhere. You deleted the column that took 11 percent of the credit and lost less than two thousandths of a point.

That was one split, and you know from Step 4 what one split is worth. Here is the same question asked thirty times.

with_it, without_it = [], []

for seed in range(30):
    a, b, c, d = train_test_split(dup, y, test_size=0.25, random_state=seed)
    tree = DecisionTreeClassifier(max_depth=6, random_state=0).fit(a, c)
    with_it.append(tree.score(b, d))

    a, b, c, d = train_test_split(gone, y, test_size=0.25, random_state=seed)
    tree = DecisionTreeClassifier(max_depth=6, random_state=0).fit(a, c)
    without_it.append(tree.score(b, d))

print(round(sum(with_it) / 30, 4))
print(round(sum(without_it) / 30, 4))

Both splits use the same seed each time round, so the two models are always being asked about the same people. That is the only fair way to compare them.

You should see

0.8844, then 0.8845

Averaged over thirty different sets of people, the table without pdays is very slightly ahead. Not meaningfully ahead: a ten-thousandth of a point is nothing. But the column that took 11 percent of the ranking is worth, as far as anyone can measure, exactly nothing at all.

This is the loop you will need in Part 7, and it is the only tool on this page that answers the question people think a ranking answers.

The two mistakes this stops you making

Deleting a zero. "It scored zero, so I dropped it" is how people delete the backup for a column they are about to lose. If your top column is one day recorded differently, or stops arriving, the zeros are what you have left.

Believing a landslide. A column at the top with three redundant columns underneath it is not four times as important as anything. It won a race by arriving first.

The fix is the loop you just ran: take the column out, run thirty splits, and see what actually happens to the score. It is slower than reading a ranking, and it is the only one of the three that answers the question you meant.

One warning about it, so you do not swing too far the other way. A column can be redundant, cost nothing to delete, and still be related to the answer. The score not moving tells you the model does not need it. It does not tell you the column has nothing to do with anything.

Part 6PredictRun20 min

The column both rankings love

One last one, and it is the reason this step cannot be the last word.

Put duration back in. You met it in Step 3: the length of the phone call, which does not exist until the call has happened.

Predict first. Two positions.

Where does duration come in the built-in ranking, and where in the permutation ranking?

You know it is leakage. The question is whether either ranking knows.

leaky = ready.copy()
leaky["duration"] = df["duration"]

L_train, L_test, y_train, y_test = train_test_split(
    leaky, y, test_size=0.25, random_state=0)

cheat = DecisionTreeClassifier(max_depth=6, random_state=0)
cheat.fit(L_train, y_train)

print(round(cheat.score(L_test, y_test), 4))

built_in = pd.Series(cheat.feature_importances_, index=L_train.columns)
print(built_in.sort_values(ascending=False).head(3).round(4).to_string())
You should see

0.9036, then

duration            0.4587
poutcome_success    0.1816
age                 0.0740
shuffled_leaky = permutation_importance(cheat, L_test, y_test,
                                        n_repeats=10, random_state=0)

honest_leaky = pd.Series(shuffled_leaky.importances_mean, index=L_test.columns)
print(honest_leaky.sort_values(ascending=False).head(3).round(4).to_string())
You should see
duration            0.0580
poutcome_success    0.0213
contact_cellular    0.0118

Top of both. By a distance, in both. Nearly half of the built-in table, and getting on for three times the next column in the permutation table.

The score is 0.9036, as high as anything this course has produced. Step 3 got a number just like it, in the same way. Every check you have agrees this is the best column you own.

This is the sentence to take out of the step

Permutation importance did not catch it, and it never will, because duration genuinely does predict the answer on rows the model has never seen. Every one of those rows is a call that already happened.

A ranking measures how well a column predicts. Leakage is a column that predicts well and will not be there. No amount of measuring the first thing will ever tell you about the second.

The only thing that catches it is the Part 0 question, asked by a person: would I have this, for somebody I am about to choose, at the moment I have to choose?

Which gives you what a ranking is actually for. The ranking is not the check. The ranking is the shortlist. Every column near the top of it gets that question asked out loud, and the higher it is, the harder you ask.

So you can answer the question after all, when somebody asks which columns matter most in your model. Not with the built-in table, and not with a shrug. You say which of the two rankings you are quoting, what its numbers mean, and which of the columns near the top you checked by hand and how. That answer is three sentences long and you now have all three.

Part 7Make60 minNo answer given

Your own table

No code here and no answer at the end. Your own table, the one you signed up in Step 0 and have carried through five steps. Give it a sitting of its own; it is the longest part in the step.

Use new names

mine, not df. my_weight, not weight. If you reuse the names above, the cells you already ran will still work and print a ranking of the wrong table.

The brief

Produce one importance table for your own model that you would be willing to show somebody, and delete one column with the reason written next to it.

It is done when

If your table has an id column in it

Go and look at where it comes in the built-in ranking before you drop it. This is the single most common version of Part 3 in real work, and seeing it happen on your own data is worth more than reading about it here.

Then drop it. An id is a name, not a fact about anybody.

If permutation importance takes too long to run

Drop n_repeats to 3 and say in your README that you did. Fewer repeats means a noisier answer, and the honest thing is to say which you took rather than to quote it as if it were the ten.

Working alone

You are the person who knows your subject. Put the README away, and tomorrow read your top five out loud, in order, before you look at the numbers again. You are listening for your own surprise: "that one is at the top?" is the beginning of finding either a leak or something real that nobody had noticed.

If you have somebody to ask

Read the same five to them, in order, without saying what the numbers are. Their surprise is worth more than yours, because they do not know what the model did and you do.

Part 8Ship15 min

Write it up

Add to the README you started in Step 0.

Then say it to somebody

Find a person who does not code. Tell them you added a column of row numbers to a model and it came third out of forty-eight.

If they laugh, they have understood it. That reaction is the correct one, and most people who read importance tables for a living have never had it.

If they do not laugh, tell them the column was 0, 1, 2, 3 counting down the page and nothing else, and try again.

Part 9Check15 min

Check yourself

Six questions. Answer them in writing before you open the answers.

One counts more than the other five. You can get five right and still not pass this step.

  1. Built-in importance put the row number third. Permutation importance put it forty-seventh. Which one was wrong, and what was the other one measuring?
  2. Your six columns at depth 3 became 47 columns at no depth limit. What changed, and what did not?
  3. Three columns scored exactly zero and none of them was useless. Explain the zero.
  4. A colleague is predicting which customers will cancel their subscription. She shows you the built-in importance table from a tree eight questions deep. Three lines matter.

    1. Top of the table, taking a third of the total: customer_id.
    2. Second: days_since_last_login, which her system records as of the day she pulled the data.
    3. Bottom, at exactly 0.0000: plan_type. She says this proves the plan somebody is on has nothing to do with whether they cancel.

    Take the three in turn. Say what is happening in each case and what you would do about it. One of them is not a discovery at all, one might be a leak and you cannot tell from the table, and one conclusion does not follow.

  5. Permutation importance gave the best column 0.0134 and built-in importance gave it 0.5974. Why can you not compare those two numbers?
  6. Both rankings put duration first. Both were right. Why is that a problem, and what catches it?
Open the answers, once all six are written down
  1. Neither was wrong. The first measured how much the column shaped the tree that was built, on the training rows, and it shaped it a lot. The second measured how much the column earns on people the model has never seen, and it earns nothing. You wanted the second, and the first is the one every model hands you for free.
  2. The model changed. The columns did not. A deeper tree asks more questions, so more columns get used for something, so more of them have a number above zero. Importance is a fact about the model you built, and the deeper that model, the more of the ranking is fitted to your particular rows.
  3. The tree met an equivalent column first and took the credit for both. By the time it looked at the other three there was nothing left to explain. Zero means "this told me nothing I did not already know", and when you delete the column that beat them, one of the zeros stops being a zero.
  4. customer_id is not a discovery, it is Part 3. An id has as many different values as there are rows, so it offers the tree thousands of places to cut, and some of them separate the training rows by luck. Built-in importance counts that as work done. Drop the column, and if she wants proof first, run permutation importance and read the size of the number rather than its place in the list: it will be a few ten-thousandths either side of zero, which is nothing, even on a run where that still leaves it partway up a table of near-zeros.

    days_since_last_login might be a leak and the table cannot tell her. If it is counted as of the day she pulled the data, then for somebody who cancelled three months ago it has been counting since they left, which is a fact about the cancellation. It would predict beautifully on held-back rows and be useless on a customer who is still here. The table will not catch it; the question will: at the moment she has to decide, for a customer still subscribed, what value would that column hold?

    The zero proves nothing about plan type. It proves the tree did not use that column, and one obvious reason is that something else in her table already says it: the price paid, the seat count, the date they joined. Take the plan out and run thirty splits. If the score does not move, what you have learned is that the model does not need the column, which is not what she said. To answer what she actually said, stop modelling and count: cancellation rate by plan, one line per plan, the way you did months in Step 1.

    Mark yourself right only if you got all three, and only if the middle one is a maybe in your answer rather than a verdict. Certainty about the second one is the failure this question is looking for.

  5. Because they are different kinds of number. Built-in importance is a share of the model's work and the whole table adds up to 1, so 0.5974 means "sixty percent of what this tree did". Permutation importance is an amount of score, so 0.0134 means "this model gets about one and a third points worse when that column is scrambled". One is a proportion of something, the other is a quantity of something else. And note what neither of them is: what you would lose by dropping the column and training again. Part 5 deleted a column worth 11 percent of the built-in table and the thirty-split average did not move.
  6. Because a ranking measures how well a column predicts, and leakage is a column that predicts well and will not be there when you need it. duration is genuinely predictive on every row in the file, including the held-back ones, because every one of those calls already happened. Nothing you can compute will tell a leak apart from a genuinely brilliant column, because the two look identical in every table. What computation can do is hand you the shortlist: duration on its own predicts far better than any other column here, and a column that far ahead of the field is always worth an interrogation. The interrogation is a person's job, and the question is whether you would have that column for a person you are about to choose, at the moment you choose.
If you got the marked one wrong

Go back to Part 6 and read the box about what a ranking measures, then re-read your own Part 0 list with the money on it.

Those three lines are the three importance-table conversations you will actually have, and the middle one is where this gets uncomfortable.

Then write it out again tomorrow, from memory, and go on. That counts.

StretchOptionalHarder

If you want more

Bronze

Run Part 3's noisy table at depth 1, 2, 3, 6 and with no limit, and write down where row_number comes each time and what its importance is.

It starts at nothing at depths 1 and 2 and ends at the top of the table. Write one line on what that curve means for the advice "use a deeper model, it will find more".

Silver

Do Part 4's permutation importance five times, with random_state set to 0, 1, 2, 3 and 4, and look at the top five each time.

The order will not be identical. Write down which columns hold their place and which move, then answer this: what would you have to do before quoting a ranking as if the order were a fact?

Gold

Take the six columns from Part 2's small tree and train a model on those six alone. Score it over thirty splits and compare it with the 47-column model at the same depth. Then do the same with the six columns from the bottom of the permutation ranking.

The top six will beat the full table and the bottom six will not, which is the opposite of a cautionary tale. So write the honest version instead: say what that result does and does not license you to conclude, and then answer this one question. What would have happened to it if row_number had been one of the six at the top?

Before you move on

All six, or it is a not yet.

The thing you built.
Both rankings for your own model, top ten of each, side by side in your README.
The question, asked five times.
Your top five columns, each with one written line saying whether you would have it at the moment you need the answer.
The deletion.
One column gone, the reason beside it, and what thirty splits said about losing it. The reason cannot be that it scored zero.
The answer you could give out loud.
Three sentences in your README for "which columns matter most in your model": the ranking you would quote, what its numbers mean, and the column you checked by hand.
Check yourself.
Five of six right, including the marked one about the three lines of the table.
The build log.
One thing that surprised you, and one prediction you wrote down that turned out wrong. Parts 3, 4 and 6 all asked you to predict a position.
It still runs.
Kernel, then Restart Kernel and Run All Cells. Nothing should turn red.

Keep your two rankings. Step 7 asks the question underneath all of them: what is a score actually measuring, who pays when it is wrong, and why accuracy has been the wrong number since Step 1.

Step 7 is being built. It is the one where accuracy finally runs out, and you learn what to say when somebody asks how good your model is.