Step 5

Making columns that were not there

Twelve shops on a grid, and no line straight up or straight across separates them. That is a picture of something a tree cannot say cheaply, and of the one column that fixes it. Then four sensible columns for the bank file, measured properly, and not one of them pays.

About 3 hours 50 minutes10 partsBest over two sittingsThe one about knowing when to stop
Read this first3 min

What this step gives you

Halfway through Step 4 a tree found October on its own, out of twelve month columns nobody explained to it. Then the rest of that step was about how little you could safely conclude from it.

This step starts from the other side of that. There are things it can only reach by building a staircase, however deep you let it go and however many columns you hand it, and they are not exotic. The commonest one is the gap between two numbers.

The one sentence

A model can only use what it can say cheaply. A tree says "is this number below this cut", one question at a time, so a relationship between two columns costs it many questions and it still only gets close. Hand the relationship over as a column and it costs one.

The two halves of this step

In the first half you meet a small table where the answer is exactly that: a gap. You watch a tree fail on it, then hand it one column you built yourself, and watch a single question get every person right.

In the second half you go back to the bank file and try the same trick four times, properly measured, and it does not work once. Then you work out why, which is worth more than the trick.

The two halves together are the actual skill: knowing which missing column is worth building, and being willing to find out that the answer is none of them.

What you actually do, in order

  1. Plot twelve shops on a grid, by hand, and try to separate them with a straight line.
  2. Open a small table with two dates in it and watch a tree struggle.
  3. Build one column, and watch one question get everybody right.
  4. Take the lid off and see what the tree was doing instead.
  5. Build four columns for the bank file that ought to help.
  6. Measure them the way Step 4 taught you, and watch all four fail.
  7. Build the one your own table needs, and measure it the same way.
How the time actually goes

Parts 0 and 1 are half an hour on paper, and you do not open the notebook until Part 2. Part 0 takes longer than it looks, because plotting twelve points by hand is slower than reading about it.

If you have an hour and a half, Parts 0 to 4 are a whole story and a good place to stop. If you have twenty minutes, this is not the night for Step 5.

If three quarters of an hour is what you always have, here is the whole step mapped to it: Parts 0 and 1 the first night, Parts 2 and 3 the second, Part 4 the third, Parts 5 and 6 the fourth, Part 7 on its own the fifth. Parts 5 and 6 are one idea in two halves and must not be split. Part 5 ends by naming the column that did the damage; Part 6 is the reason.

Part 0Unplug25 minAway from the screen

Draw the rule on a grid

Paper and pen. You want a page with room for a square about the size of your hand.

A chain of shops. For each shop you know the month it opened, counted from the month the company started, and the month a particular sale happened, counted the same way. Some sales were big and some were not.

Here are twelve real rows out of the table you will open in Part 2.

Opened, monthSale, monthBig sale
548yes
817no
1483yes
2169yes
2582yes
3073yes
3540no
4048no
4453no
58114yes
6164no
6469no

Draw it

Draw two lines for the edges of a square, one along the bottom and one up the left.

Along the bottom is opened. Mark 0 at the left, then 20, 40 and 60, evenly spaced. Up the left is sale. Mark 0 at the bottom, then 40, 80 and 120, evenly spaced. Squared paper makes this easier, and the back of an envelope is fine.

Now put a mark for each of the twelve shops: across to its opened month, up to its sale month. Use a small circle for the big sales and a cross for the rest. Twelve marks in all. Being a few millimetres out will not matter.

Now the exercise, and give it a proper try

Draw one straight line, either straight up and down or straight across, with circles on one side of it and crosses on the other.

Spend three minutes on it. Try up and down. Try across. Count how many of the twelve your best line gets right, and write the number down.

You cannot do it with one line. The best line across gets 10 of the 12, and so does the best line up and down. Never better, either way.

Now break the rule. Draw one line at any angle you like.

A line running diagonally, up and to the right, puts all six circles above it and all six crosses below. Every one of the twelve, and with room to spare: you can be quite sloppy about the angle and still get all of them.

Here is the same thing drawn properly, so you can check yours. The left panel is the best straight-across line there is, with the two shops it gets wrong. The right panel is one line at an angle.

One line across, the best there is: 10 of 12 0 20 40 60 0 40 80 120 opened, month sale, month One line at an angle: all 12 0 20 40 60 0 40 80 120 opened, month sale, month

Two straight lines, one across and one up and down, get eleven of the twelve. Three get all twelve, which is worth knowing, because that is exactly what a tree with three questions would do. It is not that a tree cannot separate these twelve. It is what it costs, and what happens to the thirteenth shop it has never seen.

Write this down and keep it

Look at the twelve rows again. Add a fourth column on your paper and, for each shop, subtract the opened month from the sale month. Twelve subtractions, done on the page rather than in your head.

Then write one sentence saying what the circles have in common that the crosses do not.

Write your sentence down before Part 2. If you get it now, the rest of the first half will feel obvious, and that is exactly the right way for it to feel.

What you just found out about trees

Every question a tree asks is one number against one cut. On a grid, that is a line straight up and down, or straight across. Never at an angle.

More questions means more lines, so a tree builds a staircase alongside a diagonal. With enough questions and enough shops the staircase gets close. It never becomes the diagonal, and every extra step costs a question and covers fewer shops.

Part 1Words5 min

Words to know

Feature
A column a model is allowed to look at. Everything in your table that is not the answer.
Feature engineering
Building a new column out of the ones you have, because the model cannot build it for itself.
Express
What a model can say, and how dearly. A tree says "balance is under 1,200" in one question. To say "sold minus opened is over 24" it needs a staircase of them, and it never quite arrives.
Interaction
When two columns only mean something together. A gap is one. So is "in October, and after the sixteenth", which Step 4's tree spent two of its three questions on.
Box
Step 4's word for a one-and-zero column standing for one answer, one per possible answer. The bank table has 38 of them, twelve for the months alone.
Domain knowledge
What somebody who does the work knows and your table does not say. It is where good features come from, and it is why you keep being sent to talk to people.
The test for a new column, and you can use it today

Before you build one, ask: could the model already say this?

Splitting age into five bands is something a tree can already do, so it gains nothing. Subtracting one date from another is not, so it might gain everything.

Part 2PredictRun20 min

A table with two dates in it

Open band-05.ipynb. Every cell on this page is already in it, in this order. You run them; you do not type them. The typing is in Part 7, on your own table.

If you are coming back on a new evening

Run every cell from the top before you start. Nothing is remembered overnight, and the cells later in this step need names that were made earlier in it.

If you see NameError: name 'extras' is not defined, that is all this is. Nothing is broken.

If your numbers are different from mine

Check that every random_state=0 on the page is in your cell too. Without it the rows are shuffled differently every run, so your numbers change every time you press Shift and Enter and no two people in the world can compare notes.

If the numbers still differ, run every cell from the top in a fresh kernel: Kernel, then Restart Kernel and Run All Cells.

If it is a while since you started JupyterLab: open a terminal, move into the course folder with cd, then type jupyter lab. Parts 6 and 7 of the setup page walk through it slowly.

Where this table comes from, so that you can trust it

The rest of this course runs on one real file. This part uses a made-up one, and you should know that before you draw any conclusions from it.

It holds 600 shops and exactly one relationship, so that you can see that relationship on its own. The bank file has no gap like this in it, for a reason that is the whole of Part 6.

It is built by tools/make_stores.py, out of real numbers from the bank file rather than out of a random generator, so that everybody gets the same 600 shops.

import pandas as pd
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split

stores = pd.read_csv("data/stores.csv")

print(stores.shape)
print(stores["big_sale"].mean().round(4))

A plain read_csv with no sep=";" this time, because this file uses commas like most of them do. The bank file was the odd one.

You should see

(600, 3), then 0.6267

600 shops, three columns, and 62.67 percent of the sales were big ones. So the lazy rule here is say yes to everybody, rather than say no as it was for the bank.

Hold that number loosely. In a moment you will keep a quarter of the shops back, and the lazy rule scores slightly differently on that quarter, the same way the bar in Step 3 was 0.878 on the test half and 0.8848 across the whole file.

print(stores.head(5))
You should see
   opened_month  sold_month  big_sale
0            34         101         1
1            48         109         1
2             6          45         1
3            15          83         1
4            46         110         1

The same shape as your paper grid. Two months and an answer. The first five all happen to be big sales, which is what a 63 percent yes rate looks like when you only glance at five rows, and is a good reason never to judge a table by its head.

Predict first. Write two numbers.

You are about to split these 600 shops the usual way and train the usual trees on the two columns as they stand.

What does a tree with two questions score on the shops it has never seen? What does a tree with no limit at all score?

You have an advantage here that you did not have in Step 3: you already know, from your paper grid, what the answer depends on. Use it. A good prediction now is worth more than a good score later.

X = stores[["opened_month", "sold_month"]]
y = stores["big_sale"]

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=0)

flat = DecisionTreeClassifier(max_depth=2, random_state=0)
flat.fit(X_train, y_train)

print(round(flat.score(X_test, y_test), 4))
You should see

0.7933

Better than saying yes to everybody, which scores 0.5667 on this particular test half. That is the 62.67 percent from two cells ago, measured on the 150 shops that were held back rather than on all 600. Better than lazy, then, and not good.

deeper = DecisionTreeClassifier(max_depth=3, random_state=0)
deeper.fit(X_train, y_train)
print(round(deeper.score(X_test, y_test), 4))

loose = DecisionTreeClassifier(random_state=0)
loose.fit(X_train, y_train)
print(round(loose.score(X_train, y_train), 4))
print(round(loose.score(X_test, y_test), 4))
print(loose.get_n_leaves())
You should see

0.8733, then 1.0, then 0.9467, then 20

Read the last three together, because you have seen this shape before. Perfect on the shops it learned from, 0.9467 on the ones it has not, and twenty endings to do it. That is Step 3's memorising tree, in a new place, and this time you can see exactly what it is memorising: a staircase, climbing beside a diagonal it cannot draw.

Part 3PredictRun20 min

One column you build yourself

Now the thing you worked out on paper.

stores["months_open"] = stores["sold_month"] - stores["opened_month"]

print(stores["months_open"].min(), stores["months_open"].max())
print(stores.head(3))

One line. Subtract one column from another and pandas does it for all 600 rows at once, the same way multiplying a column by a hundred did in Step 2.

You should see

0 71, then a table with a fourth column on the end holding 67, 61 and 39.

Predict first. One number.

Now a tree with one question, on all three columns, including the new one.

One question. Write down what it scores on the shops it has never seen.

X2 = stores[["opened_month", "sold_month", "months_open"]]

X2_train, X2_test, y_train, y_test = train_test_split(
    X2, y, test_size=0.25, random_state=0)

one_question = DecisionTreeClassifier(max_depth=1, random_state=0)
one_question.fit(X2_train, y_train)

print(round(one_question.score(X2_train, y_train), 4))
print(round(one_question.score(X2_test, y_test), 4))
You should see

1.0, then 1.0

Every shop it learned from, and every shop it never saw. One question.

In Step 3 a perfect score was a warning. Here it is not, and there are two separate reasons, and you should have both.

The first reason, and I am telling you because you would be right to ask

This table was built so that a big sale is exactly a gap of more than 24. That is in tools/make_stores.py, and you can read it. There is no noise in it at all, so a model that can say the rule gets every row, and nothing you ever do on real data will be this clean.

That is on purpose. You are being shown a mechanism, not a result. Part 5 is the same idea meeting a real file, and it goes very differently.

The second reason is that nothing leaked, and here is the test, which is narrower than you might guess.

When a new column is safe

A column is safe when each row is worked out from that row alone, out of things you would have for a new person on the day, using arithmetic you can read.

All three parts matter. Row by row is the one people miss: the moment a column is worked out across the whole table, the other rows have had a say in it, and some of those rows are your test set. The marked question at the end of this step is exactly that mistake.

months_open passes on all three. Every shop's gap is its own two numbers subtracted, and no other shop is consulted.

print(export_text(one_question, feature_names=list(X2.columns)))
You should see
|--- months_open <= 24.50
|   |--- class: 0
|--- months_open >  24.50
|   |--- class: 1

Two years. A shop that has been open more than two years makes big sales, and one that has not does not.

The tree threw away both of the columns it was given and used only the one you built. It was never going to find that rule in one question, and it found it immediately once somebody wrote it down.

This is the sentence to take out of the step

You did not add any information. Both numbers were already there, in every row, in front of the model, from the beginning.

What you added was a way of saying it.

Part 4Investigate20 min

What a tree can and cannot say

Look at what the unlimited tree was doing instead. This is the first four questions of the twenty-ending tree from Part 2.

print(export_text(loose, feature_names=list(X.columns), max_depth=2))

max_depth=2 here is a limit on the printing, not on the tree. The tree is unchanged; you are just asking to see the top of it.

You should see the top of a tree that starts
|--- sold_month <= 51.50
|   |--- opened_month <= 17.50
|   |   |--- sold_month <= 36.00

and six questions in all at this printing depth, with truncated branch of depth lines where it has stopped showing you.

Read the first three cuts as lines on your paper grid. One across at 51.5, one up and down at 17.5, then another across at 36. Across, up and down, across, and on it goes: twenty boxes in all, each a small rectangle with a single answer in it.

That is a staircase. It follows the diagonal, roughly, and every extra step costs a question and covers fewer shops. Sorted by size, the twenty boxes hold 1, 1, 2, 2, 4, 5 shops and upwards. Two of them hold a single shop each. Step 2 put it this way: a rule built around one person is not a rule, it is a note about that person.

The general shape of it, stated carefully

It is not that a tree cannot get there. Given enough questions and enough rows it will build a staircase against almost any boundary, and Part 2 watched it climb to 0.9467 doing exactly that.

What is true is narrower and more useful:

So the column does not tell the model anything new. It makes something cheap that was expensive, and cheap is what you are buying.

And it depends on the model, not just the data

A tree is unusually bad at gaps because it looks at one column at a time. The model in Step 9 is not: it adds the columns up with weights it works out for itself, and a gap is a subtraction, which is an addition with one weight turned negative. On this shop table that model gets every shop right from the two raw columns, without your help.

Gold at the bottom asks you to run it. The lesson is not that trees are poor. It is that "the model cannot express this" is a question about a particular model, and you have to know which one you are holding.

Which is why the phone calls matter

You cannot find months_open by staring at a table. You find it by knowing that shops take time to establish, which is something a shopkeeper knows and a spreadsheet does not.

Every step has asked you to show your work to somebody who does not code. This is the step where that stops being a courtesy and starts being the method.

Two kinds of new column, and one of them is nearly always a waste

Ask which kind you are about to build before you write any code. It will save you most of the time people waste on this.

Part 5Modify25 min

Four columns for the bank file

Back to the real file, and to the 47-column table you built in Step 4.

Here are four columns a sensible person would build for it. Each one is a real idea, and one of them is yours.

One is a sum and one is a ratio, which Part 4 says can pay. Two are things one question could arguably already say. Before you run anything, write down which is which, and which single one you would bet on.

bank = pd.read_csv("data/bank.csv", sep=";")
y_bank = (bank["y"] == "yes").astype(int)

plain = ["age", "balance", "day", "campaign", "pdays", "previous"]
many = ["job", "marital", "education", "contact", "month", "poutcome"]

ready = bank[plain].copy()
for name in ["default", "housing", "loan"]:
    ready[name] = (bank[name] == "yes").astype(int)

ready = pd.concat([ready, pd.get_dummies(bank[many], dtype=int)], axis=1)

print(ready.shape)

Step 4's Part 4, in one cell. Every piece of it is from that step: .copy(), the loop turning three yes-or-no columns into ones and zeros, get_dummies making 38 boxes, and concat joining them on sideways.

If you are wondering why this uses get_dummies

You should be. Step 4's Part 7 showed you that get_dummies breaks the moment you send new rows through a trained model, and had you write down when you would still use it.

This is that case. Nothing here is ever asked about a row from outside this table: every score comes from splitting these same 4,521 people. For looking at a table and comparing models on it, get_dummies is fine and shorter. The encoder is for the day you use the model.

If that made you stop and check, you did the right thing, and it is exactly what Step 4 was for.

You should see

(4521, 47)

extras = ready.copy()

extras["peak_month"] = bank["month"].isin(["mar", "sep", "oct", "dec"]).astype(int)
extras["contacts_total"] = bank["campaign"] + bank["previous"]
extras["balance_per_year"] = (bank["balance"] / bank["age"]).round(2)
extras["contacted_before"] = (bank["pdays"] != -1).astype(int)

print(extras.shape)
print(extras["peak_month"].sum())

.isin([...]) is from Step 1, where you used it to write the wider rule. The division and the addition work a column at a time, like the subtraction in Part 3.

You should see

(4521, 51), then 201

Only 201 of the 4,521 calls were made in one of your four months. That is worth noticing now rather than in Part 6: the column you were most confident about is switched on for one call in twenty-two.

Predict first. Two numbers, and be honest about them.

You are about to run Step 4's thirty splits on the 47-column table and the 51-column one, side by side.

What does each average? And how many of the thirty does the wider one win?

base_scores = []
extra_scores = []
wins = 0

for seed in range(30):
    a, b, c, d = train_test_split(ready, y_bank, test_size=0.25, random_state=seed)
    plain_tree = DecisionTreeClassifier(max_depth=3, random_state=0)
    plain_tree.fit(a, c)
    base_score = plain_tree.score(b, d)

    a, b, c, d = train_test_split(extras, y_bank, test_size=0.25, random_state=seed)
    extra_tree = DecisionTreeClassifier(max_depth=3, random_state=0)
    extra_tree.fit(a, c)
    extra_score = extra_tree.score(b, d)

    base_scores.append(base_score)
    extra_scores.append(extra_score)

    if extra_score > base_score:
        wins = wins + 1

print(round(sum(base_scores) / 30, 4))
print(round(sum(extra_scores) / 30, 4))
print(wins)

The same loop as Step 4's Part 6, with one thing changed: both models get three questions, so the only difference between them is the four new columns.

The four short names come out of train_test_split in its fixed order, and it is worth saying them out loud once more: a is the training columns, b the test columns, c the training answers, d the test answers. Short names here because they are used twice and thrown away.

You should see

0.8902, then 0.8882, then 5

Four thought-out columns, built from a real finding and two honest ratios, and the average went down. They won 5 of the 30.

And this time the single split agrees

In Step 4 the one split you were handed happened to flatter the new columns. Not this time. I ran random_state=0 on its own, and the wider table scores 0.8912, catching 23 of the people who would say yes for 8 wasted calls, against Step 4's 0.8983 catching 28 for 5. You do not need to run it; the cell above already told you the important thing.

Worse on the average and worse on the split. There is nothing to argue about, which is a comfortable place to be and does not happen often.

I also ran them one at a time, because four columns failing together could hide one that helps. The averages, against the 0.8902 to beat: peak_month 0.8882, contacts_total 0.8902, balance_per_year 0.8901, contacted_before 0.8902. Bronze at the bottom asks you to reproduce that.

One of the four is doing all the damage, and it is the one you were most sure of.

Part 6Investigate20 min

Why none of them paid

This is the part that is worth the step, and it starts with a question you can put to the model directly rather than guessing at: did it even use the column?

new_names = ["peak_month", "contacts_total",
             "balance_per_year", "contacted_before"]

used = pd.Series(0, index=new_names)

for seed in range(30):
    a, b, c, d = train_test_split(extras, y_bank, test_size=0.25, random_state=seed)
    tree = DecisionTreeClassifier(max_depth=3, random_state=0)
    tree.fit(a, c)
    weight = pd.Series(tree.feature_importances_, index=extras.columns)
    used = used + (weight[new_names] > 0)

print(used.to_string())

Three new things, all small.

You should see
peak_month          29
contacts_total       1
balance_per_year     5
contacted_before     0

Not one of them is what I expected when I wrote this page, and the first line is the opposite of it. Take them in turn.

peak_month did not get ignored. It got used, and it did harm.

This is the one I got wrong when I first wrote this page, and it is more interesting than what I expected.

I assumed the tree would leave it alone, because it already had twelve month columns and had picked October out of them. It does the opposite. It reaches for peak_month in 29 of the 30 splits, and it reaches high, because lumping four good months together makes a bigger, cleaner-looking split than any single month can.

Print the tree at random_state=0 and you can watch the damage. In Step 4 the second question was month_oct, and under it sat the leaf that caught six people in October after the sixteenth. With peak_month available, the tree asks that instead, and then both leaves underneath say no. The leaf that was catching people is gone, and the catches drop from 28 to 23.

So the coarse column won the argument and then had nothing useful to say. That is worth remembering: a tree takes the biggest split it can see right now, and the biggest split now is not always the best three questions later. Nobody warns you about this, and it is why Step 12 exists.

One correction to something you may believe from Step 1, while you are here. October is not special. In October 37 of 80 calls ended in a yes. In December it was 9 of 20 and in March 21 of 49, which are the same kind of rate. October wins because it holds 80 calls to December's 20, so the tree can build a bigger, steadier leaf out of it. Four months are genuinely good, and the tree picks the one it can lean on hardest.

contacts_total and balance_per_year were the right shape and the wrong idea

These two are the kind Part 4 said to look for: a sum and a ratio, neither of which one question can say.

The tree barely touched them: contacts_total once out of thirty and balance_per_year five times, and the average did not move.

The reason is simpler than the shape. Whether somebody opens a savings account has very little to do with their total number of phone calls, or with their money divided by their age. The arithmetic was legal. The idea was wrong.

A feature is a guess about how the world works, written as a column. Getting the shape right does not make the guess right, and no amount of care with the shape will rescue a guess that is not true.

contacted_before was already in the file three times over

Zero out of thirty. That bottom line is what a column with nothing whatever to add looks like, and Step 4 half told you already: it said contacted_before "never won a place in the top three questions". Now you have watched it lose thirty times.

It is the same 816 people, three times. pdays is -1 for exactly the other 3,705, so a cut anywhere between -1 and 1 makes the same two groups. poutcome_unknown, one of your 38 boxes, marks exactly the same 3,705, not almost. And previous is 0 for exactly those people too, so previous <= 0.5 is a third way to say it.

Three columns already said it in one question each. You built a fourth.

Now the one that is not in the file at all

The gap that would matter for this data is time. How long since this campaign began, how long between the last campaign and this one, whether this call came before or after the neighbouring bank went under.

You cannot build any of it. The file has a day and a month and no year, and you found that out in Step 0, in Part 6, when you went looking for a year column and there wasn't one.

That is what a missing column actually looks like in real work. Not a gap you can fill with arithmetic. A question the table cannot answer, which no amount of cleverness will fix, and which somebody has to be told.

One column here was engineered, and somebody else did it

Look at pdays again. Days since this person was last contacted. That is not something the bank's computer recorded on the day. It is a subtraction between two dates that somebody did before they gave you the file, which makes it the one hand-built gap in the whole table.

Hold it to the same standard as your four.

without_pdays = ready.drop(columns=["pdays"])

scores = []
for seed in range(30):
    a, b, c, d = train_test_split(without_pdays, y_bank, test_size=0.25, random_state=seed)
    tree = DecisionTreeClassifier(max_depth=3, random_state=0)
    tree.fit(a, c)
    scores.append(tree.score(b, d))

print(round(sum(scores) / 30, 4))

.drop(columns=["pdays"]) hands back a copy of the table with that column missing. It leaves ready alone, the way .copy() does.

You should see

0.891

Against 0.8902 with it. Very slightly better without.

So somebody did the job you are learning, on the right column, for the right reason, and on this file it did not pay either. That is the honest shape of this work, and it is why the measuring matters more than the ideas.

Write this in your README

Three lines. The two averages and the win count. Which of the four columns you would defend in a review even though it lost. And one sentence naming a column this table would need that cannot be built from what is in it.

Part 7Make60 minA sitting of its own

Your own table

No code here and no answer at the end. Your own table, the one you signed up in Step 0 and got into shape in Step 4.

Use new names

mine, not bank. my_extras, not extras. If you reuse the names above, the cells you already ran will still work and print numbers about the wrong table.

The brief

Build the columns your table needs, decide which of them are worth anything, and prove it with thirty splits rather than one.

It is done when

If your table has two dates in it

Start there. The gap between two dates is the feature your model is least able to find for itself, so it is the first one worth testing. Days between, months between, and whether one came before the other. Whether it pays is a separate question, and Part 6 is what that looks like when the answer is no.

Step 17 is where dates get their own step, and everything you write down now feeds it.

If a column of yours helps a lot

Be suspicious in the direction Step 3 taught you. Ask the Part 6 question of it: would you have this, for a new row, at the moment you have to answer?

A feature built out of columns you will not have on the day is still leakage, however honest the arithmetic that made it.

Working alone

If the table is about work you do, you are the person this step keeps telling you to go and find. Sit down with your three ideas before you build them and answer two questions in writing: what would I look at, and what do I wish somebody had recorded?

The second question is the one that pays, and it is the fifth item on the list above. Domain knowledge is not a thing other people have. It is what you know about this data that the table does not say.

If you know somebody else who works with this data

Ask them both questions before you build anything, and write their answer down word for word rather than in your own words. The wording is where the column hides.

Part 8Ship15 min

Write it up

Add to the README you started in Step 0.

Then say it to somebody

Find a person who does not code and draw them the grid from Part 0, with the dots and the diagonal.

Then tell them that the computer is only allowed to draw lines straight up and straight across. If they say "so you had to do the diagonal for it", you have explained the whole step in twenty seconds, and so have they.

Part 9Check15 min

Check yourself

Six questions. Answer them in writing before you open the answers.

One counts more than the other five. You can get five right and still not pass this step.

Two of the answers below offer half marks. Half is not a pass on the marked one: if you got half there, read the answer, then come back tomorrow and write it out again from memory before you go on. Half on any of the other five is a pass.

  1. A tree with no limit scored 0.9467 on the shops. One question on a column you built scored 1.0. What did the new column give the tree that it did not have before?
  2. Your model is a tree. You replace age with five age bands and the score does not move at all. A colleague says the bands must be wrong. What do you say?
  3. contacts_total is a sum of two columns, which is exactly the shape Part 4 said to look for, and it did nothing. Why not?
  4. A colleague has a table of hospital visits, one row per visit. It holds admitted_on and discharged_on as dates, the ward, the patient's age, and twenty other columns. She is predicting whether a patient will be readmitted within a month. Her tree scores badly and she asks you which of three new columns to build.

    1. The length of stay in days, from the two dates.
    2. Age in five bands instead of the raw number.
    3. The share of visits to that ward that ended in a readmission, worked out across her whole table.

    For each of the three, say whether it can help her and why. One of them has almost nothing to offer a tree, and one of them is not safe to build the way she has described it.

  5. The bank file has a day and a month and no year. Name one column you would build if it had a year, and say what a tree would do with it that it cannot do now.
  6. You build a new column and your score improves by half a point on one split. Name the two things you do before you tell anybody.
Open the answers, once all six are written down
  1. Nothing it did not have. Both numbers were in front of it all along. What the column gave it was a cheap way of saying the relationship between them: one question instead of a twenty-box staircase that was fitted to those particular shops and got the edges slightly wrong.
  2. That the bands are probably fine and that this is what should have happened. A tree can already cut age wherever it likes, so bands say nothing new. Columns of that kind usually do nothing, and when one does help it is because it saved your model a question rather than because it added anything. Say all that before you build it, not after.
  3. Because the shape of a feature and the truth of a feature are different things. A sum is something one question cannot say, so it was allowed to help. Whether somebody opens a savings account has very little to do with how many times they have been rung altogether, so it had nothing to say. The tree used it in one split out of thirty.
  4. Length of stay: build this one. It is the gap between two columns, which costs a tree a staircase and costs you one subtraction, and it is the thing a nurse would tell you matters. It is the shop table with a different hat on. It is also safe: worked out row by row, from two things she has the moment a patient leaves.

    Age bands: very little to offer. A tree can already cut age anywhere it likes, so the bands say nothing it could not say. The one case where they help is if her tree is short of questions and a band saves it one, and that is a reason to give the tree another question rather than to build the column. Full marks if you said that; half if you said they cannot help at all.

    The ward readmission rate: not like that. It is built from the answers of other rows, and worked out across the whole table, so rows in her test set have helped build a column her training rows learn from. That is Step 4's scaler with a new face, and it is worse than the scaler, because a scaler learns a range and this learns the answer.

    How badly does it flatter her? Nobody can tell her without measuring, and how much depends on how many visits each ward has: a ward with 40 visits leaks a lot into each row, a ward with 4,000 leaks almost nothing, and the same code gives both. If she wants the column, it is worked out on the training rows only, applied to the rest, and measured the Part 5 way over thirty splits before she believes a word of it.

    Mark yourself right only if you said the third one is built out of answers and that it is worked out across the whole table. Either fault alone is half.

  5. Anything time-shaped. Days since the campaign started, days between the previous campaign and this call, or whether the call came before or after a particular date. A tree cannot subtract two dates, so with a year in the file you would build the gap yourself and hand it over as one column, exactly as you did for the shops.
  6. Run it over thirty splits and see whether the average moves, and check whether the column would exist for a new row on the day you need the answer. One tells you whether it is real, the other tells you whether it is honest.
If you got the marked one wrong

Go back to Part 4 and read the two kinds of new column again, then read Step 4's Part 9 about the scaler.

Those three columns are the three mistakes people actually make with features, and you will be handed all three within a month of starting work.

StretchOptionalHarder

If you want more

Bronze

Run the four bank columns one at a time rather than all together: four loops of thirty splits, four averages against the 0.8902 to beat.

You should get 0.8882, 0.8902, 0.8901 and 0.8902. Then answer this: if one of them had come out at 0.8915, what would you have needed to do before believing it?

Silver

Take months_open away again and find out what the tree buys with extra rows instead. Split the shops the usual way first, so that you have 450 to train on and 150 held back. Then train an unlimited tree on the first 100 of the training rows, then 200, then 300, then all 450, scoring every one of them on the same 150.

Take the training rows out of the training half, not out of the whole table. Slicing the first 100 rows off the full 600 puts some of your held-back shops into the training set, which is Step 3's duplicate trap in a new coat, and your last score would be a memory test.

Write down the four scores and the four leaf counts. It climbs and then flattens out somewhere below 1.0, and the flattening is the point: one column got there with one question and 450 rows. Say in one line what a feature buys you, in the currency of rows.

Gold

Go back to the shop table, the two raw columns, no months_open, and train a different model on them:

from sklearn.linear_model import LogisticRegression

straight = LogisticRegression()
straight.fit(X_train, y_train)

print(round(straight.score(X_test, y_test), 4))
print(straight.coef_.round(3))

You have not met this model and Step 9 is where it belongs, so run it, do not study it. It scores 1.0 on shops it has never seen, from the two columns a tree needed twenty boxes for, and the two numbers it prints are close to equal and opposite.

Write the paragraph explaining why. What does one weight of about -1.7 and one of about +1.7 do to a pair of columns, and what does that have to do with the column you built by hand in Part 3? Then say what it costs: name one thing the tree could do on this table that this model cannot.

Before you move on

All six, or it is a not yet.

The thing you built.
Three columns for your own table, each with the sentence that justified it written before the code, and the thirty-split average for each.
The one you did not build.
A column your model could already express, named, with the reason you left it alone.
The one you cannot have.
A column your table would need and does not support, and what somebody would have to start recording.
Check yourself.
Five of six right, including the marked one about the hospital visits.
The build log.
One thing that surprised you, and one prediction you wrote down that turned out wrong. Parts 2, 3 and 5 all asked you to predict.
It still runs.
Kernel, then Restart Kernel and Run All Cells. Nothing should turn red.

Keep the sentence you wrote in Part 0, and the shop grid. Step 6 asks which of your columns are pulling their weight, and it starts by handing you a way to ask the model directly instead of guessing.