Step 3

Cheating, on purpose

Your Step 2 tree scored 0.9998 and could not be checked, because it had already seen everybody. Train the same kind of tree on three quarters of the people and it scores 0.7975 on the quarter it never met. Then you build three better lies on purpose.

About 3 hours 45 minutes11 partsBest over two sittingsThe one that matters
Read this first3 min

What this step gives you

Before anything else, find the note you wrote

In Part 5 of Step 2 you wrote down, in your own words, why you did not believe that 0.9998. Go and find it. Read it back to yourself now.

If you cannot find it, or you never wrote one, you can rebuild it in a minute. Here is the fact you were given: that tree had 709 endings for 4,521 people. Work out roughly how many people sit in each ending, and write two sentences saying what bothers you about a rule built for that few. Do that now, before you read another line.

Everything below either confirms what you wrote or corrects it. Both are worth a lot. Neither is worth anything if you cannot remember what you thought.

By the end of this step you will never again trust a score without asking one question first: which people was this measured on?

That sounds small. It is the difference between a model that works and a model that embarrasses you in front of the people who paid for it.

The thing that happens in Part 3

In Step 2 your unlimited tree scored 0.9998. It found 520 of the 521 people who said yes and wasted no calls at all.

You cannot check that tree. It was shown all 4,521 people, so there is nobody left to test it on. That is the first thing to understand: the problem is not that the number is high, it is that no honest number exists for it.

So in this step you keep a quarter of the people back and train the same kind of tree on the other three quarters. It scores 0.9997 on the people it learned from and 0.7975 on the quarter it never met.

Saying no to every single person would have scored 0.878. Your near-perfect model does worse than a rule a child could write on a napkin.

Then you cheat three times, on purpose

The first cheat is loud. You leave the answer sitting in the table and score a flat 1.0 on people the tree has never seen. Nothing goes red. Nothing warns you.

The second is quiet, and it is the one that costs real people their jobs. You add a column that is allowed to be there, that is not the answer, that no check will ever complain about, and your score goes up. You will only catch it by asking a question about the world, not about the data.

The third attacks the split itself. You put every person in the file twice, change nothing else, and watch the same broken tree look respectable. That one is the marked question at the end.

What you actually do, in order

  1. Memorise eight real people on paper, score full marks, then fail on four new ones.
  2. Cut the file in two and check the two halves.
  3. Watch the big tree fall.
  4. Watch the small tree hold.
  5. Leave the answer in the table, and see what a model does with it.
  6. Add a column that is not the answer, and score higher than you have ever scored.
  7. Enter every person twice, and score the same broken tree again.
  8. Split your own table and get your first honest number from it.
If you have not done Step 2

Do it first. This step starts from a tree you built and a number you did not believe. Without those, the whole thing is just words about splitting.

Part 0Unplug15 minAway from the screen

Memorise eight people

Paper and pen. Laptop shut.

Here are eight real people from the file. Two things about each, and whether they said yes.

AgeMoney in the bankSaid yes
29751no
31574no
37480yes
433215no
487195yes
662262no
67701yes
798556yes

First, try to find a rule

Spend three minutes looking for one. Is it the young ones? The rich ones? Read the ages down the page: no, no, yes, no, yes, no, yes, yes. Now read the money in order of size: 480 yes, 574 no, 701 yes, 751 no, 2262 no, 3215 no, 7195 yes, 8556 yes.

There is no single cut of either column that gets more than six of these eight right. Check that yourself before you believe it.

Now stop looking and just memorise

Take two minutes and learn the eight answers by heart. No, no, yes, no, yes, no, yes, yes.

Cover the table with your hand. Write the eight answers from memory.

You will get eight out of eight. Full marks. A perfect score.

Now the four you have not met

Four more real people from the same file. Write down yes or no for each, using whatever you have in your head.

  • Aged 34, with 215 in the bank.
  • Aged 23, with 4 in the bank.
  • Aged 56, with 3021 in the bank.
  • Aged 58, with 3382 in the bank.

Write your four answers down before you read on. Written down, on paper.

Open this once your four answers are on paper

The true answers are yes, yes, no, no.

You will probably have got about two of the four. Two out of four is what tossing a coin gets you on average, so two hours of memorising bought you nothing.

If you gave in and used a rule anyway, look at what happened. The best age rule you could have drawn from the eight, say yes above 31, gets one of the four. The best money rule gets one as well. Both do worse than the coin, because the pattern in the eight was never a pattern. It was eight people.

Write one line, and keep it

Your score on the eight was 100 percent. Your score on the four was 50 percent.

Both were real scores, honestly counted. Write down which of the two tells somebody what you are worth on the next person who walks in.

You have just done, with your hands, what your unlimited tree did in Step 2. It had 709 endings for 4,521 people. That is about six people per ending, and many endings held exactly one person.

An ending built around one person is not a rule. It is a note about that person.

Part 1Words5 min

Words to know

Six words. You will hear all six in any job interview for this work.

Training set
The rows you let the model learn from. Usually most of them.
Test set
The rows you keep back and never let it see. You only use them to measure.
Split
Cutting your rows into those two piles. In Step 2 this word meant one question in a tree. It means both things, and people rely on you to tell which from the sentence around it.
Overfitting
Learning the rows you were given so closely that you learn things which are only true of those rows. Memorising instead of understanding.
Leakage
Anything that lets information reach your model that it would not have in real use. Usually a column you will not have at the moment you need the answer. It makes your score go up and the score worthless, and the model with it, until you take the column out and measure again.
Honest score
Not a technical term. It is the number you would be willing to put in writing to somebody who is going to spend money on it.
The one sentence to keep

A score measured on rows the model learned from is not a measurement. It is a memory test, and the more a model is able to memorise, the less that number means. A model too small to memorise anything will score about the same either way, which is a fact about that model and not a reason to trust the practice.

Part 2Run20 min

Cut the file in two

Open band-03.ipynb. Every cell on this page is already in it, in this order. You run them; you do not type them. The only typing you do in this step is in Part 8, on your own table.

If you are coming back on a new evening

Run every cell from the top before you start. Nothing is remembered overnight, and the cells later in this step need names that were made earlier in it.

If you see NameError: name 'X_train' is not defined, that is all this is. Nothing is broken.

If JupyterLab is not open in your browser

It does not stay running between days. Open a terminal, move into the course folder, type jupyter lab and press Enter. Parts 6 and 7 of the setup page show you how.

If the first cell says FileNotFoundError

Jupyter is not standing in the course folder. Close it, move into the folder that holds data, and start it again from there. The path data/bank.csv is read from wherever Jupyter was started, not from where the notebook sits.

The first cell rebuilds everything you had at the end of Step 2: said_yes_before is 1 for the people who said yes to the last campaign and 0 for everybody else, X is the seven columns the model is allowed to look at, and y is the answer, 1 for yes and 0 for no.

Two things about the imports. The first line brings in two names at once, separated by a comma, instead of the two separate lines you wrote in Step 2. Same effect, less typing. The second line is new: scikit-learn keeps its tools in labelled drawers, and model_selection is the drawer for deciding which rows a model gets to see. That is where the splitter lives.

import pandas as pd
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split

df = pd.read_csv("data/bank.csv", sep=";")
df["said_yes_before"] = (df["poutcome"] == "success").astype(int)

cols = ["age", "balance", "day", "campaign",
        "pdays", "previous", "said_yes_before"]

X = df[cols]
y = (df["y"] == "yes").astype(int)

print(X.shape, y.sum())
You should see

(4521, 7) 521

Same table, same seven columns, same 521 people who said yes. You are starting from where Step 2 left you.

Now the cut.

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=0)

print(X_train.shape, X_test.shape)
print(y_train.sum(), y_test.sum())

Four new names on the left of one equals sign. train_test_split hands back four things at once, in that fixed order: the training columns, the test columns, the training answers, the test answers. Swap two of those names and everything below will run and be wrong, so copy the order exactly.

That line ends in an open bracket and carries on underneath. Python allows this: while a bracket is open, the line has not finished. The setup page warned you that a bracket you forget to close is an error, and it still is. A bracket you close on the next line is not. The indent on the second line is only there to make it easy to read.

test_size=0.25 means a quarter of the rows go into the test pile. random_state=0 means the rows are shuffled the same way every time you run it. That is why your numbers should match the ones on this page. If a future version of scikit-learn ever changes how it breaks a tie, they could drift by a digit, and the numbers in requirements.txt are the ones every figure here was produced with.

You should see

(3390, 7) (1131, 7), then 383 138

3,390 people to learn from. 1,131 kept back. Of those kept back, 138 said yes.

If your numbers are different from mine

You left out random_state=0. Without it the rows are shuffled differently every single run, so your score changes every time you press Shift and Enter, and no two people in the world can compare notes. Put it back and run the cell again.

Before you measure anything, work out what the score to beat is. On the test pile only.

print(round((y_test == 0).mean(), 4))

This is the lazy rule from Step 1, aimed at the test pile: say the commonest answer to everybody and never think again. (y_test == 0) is True for everybody who said no. Taking the mean of a column of True and False gives you the share that are True, because True counts as 1 and False counts as 0.

The , 4 inside round is how many figures after the dot you want to keep. Every score on this page is rounded to four, so if you drop it you will see a number twenty digits long and think something has gone wrong. Nothing will have.

You should see

0.878

Saying no to all 1,131 of them scores 0.878. That is your baseline, worked out on the test pile rather than on the whole file, and from here on this course calls it the bar. Same idea as Step 1, measured where it now has to be measured. Anything below it is worse than not bothering.

Part 3PredictRun25 min

The tree that knew everything

Now the thing this step exists for.

Predict first. Write two numbers.

You are about to train the same unlimited tree from Step 2, but only on the 3,390 training people. Then you score it twice: once on the people it learned from, once on the 1,131 it has never seen.

Write down both numbers before you run it. Be specific. A prediction of "lower" is not a prediction.

big = DecisionTreeClassifier(random_state=0)
big.fit(X_train, y_train)

print(round(big.score(X_train, y_train), 4))
print(round(big.score(X_test, y_test), 4))

Look at what changed from Step 2. You train on X_train, y_train. You score twice, and the second scoring hands it X_test, y_test, which the tree has never met.

You should see

0.9997, then 0.7975

Read those two numbers again.

On the people it learned from: 0.9997, which is as near perfect as makes no difference.

On the people it has never seen: 0.7975.

The bar was 0.878. Your model is a long way under it. On this measure, saying no to every single person, without a computer, without a course, without any of this, beats what you just built.

Now find out what that actually costs.

guess = big.predict(X_test)

print(((guess == 1) & (y_test == 1)).sum())
print(((guess == 1) & (y_test == 0)).sum())
You should see

25, then 116

Out of the 138 people in the test pile who would have said yes, it found 25. To find those 25 it told you to ring 116 people who said no.

Say this one out loud

In Step 2 a tree like this one found 520 of 521 and wasted nothing. This one finds 25 of 138 and wastes 116 calls.

Trees of this kind were always this bad at the job. Step 2 could not show you, because it had no honest people left to ask about. Nothing got worse between the two steps except your information.

Why it happened

The tree kept asking questions until every one of the 3,390 training people sat alone or nearly alone in their own ending. Print it if you like: the first five questions down the left-hand side are said_yes_before <= 0.50, then age <= 60.50, then pdays <= 401.00, then day <= 1.50, then balance <= 19.50. It goes 31 questions deep in places and finishes with 529 endings.

That wall is real. It is also about one person, not about people. The next person who arrives is on the wrong side of a wall that was built for somebody else.

This has a name and you now own it: overfitting. The tree learned the rows instead of the pattern.

Part 4RunInvestigate20 min

The small tree keeps its word

Same split. Same data. Two questions instead of thirty-one.

small = DecisionTreeClassifier(max_depth=2, random_state=0)
small.fit(X_train, y_train)

print(round(small.score(X_train, y_train), 4))
print(round(small.score(X_test, y_test), 4))
You should see

0.8941, then 0.893

The two numbers are almost the same. That tells you one specific thing: this tree is not memorising. It does about as well on strangers as it did on its own homework.

It does not tell you the tree is good, or honest, or worth shipping. Part 5 hands you a model with no gap at all between the two numbers that is completely worthless. A small gap rules out one fault. It is not a certificate.

It is also above the 0.878 bar, which the big tree was not.

guess_small = small.predict(X_test)

print(((guess_small == 1) & (y_test == 1)).sum())
print(((guess_small == 1) & (y_test == 0)).sum())
You should see

22, then 5

Put the two side by side

On the 1,131 people it never sawBig treeSmall tree
Score0.79750.893
Yes-people found, out of 1382522
Calls wasted1165

The big tree finds three more people. It burns 111 extra phone calls to do it.

If a call costs your bank five euros, the big tree spent 555 euros to find three extra customers. Somebody has to decide whether that was worth it, and that somebody is you, and you cannot decide it from a score alone.

One warning about those three, and it is the habit this whole step is for. Three is what this split gave you, and it happens to be the smallest gap of the thirty splits I ran: the usual gap is about thirteen customers for about a hundred wasted calls. The trade is the lesson. The exact price is not.

Write this in your README before you go on

One line: which of the two you would hand to the bank, and what the extra three customers cost.

There is a defensible answer either way. What is not defensible is picking the bigger score without knowing what it bought.

Part 5PredictModify20 min

Cheat one: leave the answer in

You have just seen a split catch a bad model. Now watch a split fail to catch a worse one.

You are going to put the answer itself into the table of columns, as if you had copied it in by mistake. People do this every week, and it is one of the two standard ways a real project produces a beautiful number that means nothing. Part 6 has the other one.

Predict first. Write two numbers.

The tree gets a column that is literally the answer. It is trained on 3,390 people and scored on 1,131 it has never seen.

What does it score on the people it learned from? What does it score on the strangers? Write both down.

X_leak = X.copy()
X_leak["answer"] = y

L_train, L_test, y_train, y_test = train_test_split(
    X_leak, y, test_size=0.25, random_state=0)

leaky = DecisionTreeClassifier(max_depth=2, random_state=0)
leaky.fit(L_train, y_train)

print(round(leaky.score(L_train, y_train), 4))
print(round(leaky.score(L_test, y_test), 4))

X.copy() again, so that the table you have been using is left alone. Then one new column called answer, holding exactly the thing you are trying to guess.

In the printing line further down you will see cols + ["answer"]. A plus between two lists joins them end to end and hands back a new list of eight names. It does not change cols, which still holds seven and is used again in Part 7.

The split runs with the same random_state=0, so the same people go to the same side as before. y_train and y_test come out holding exactly what they held before. That is the whole point of random_state.

You should see

1.0, then 1.0

A flat, perfect score on 1,131 people it has never met.

Nothing turned red. No warning appeared. The split, the thing that caught the bad model in Part 3, did not make a sound here.

Now look at it

print(export_text(leaky, feature_names=cols + ["answer"]))
You should see
|--- answer <= 0.50
|   |--- class: 0
|--- answer >  0.50
|   |--- class: 1

One question. It asks whether the answer is the answer.

You gave it seven columns of real information and one column of the answer, and it threw the seven away. Which is what it is built to do. It tries every cut of every column and keeps the best one, and nothing beats the answer.

This is the lesson, and it is not about this cell

Nobody types X["answer"] = y on purpose. It happens because somebody joined two files on a customer number, and the second file already had the outcome in it, and it arrived under a name like status_final or closed_flag.

It will not look like the answer. It will look like a column.

The tell is the score. When a model you did not expect much from returns 1.0, or 0.99, your first thought is not "it worked". Your first thought is "what did I hand it".

How you catch it

Two habits, and you already have both of them.

  1. Print the tree, the way you just did. If one column is doing all the work, look hard at that column.
  2. Take that column out and run it again. Taking a column out means making a shorter list and passing that instead, the mirror of the plus you just met: seven_minus = ["age", "balance", "day", "campaign", "pdays", "previous"], then train on X[seven_minus].
What a cliff does and does not prove

If the score collapses when you remove one column, you have learned that the column carries almost everything. That is a reason to look at it. It is not a verdict.

A column can be the most useful thing in your table and still be perfectly honest. The most useful single question you have is said_yes_before. It is the first question every tree in this course has asked, on all thirty splits I tried, and there is nothing wrong with it, because you know who said yes to the last campaign long before you pick up the phone today.

So the cliff tells you where to look. The next part tells you what to look for.

Part 6ModifyInvestigate20 min

Cheat two: the one that does not look like cheating

That last one was easy to spot once you looked. This one is not, and this one is the reason this step is in the course.

The file has a column that has never been in your seven: duration. It is how many seconds the phone call lasted. If you did Step 1's Silver task you met it once and were asked to write down why the bank could never use it. This is that, with numbers.

It is not the answer. Nobody copied it in by mistake. It was collected honestly, it is in the file the bank shipped, and every check you have ever run would let it through.

cols_d = cols + ["duration"]

D_train, D_test, y_train, y_test = train_test_split(
    df[cols_d], y, test_size=0.25, random_state=0)

deep = DecisionTreeClassifier(max_depth=3, random_state=0)
deep.fit(D_train, y_train)

print(round(deep.score(D_test, y_test), 4))

guess_d = deep.predict(D_test)
print(((guess_d == 1) & (y_test == 1)).sum())
print(((guess_d == 1) & (y_test == 0)).sum())

cols + ["duration"] makes a new list of eight names without touching cols. max_depth=3 gives it one more question than the small tree had, which it needs to use the new column properly.

You should see

0.9072, then 64, then 31

0.9072 on strangers. That is the best honest-looking number in this course so far. It beats the 0.878 bar properly, and it finds 64 of the 138 instead of 22.

Being fair to the thing you are about to disqualify: that is the friendliest of thirty splits for it. Over all thirty it averages 0.8933 and 41 catches, so the honest headline is "about twenty more customers", not forty-two.

It is not free. Wasted calls go from 5 to 31, so by the five-euro arithmetic in Part 4 you spent 130 euros to find 42 more customers on this split. That is a good trade, and you should be able to say so with the numbers in your hand rather than by pointing at the score. In a real job you would be tempted to send this in an email with an exclamation mark in it.

Do not run anything. Answer this.

The bank wants to use your model on Monday morning to pick who to ring.

It is Monday morning. You have not rung anybody yet. What is the value of duration for the person you are about to ring?

There isn't one. The call has not happened. Its length is not a fact about the customer, it is a fact about a call you have not made.

Long calls and yes answers do go together. Which one causes the other is a good argument to have over lunch and it does not matter here. What matters is the order in time: the length exists only after the call, and you have to choose who to ring before it.

The column is real. The number is real. The model is worthless, because on Monday morning the column is empty for every person you care about.

The question that catches this, and the only one that does

For every column: would I know this, for a new person, at the moment I have to answer?

Not "is it in the file". Not "is it allowed". Not "does it help". Would I have it, then, for somebody I have not dealt with yet.

No script can answer that for you. It is a question about how the work happens, and you get the answer by asking whoever does the work.

This is leakage, the fifth word in Part 1, and now it has a face.

Nobody hid it. Open data/bank-names.txt, the notes that came with the file, and just above the list of columns there is a heading: related with the last contact of the current campaign. Four columns sit under it. Whoever wrote that heading was telling you exactly which columns belong to the call itself. That is why Step 0 made you open the notes.

Now do it to yourself. This one is not rhetorical.

Four columns sit under that heading. One of them is duration. One of the other three has been in the seven you have been calling honest all step.

Open the notes, find it, and ask the Part 6 question of it. Write down what you decide and why, before you start Step 4.

Then do the same for campaign, which is not under that heading but whose own description gives it away if you read to the end of the line.

There is a defensible answer for both, and it is not the same answer as for duration. What is not defensible is using them for three steps without noticing. Nor did I, until somebody read this page looking for exactly this.

Write this in your README

Three lines. What duration did to the score, the sentence you would say to a manager who saw 0.9072 and asked why you are not shipping it, and what you decided about the other two columns.

Part 7PredictModify25 min

Cheat three: the same person twice

One more, and this one attacks the split itself.

First, check the file you have.

print(df.duplicated().sum())

.duplicated() marks every row that is an exact copy of a row further up. Summing it counts them.

You should see

0

Clean. Not one repeated row in 4,521. Real files are rarely this tidy, so you are going to break it on purpose.

twice = pd.concat([df, df], ignore_index=True)

print(twice.shape)
print(twice.duplicated().sum())

pd.concat stacks tables on top of each other. Handing it [df, df] stacks the file on itself, so every person now appears twice. ignore_index=True renumbers the rows from 0 so there are no repeated row numbers.

You should see

(9042, 18), then 4521

Predict first. One number.

Now the unlimited tree again, the one that scored 0.7975 on strangers. Same settings, same quarter kept back. The only change is that every person is in the file twice.

What does it score on the test pile now?

X2 = twice[cols]
y2 = (twice["y"] == "yes").astype(int)

X2_train, X2_test, y2_train, y2_test = train_test_split(
    X2, y2, test_size=0.25, random_state=0)

big2 = DecisionTreeClassifier(random_state=0)
big2.fit(X2_train, y2_train)

print(round(big2.score(X2_train, y2_train), 4))
print(round(big2.score(X2_test, y2_test), 4))
You should see

0.9999, then 0.9447

0.7975 became 0.9447. The tree is exactly as bad as it was. The data is exactly the same data. Nobody added a single new fact.

There is a tool for exact copies, and it is worth meeting now.

print(twice.drop_duplicates().shape)

.duplicated() marks the copies. .drop_duplicates() throws them away and hands back the table without them. It matches on the whole row: every value in every column has to be identical.

You should see

(4521, 18)

Back to the file you started with. That is the fix, when the copies are exact.

This does not happen to every model

Run the small tree on the doubled table and watch what it does. Predict first: up, down, or unmoved?

small2 = DecisionTreeClassifier(max_depth=2, random_state=0)
small2.fit(X2_train, y2_train)

print(round(small2.score(X2_test, y2_test), 4))
You should see

0.8859

It went down, from 0.893.

Do not read too much into that one number, and you already know why. Run the same comparison over thirty splits and the small tree gains nothing on average: it goes up on fifteen of them and down on the other fifteen. The big tree goes up on all thirty, by about fourteen points every time.

So the rule is not "duplicates raise your score". It is narrower and more useful than that. Copies pay off only for a model that can memorise, because memorising is the only way to profit from seeing the same person twice. The bigger the model, the more it gains. Yours gained fifteen points and went from useless to respectable-looking.

What happened

The split shuffles rows, not people. When a person appears twice, one copy can land in the training pile and the other in the test pile.

The tree memorises the copy it trained on. Then you test it on the twin, and it recognises it perfectly, because it has seen that exact person before.

About three quarters of the test pile has a twin sitting in the training pile. Your test set is not a test. It is most of the homework again, handed back in a different order.

The rule you take away

A split is only honest if a person cannot be on both sides of it.

Rows are not people. Split by whatever a person is in your table: a customer number, a household, a patient, a school. Doing that properly needs tools you do not have yet, so for now the job is to notice when you are exposed, and write it down. Step 14 gives you the tools, and Step 17 handles the version of this trap that time creates.

Part 8Make40 minNo answer given

Split your own table

No code here and no answer at the end. Use the table you signed up in Step 0 and modelled in Step 2.

Use new names

mine, not df. my_X and my_y, not X and y. If you reuse the names above, the cells you already ran will still work and print numbers that look fine and are about the wrong table.

Your own file is not in the data folder, so the path is not data/bank.csv. Put a copy of your file into that folder and read it as data/your-file.csv. That is the least fiddly way, and paths are Step 20's problem, not today's.

The brief

Split your own table, train your depth 2 tree on the training half only, and get your first honest number.

It is done when

If your two scores are almost the same and both are poor

That is not overfitting, and splitting will not fix it. It is one of two things. Either the columns you have do not say enough about the answer, or the model is too small to use what they say. Try an unlimited tree: if it too is poor on both halves, it is the columns. Either way it is a real finding and worth a line in your README, and Step 5 is where you learn to build the column your table is missing.

If your test pile has almost no yes-people in it

With a small table, a quarter of a rare answer can be five or six people, and a score built on six people moves wildly. Note the number in your README. Step 7 is where you learn the right way to score a rare answer, and it starts from exactly this complaint.

Working alone

Somebody who works in the area your data comes from is the person to ask: a nurse for hospital data, a pharmacist for prescriptions, a shopkeeper for sales, a teacher for school data. If your table is about work you do, that person is you, and this is one of the few advantages you have over somebody who came to this from computing.

Go down your column list and ask one question of each: would I know this before the thing happens? Write the answers down in your own hand. You are looking for the one you hesitate over, and the hesitation is the finding.

If you know somebody else who works with this data

Ask them the same question, column by column, and write down what they say rather than what you expected them to say.

Part 9Ship15 min

Write it up

Add to the README you started in Step 0.

Then say it to somebody

Find a person who does not code. Tell them, in under a minute, how a computer can score 100 percent and still be useless.

If they say "like cramming for an exam", you have explained it. If they go quiet, you have not, and the fix is Part 0: tell them about the eight people.

Part 10Check15 min

Check yourself

Six questions. Answer them in writing before you open the answers.

One counts more than the other five. You can get five right and still not pass this step.

  1. Your tree scores 0.9997 on the training pile and 0.7975 on the test pile. Which of those two numbers do you put in an email, and why the other one is worthless.
  2. The small tree found 22 of 138 and the big tree found 25. Why is the small tree still the better model?
  3. You add a column and your test score jumps from 0.89 to 0.99. Name the first two things you do, in order.
  4. A colleague splits her rows, trains on the training rows, and scores 0.91 on the test rows. She shows you the code and the split is done correctly.

    Then she finds that every customer was entered into the system twice, once with their name misspelled, so the two rows are not exact copies of each other.

    Is her 0.91 still honest? Say what is happening to it, and say what she should do. Then say why drop_duplicates() does not fix it.

  5. duration is a real column, honestly collected, and it raises the score. In one sentence, what is wrong with using it?
  6. You split by rows. Give one example, from any kind of work, where a person can end up on both sides of that split.
Open the answers, once all six are written down
  1. The 0.7975. It is the only one measured on people the tree did not learn from. The 0.9997 is a memory test: it tells you the tree can recall its own homework, which nobody is paying for.
  2. Because it found nearly as many for 5 wasted calls instead of 116. Twenty-two customers for five wasted calls is a business. Twenty-five for 116 is a phone bill. The small tree is also not memorising, which its two nearly equal scores tell you, so at least it is doing on strangers what it did in training. Whether it still works next month is a different question, and no split you have made can answer it.
  3. Print the model and see which column is doing the work. Then take that column out and run it again. If the score collapses, that column is carrying nearly everything, which tells you where to look and nothing more. The verdict comes from asking whether you would have that column, for a new person, at the moment you have to answer. A column can carry nearly everything and be perfectly honest.
  4. No, it is not honest, and it is inflated for exactly the reason yours was in Part 7. The same customer is on both sides of the split. The model memorises one copy and recognises the twin, so a chunk of her test pile is homework she has already marked.

    drop_duplicates() looks for rows that match exactly. Her two rows do not match, because one name is misspelled. It will delete nothing and report nothing, and she will believe her data is clean.

    She has to split on the person, not the row: work out who is who first, using whatever identifies a customer, then send all of one person's rows to the same side. Mark yourself right only if you said the misspelling is what defeats the exact-match check. Saying "her data is dirty" is half.

  5. You will not have it when you need it. On the morning you have to choose who to ring, no call has happened, so the column is empty for every person you care about.
  6. Any answer where one real thing produces several rows. A patient with several visits. A household with several bank accounts. A shop with a row per week. A student sitting the same exam twice.
If you got the marked one wrong

Go back to Part 7 and run it again, and this time write down what the split actually shuffles.

Then read the answer, write it out again tomorrow from memory, and go on. That counts.

Every step from here on measures itself on a test pile. If you do not fully believe that a test pile can lie to you, every number in the rest of this course is decoration.

StretchOptionalHarder

If you want more

Bronze

Train trees at depth 1, 2, 3, 5, 10 and unlimited. For each one write two numbers: the training score and the test score.

One column climbs the whole way. The other climbs and then turns round. Find where it turns, and write down what that turning point is worth to somebody choosing a model.

Silver

Run the Part 2 split five times with random_state set to 0, 1, 2, 3 and 4, and score the depth 2 tree on the test pile each time.

Write down the five numbers and the gap between the largest and the smallest. Then answer this: if you had run it once and got the highest of the five, and put that in a report, what exactly would you have done wrong?

Gold

Go back to the duplicates in Part 7, but instead of copying the whole file, copy only the 521 people who said yes.

Split, train the unlimited tree, and write down the test score and the number found. Then explain what happened, and why somebody trying to help a model find rare answers might do this without realising what it does to their score.

Before you move on

All seven, or it is a not yet.

The thing you built.
Your own table split, with both trees scored on both halves, four numbers in your README.
The column list.
Every column in your own table, marked for whether you would have it at the moment you need the answer, and what happened when you removed the ones you marked.
The comparison.
Big tree against small tree on the test pile: score, found, wasted, for both.
Check yourself.
Five of six right, including the marked one about the misspelled names.
The build log.
One thing that broke or surprised you, and one prediction you wrote down that turned out wrong. Part 5 and Part 7 both ask you to predict; at least one of those predictions was wrong.
It still runs.
Kernel, then Restart Kernel and Run All Cells. Nothing should turn red.
Ninety seconds, out loud.
Record yourself explaining why a score of 1.0 is a reason to worry. Every step ends this way, and this is the one an interviewer is most likely to ask you.

Keep your Part 4 table. Step 7 comes back to it and shows you that both of those scores were answering the wrong question.