Step 2

Rules the computer works out

You hand it seven columns and let it find its own rule. What it finds first is the one you scored by hand in Step 1.

About 3 hours 30 minutes10 partsBest over two sittingsYour first model
Read this first3 min

What this step gives you

By the end you will have built a model, drawn it on paper, and been able to say exactly what it does. Not roughly. Exactly, question by question, the way you would explain a form.

That matters more than it sounds. Most people who use models cannot say what theirs does. They can say what it scored. You are going to be able to point at yours and read it out loud.

The thing that happens in Part 3

In Step 1 you read a rule, scored it yourself, and got 0.8929, with 83 people caught and 46 calls wasted. Step 1 said at the time that it was a rule you could have thought of.

In this step you give a computer seven columns and no rule to follow, and let it work one out.

It picks yours. The same rule, the same 83 people, the same 46 wasted calls. Not close to it. Identical.

Then Part 3 asks you to notice the catch in that, because one of the seven columns is your Step 1 rule with the words taken out, and you are the one who put it there. What the computer did was confirm your rule was the best of the seven. What it could not do is find it in two questions, and the Silver task at the bottom has you watch that happen.

That is the best possible introduction to a model, and it is why Step 1 came first. A model is not a wiser thing than you. It is a very fast, very patient thing that tries every split of every column and keeps the best one.

What you actually do, in order

  1. Sort twenty objects by asking yes and no questions, on a table, with your hands.
  2. Turn your Step 1 rule into a column of ones and zeros.
  3. Let the computer choose its own rule, with no help, and see what it picks.
  4. Let it go one level deeper, and find something you missed.
  5. Let it go all the way down, and read the score it gives you very carefully.
  6. Change the units of a column, retrain, and find out something surprising about what a tree really is.
  7. Try to hand it a column of words, and get told no. That refusal is the whole of Step 4.
How the time actually goes

Part 4 is one idea and wants half an hour together: the middle of it shows you a group of people the tree says no to for a reason you cannot see yet, and the end of it is the reason. If you have less than half an hour left, stop at the end of Part 3.

In three-quarter-hour evenings: Parts 0 to 2, then Part 3, then Part 4, then Parts 5 and 6, then Part 7 on its own.

Before you start

Step 1 finished, including your own table. You will use the same file and the same laptop set-up.

One thing you will notice. The files are called band-00.ipynb, band-01.ipynb and so on. The steps were called bands while this was being written and the file names kept it. Same thing, and nothing you have done changes.

Part 0Unplug15 minAway from the screen

Sort twenty things by asking questions

Read this whole part first. Then do it on a table, with real objects. It takes about ten minutes and it is the only explanation of a tree you will need.

Get twenty things

Anything, as long as they are different from each other. Things off your desk. Things out of a kitchen drawer. Coins, keys, pens, fruit, cups, tins.

Now pick something you want to sort them into. Two piles. Things I would put in a bag versus things I would not. Things that are mine versus things that are not. It does not matter, as long as you know the answer for all twenty.

Then do this

  1. Push them into one heap. Look at the heap and think of a single yes or no question you could ask about any one object. Is it made of metal. Is it bigger than my hand. Does it have writing on it.

    Write your question on a piece of paper.

  2. Ask your question of every object and push the heap into two smaller heaps, yes and no.

    Now look at your two heaps. You want each one to be as pure as possible. A heap that is all one answer is perfect. A heap that is half and half taught you nothing.

  3. Now do each heap again, separately. Pick a new question for the left heap, and a different new question for the right heap. Split each of them into two.

    Write both questions on your paper underneath the first, joined to it by lines.

  4. Stop there. Three questions in total: one at the top and two underneath it. Do not go further, even though you will want to.

  5. Look at what you have drawn. One question at the top, splitting into more questions, ending in heaps. Turn the paper so it hangs downwards.

You should see

A drawing that branches: one question, then two, then four. And at the bottom, four heaps of objects.

Count each heap. Write two numbers on each: how many things are in it, and how many of those were the answer you were looking for.

What you just built

That drawing is a decision tree. It is upside down, with the root at the top, which is how everybody draws them and nobody explains why.

You chose your questions by looking at the heap and using judgement. In Part 3 a computer will choose its questions by trying every possible one and measuring which leaves the tidiest heaps. That is the main difference between you and it, and there are others worth knowing: it can only ask whether a number is below a cut, it never goes back and revises its first question after seeing where the second one led, and it cannot invent a question you did not give it a column for. Steps 3 and 5 are largely about those three.

Keep the paper. Part 4 asks you to draw a second tree next to it.

Part 1Words5 min

Words to know

Six words. Read them once, then write each in your own words in your notebook.

Model
A set of rules a computer worked out from examples, instead of a person writing them.
Train
Show the computer the examples and let it work the rules out. It is also called fitting.
Split
One yes or no question that cuts a heap into two.
Leaf
A heap at the bottom with no more questions under it. Every leaf gives one answer.
Depth
How many questions get asked before you reach a leaf. Your paper tree from Part 0 has a depth of two.
Cut
The number a split turns on. In balance <= 10537.50 the split is the whole question, and 10537.50 is the cut. Some people say threshold. Same thing.
Part 2Run20 min

Turn your rule into a column

Open band-02.ipynb, in the course folder. Run the cells from the top.

If you are coming back on a new evening

Run every cell from the top before you start. Nothing is remembered overnight, and the cells later in this step need names that were made earlier in it.

If you see NameError: name 'X' is not defined, that is all this is. Nothing is broken.

If your numbers are different from mine

Check that every random_state=0 on the page is in your cell too. Without it the rows are shuffled differently every run, so your numbers change every time you press Shift and Enter and no two people in the world can compare notes.

If the numbers still differ, run every cell from the top in a fresh kernel: Kernel, then Restart Kernel and Run All Cells.

If JupyterLab is not open in your browser

It does not stay running. Open a terminal, move into the course folder, type jupyter lab and press Enter. Parts 6 and 7 of the setup page show you how.

One new thing to install, once

The setup page installed this for you, so most likely you already have it and the command below will just say so. Open a second terminal, move into the course folder the way setup Part 6 shows, and run it. Use python instead of python3 on Windows.

python3 -m pip install scikit-learn

You should see a line starting Successfully installed, or Requirement already satisfied, which is just as good. Then in JupyterLab choose Kernel, then Restart Kernel.

The name you install is scikit-learn. The name you type to use it is sklearn. Two names, one thing, and nobody ever warns you.

If you get ModuleNotFoundError: No module named 'sklearn'

The install did not take, or your kernel is still the old one. Put %pip install scikit-learn in a cell and run it. The % matters. Then Kernel, Restart Kernel, and run from the top.

  1. Load the file, the same way as before.

    import pandas as pd
    
    df = pd.read_csv("data/bank.csv", sep=";")
    
    print(df.shape)
    You should see

    (4521, 17)

  2. Turn your Step 1 rule into a column.

    In Step 1 the rule you scored was df["poutcome"] == "success", which gives a column of True and False. A tree wants numbers, so put .astype(int) on the end. That turns every True into 1 and every False into 0.

    df["said_yes_before"] = (df["poutcome"] == "success").astype(int)
    
    print(df["said_yes_before"].value_counts())

    The left hand side is new. df["said_yes_before"] = ... makes a brand new column in the table and fills it.

    You should see

    0 4392 and 1 129

    129 ones. Those are the same 129 people your hand rule said yes to in Step 1.

  3. Choose the columns the tree is allowed to see.

    cols = ["age", "balance", "day", "campaign",
            "pdays", "previous", "said_yes_before"]
    
    X = df[cols]
    y = (df["y"] == "yes").astype(int)
    
    print(X.shape, y.sum())

    Three things here.

    • df[cols], with a list inside the brackets, gives you a smaller table: just those seven columns, all 4,521 rows. One name in brackets gives one column. A list of names gives a table.
    • X is what the model may look at. y is the answer you want it to guess. Everybody uses those two letters, so you may as well start now.
    • y holds 1 for yes and 0 for no, made with the same .astype(int) you used a moment ago.
    You should see

    (4521, 7) 521

    Seven columns to look at, and 521 real yes answers to find.

    Two honest notes about those seven columns

    pdays still has its 3,705 minus ones in it, from Step 0. It does not choke a tree the way it choked an average, because all a tree ever asks is whether a number is below some cut. That is not the same as harmless, and Step 4 goes back and deals with it properly. Step 4 deals with them properly.

    duration is missing on purpose. It is the length of the call, so you only know it once the call has ended. Step 1's Silver task, if you did it, asked you to write down why the bank could never use it. Step 3 puts numbers on it, and Step 6 comes back to it.

Part 3PredictRun25 min

Let it choose, and watch what it chooses

You are going to allow the tree exactly one question. One split, then it must answer. It has seven columns to choose from and it may cut any of them anywhere.

Predict first. Write it down.

Which of the seven columns will it pick, and where will it cut?

Then write down what you think it will score. Your hand rule got 0.8929. The lazy rule got 0.8848.

from sklearn.tree import DecisionTreeClassifier

tree = DecisionTreeClassifier(max_depth=1, random_state=0)
tree.fit(X, y)

print(round(tree.score(X, y), 4))

Four new things, one line each.

You should see

0.8929

Look at that number. Then look at Step 1's table, where you wrote your hand rule down.

Now find out exactly what it built.

from sklearn.tree import export_text

print(export_text(tree, feature_names=cols))

export_text prints the tree as text. feature_names=cols hands it your column names, and without it the tree would call them feature_0, feature_1 and so on. It draws the tree lying on its side, using |--- for a branch and indenting one level per question.

A line saying class: 0 is a leaf, and the number is the answer given there. 0 means no and 1 means yes, because that is what your y column holds.

You should see

Two branches. said_yes_before <= 0.50 gives class 0, and above it gives class 1.

In plain words: if they said yes last time, say yes. Otherwise say no.

Now count what it caught. tree.predict(X) runs all 4,521 rows down the tree and hands back one answer each, a 1 or a 0, in the same order as the table. So guess.sum() is how many it said yes to, exactly the way rule.sum() worked in Step 1.

guess = tree.predict(X)

print(guess.sum())
print(((guess == 1) & (y == 1)).sum())
print(((guess == 1) & (y == 0)).sum())
You should see

129, then 83, then 46

Stop and take this in

Says yes to 129. Catches 83. Wastes 46 calls. Scores 0.8929.

Those are your four numbers from Step 1. Not close to them. The same four.

It tried every cut of every one of the seven columns and landed exactly where you landed by thinking. A model is not cleverer than you. It is faster and more patient, and it never gets bored halfway through.

Write one line in your notebook: how does that make you feel about the rule you wrote in Step 1?

Now the catch, and it is the more useful half

Go back and look at what you handed over in Part 2. One of those seven columns is said_yes_before, and you built it yourself, out of poutcome, using your Step 1 rule.

So the tree did not go and find your rule in the raw data. It was given your rule, as a column, along with six others, and it worked out that yours was the best of the seven. That is a real and useful thing to have learned, and it is not the same thing as discovering it.

Take that column away and a tree with two questions finds nothing at all. Not a worse rule: nothing. It says no to all 4,521 people and scores 0.8848, which is the lazy rule from Step 1. The Silver task at the bottom has you run it.

Two questions is the limit that matters, not the tree. Give the same six columns a third question and it starts finding people the hard way: 66 of them, at 0.8887. It got there without you, and it needed a bigger tree to do it. That trade is the whole of Step 3.

The reason is the whole of Step 4. The column it needed was poutcome, and poutcome holds words. You did the one piece of work by hand that a tree cannot do at all, in one line, in Part 2, and it looked like nothing.

One cell in this notebook is wrong on purpose

Somewhere below, a cell works out a score and prints 0.0. It does not crash. A score of zero would mean the model got every single row wrong, which cannot be true of a model that just scored 0.8929.

Find it. It is comparing two things that are not the same kind of thing. Write one line on what those two things are, then fix it so it agrees with the 0.8929 above.

The fix is to change one name. The mended cell still ends in .mean() and prints 0.8929. Replacing the whole thing with tree.score(X, y) does not count, because the point is seeing what the two sides were.

One thing you have not met: .mean() on a column of True and False gives the share that are True, because True counts as 1. So (a == b).mean() asks what fraction of the time a and b agreed. That is another way of writing a score, and you will use it constantly.

Part 4ModifyInvestigate30 min

Give it one more question

Predict first. Write two numbers.

You are about to allow it two questions instead of one. What will the score become, and how many of the 521 will it catch?

tree2 = DecisionTreeClassifier(max_depth=2, random_state=0)
tree2.fit(X, y)

print(round(tree2.score(X, y), 4))
print(export_text(tree2, feature_names=cols))
You should see

0.8938, then a tree with four endings:

  • said_yes_before <= 0.50 then age <= 60.50 gives class 0
  • said_yes_before <= 0.50 then age > 60.50 gives class 0
  • said_yes_before > 0.50 then balance <= 10537.50 gives class 1
  • said_yes_before > 0.50 then balance > 10537.50 gives class 0

Read what it found

The second question on the yes side is about money. Among the 129 people who said yes last time, it has pulled out the ones with more than 10,537 euros in the bank and decided they will say no.

Check whether that was worth doing.

guess2 = tree2.predict(X)

print(guess2.sum())
print(((guess2 == 1) & (y == 1)).sum())
print(((guess2 == 1) & (y == 0)).sum())
You should see

125, then 83, then 42

It still catches all 83. It has stopped ringing 4 people who were never going to say yes. It kept every win and dropped four losses.

That is a real improvement and it is a small one. Write both numbers down. Nine ten-thousandths of a score, and four fewer wasted phone calls. Whether that is worth anything depends on what a phone call costs, which is the argument you had with yourself in Step 1.

The split that changed nothing

Look at the other side of the tree again. On the no side it asked whether age is over 60.5, and then answered no on both branches.

So why did it bother? Find out how many people are in each of those two heaps.

older = (df["said_yes_before"] == 0) & (df["age"] > 60.5)

print(older.sum())
print(y[older].sum())

older is a column of True and False, one per row, built the same way as your rules in Step 1. y[older] means take the answers column and keep only the rows where older is True. It lines up because y and df are still in the same row order.

You should see

110, then 38

110 people, of whom 38 said yes. That is more than one in three, in a file where the overall rate is about one in nine.

Count all four boxes

That was one box out of four. Here is the whole set. You will need this exact shape again on your own table in Part 7, so read it rather than just running it.

yes_before = df["said_yes_before"] > 0.5
old = df["age"] > 60.5
rich = df["balance"] > 10537.5

print((~yes_before & ~old).sum(), (~yes_before & old).sum())
print((yes_before & ~rich).sum(), (yes_before & rich).sum())

Each line of the tree becomes one True and False column, and then & and ~ combine them into the four boxes, exactly the way you counted catches and wasted calls in Step 1.

You should see

4282 110, then 125 4

Four boxes. Add them up: they must come to 4,521, because every person is in exactly one box.

This is worth understanding properly

The tree found a group that is three times more likely than average to say yes, and then said no to all of them anyway.

It did that because 38 out of 110 is still fewer than half, and this tree gives each leaf whichever answer is commonest in it. It is being asked for one answer per heap, so it gives the safe one.

That is not a fault in the tree. It is a fault in the question we asked it. Step 9 is where you learn to ask for a chance instead of an answer, and Step 7 is where you learn to move the line. Write this leaf down. You will come back to it.

Draw it

On the same paper as your object tree from Part 0, draw this one by hand.

You should see

The two boxes under said_yes_before <= 0.5 holding 4,282 and 110, so one of them is enormous and one is small. The two on the other side holding about a hundred between them, and the box under balance > 10537.5 holding fewer than ten.

All four adding to 4,521. And only one of the four boxes saying yes.

If your two big heaps are on the right hand side, your branches are the wrong way round. <= 0.5 is the people who did not say yes before.

Now put it beside the tree you drew with your hands in Part 0. Same shape: one question, two questions, four boxes.

Part 5PredictModify20 min

Take the lid off

Two questions gave 0.8938. Now take the limit off entirely and let it ask as many questions as it likes.

Predict first. Write one number.

What will it score with no limit at all?

big = DecisionTreeClassifier(random_state=0)
big.fit(X, y)

print(round(big.score(X, y), 4))
print(big.get_n_leaves())
print(big.tree_.max_depth)

Look at what is not in that first line. Last time you wrote max_depth=2. Leave it out and there is no limit on the depth: it keeps asking questions until every ending holds people who all gave the same answer, or holds a single person. Leaving a setting out does not mean off, it means use the usual, and here the usual is as deep as you like.

Two more new things. get_n_leaves() has brackets because it is a job you are asking the tree to do, counting its endings. tree_.max_depth has none, because it is not a job, it is a fact the tree wrote down while it trained. This is the mirror of Step 0's planted bug, where brackets were missing from a job and you got a <bound method line instead of a number.

Do not turn that into a rule about which things are jobs. Whoever wrote the library decided, and not always consistently: this same tree gives you its number of endings as tree_.n_leaves with no brackets, and its depth as get_depth() with them. The rule worth keeping is narrower and never fails: if what comes back starts with <bound method, you left the brackets off.

You should see

0.9998, then 709, then 30

guess3 = big.predict(X)

print(((guess3 == 1) & (y == 1)).sum())
print(((guess3 == 1) & (y == 0)).sum())
You should see

520, then 0

It found 520 of the 521 people who said yes, and it did not waste a single phone call.

Sit with that for a moment. It is nearly perfect. Every hard thing this course has told you is difficult, it has just done, in under a second, on a laptop.

Something is wrong and you should be able to feel it

Write down, in your own words, why you do not believe that 0.9998.

Do not look anything up. Do not ask anybody. Just write what bothers you. One or two sentences.

Here is one clue, and it is the only one you get. The tree has 709 endings for 4,521 people. Work out roughly how many people are sitting in each ending.

Keep what you wrote. The first page of Step 3 asks you to read it back.

Part 6Modify20 min

Change the money and see what happens

Your two-question tree cuts the money column at 10,537.50 euros. Suppose the bank had recorded every balance in cents instead of euros, so every number is a hundred times bigger. Nothing about any person has changed. Only the units.

Predict first, and be honest about it

Train the same tree on that table. Write down what happens to the tree, and what happens to its score.

Most people get this wrong, and getting it wrong here costs you nothing.

X_cents = X.copy()
X_cents["balance"] = X_cents["balance"] * 100

cents_tree = DecisionTreeClassifier(max_depth=2, random_state=0)
cents_tree.fit(X_cents, y)

print(round(cents_tree.score(X_cents, y), 4))
print(export_text(cents_tree, feature_names=cols))

Two new pieces in there. X.copy() makes a separate table holding the same things, so that changing the copy leaves X alone. Without it you would quietly damage the table Parts 3 to 5 were built on, and nothing would tell you. And X_cents["balance"] * 100 multiplies all 4,521 balances at once. You do not have to go round them one at a time.

You should see

0.8938, and exactly the same tree, except that the money cut now reads 1053750.00 instead of 10537.50.

Same questions, same order, same answers, same score. The cut moved to the same place in the new units.

Why nothing moved

A tree never adds, multiplies or compares one column against another. The only thing it ever asks is whether one number is below some cut.

Multiplying a column by a hundred does not change which people are below which other people. The order is untouched, so every possible cut is untouched, so the tree is untouched.

Remember this, because it is not true of every model. In Step 10 you meet one that judges how alike two people are by adding their columns up, and for that one the units are everything: a column measured in thousands drowns a column measured in ones. Scaling changes what that model does. Scaling changes nothing here.

One last thing, and it is the whole of Step 4

You gave the tree seven columns. The file has seventeen, one of which is y, the answer, so sixteen were on the table.

Count how many it could actually have used, though. Nine of the sixteen hold words, and you are about to see what happens with those. Of the seven that hold numbers, six are already in your list. There was exactly one column the tree could have taken and did not, and the Gold task at the bottom is about it.

Now try handing it one of the nine.

try:
    DecisionTreeClassifier().fit(df[["age", "job"]], y)
except ValueError as error:
    print("ValueError:", error)
You should see

ValueError: could not convert string to float: 'unemployed'

You are only running that one, not writing it, so do not worry about its shape. try and except mean: attempt this, and if it goes wrong, catch the complaint and print it as ordinary text instead of turning the screen red. It is written that way so the rest of your notebook still runs afterwards. The words after ValueError: are exactly what you would have seen in red.

There it is. The tree will not take a column of words. Not a single one.

You got around it once, by hand, for poutcome, using .astype(int) in Part 2. There are eight more columns of words in that file, and most of them have more than two answers, so that trick will not stretch.

Step 4 is nothing but this problem, solved properly.

Part 7Make40 minNo answer given

Your own table, your own tree

No code here and no answer at the end. Use the table you signed up in Step 0.

Use new names

mine, not df. my_X and my_y, not X and y. If you reuse the old names the cells above will still run and print numbers that look fine and are wrong.

The brief

Train a tree of depth 2 on your own table, draw it by hand, and put its numbers next to the rule you wrote in Step 1.

It is done when

If your tree turns out to be a single box that says no to everything

That happens, and it is a real result, not a failure. You may well see four boxes rather than one, because the tree splits anyway and then every box gives the same answer, which comes to the same thing.

It means none of the columns you gave it separates your two answers well enough to change the commonest answer in any box. The Silver task at the bottom does exactly this on the bank file.

Keep it and write it down. Then try adding one more column and see if the tree wakes up. If it does not, say so in your README. Step 5 is where you learn to build the column it needed.

Working alone

Put both drawings away for a day. The object tree from Part 0 and your own tree from this part.

Come back to one of them cold and follow the branches with your finger, saying out loud what happens to a new person at each one. If you can get from the top to an answer without stopping, your drawing is right. If you stop, the place you stopped is the thing to fix.

If you have somebody to show it to

Explain one of the drawings to a person who has never seen it, the same way. If they can tell you what would happen to a new person, that is the stronger version of the same test.

Part 8Ship15 min

Write it up

Add to the README you started in Step 0 and extended in Step 1.

  1. Your three-line table: lazy rule, hand rule, tree. Score and catches for each.
  2. Your tree written out in words, as questions. Not a picture. Words.
  3. Which columns you gave it, and which you held back and why.
  4. One line on what the tree found that you did not.
  5. One line on what still worries you about the number at the bottom of Part 5.
Then show the drawing to one person

Show your hand-drawn tree to somebody who does not code. Do not explain it first.

Ask them one question: what would this do with somebody like you? Write down what they say, word for word.

A tree you can draw is the easiest kind of model for a stranger to check, and drawing it is the cheapest way there is. Bigger models can be checked too, by other means, and Step 19 is the step about doing that when nobody can draw the thing at all.

Part 9Check15 min

Check yourself

Six questions. Answer them in writing before you open the answers.

One counts more than the other five. You can get five right and still not pass this step.

  1. What does .astype(int) do to a column of True and False, and why did the tree need it?
  2. Your depth 1 tree and your Step 1 hand rule gave the same four numbers. What does that tell you about what a model is?
  3. A leaf holds 110 people and 38 of them said yes. The tree answers no for that leaf. Why?
  4. Your tree's second question is balance <= 10537.50.

    You change every balance in the file from euros to cents, so 10537.50 becomes 1053750. Nothing about any person changes. You train the same tree again, with the same settings.

    What happens to the tree, and what happens to its score? Say why.

  5. The unlimited tree scored 0.9998 with 709 endings. Roughly how many people are in each ending, and why does that bother you?
  6. You hand the tree a column of job titles and it refuses. In one sentence, what is it actually complaining about?
Open the answers, once all six are written down
  1. It turns every True into 1 and every False into 0. The tree needed it because it only works with numbers, and it cannot compare a True to a cut.
  2. That the model is not cleverer than you: given those seven columns it tried every cut of every one and arrived where you arrived by thinking. Full marks if you also said what it was given. One of the seven was your Step 1 rule, turned into ones and zeros by you in Part 2, so it confirmed your rule rather than discovering it. Take that column away and a two-question tree finds nothing at all, though a three-question one starts finding people again.
  3. Because 38 is fewer than half of 110, and this kind of tree gives each leaf whichever answer is commonest in it. It is only allowed one answer per leaf, so it gives the safer one, even though that group is three times more likely to say yes than the file as a whole.
  4. Nothing happens. The tree is identical and the score is identical, 0.8938. The only visible change is that the cut is now written 1053750 instead of 10537.50, which is the same place in the new units. A tree only asks whether a number is below a cut, and multiplying a column by 100 does not change which people are below which. Mark yourself right only if your answer says the order of the people is unchanged. Multiplying by a hundred keeps that order; multiplying by minus a hundred would turn it upside down, which is a different question. Saying "it stays the same" without saying why is half.
  5. About six people per ending on average, which sounds survivable until you look at the spread. The middle ending holds two people, and 283 of the 709 hold exactly one. A rule built around one person is not a rule, it is a note about that person. That is what should bother you, and it is the whole of Step 3.
  6. That it cannot turn the word into a number. It has no way to ask whether "unemployed" is below a cut, because words do not sit in an order it can cut.
If you got the marked one wrong

Go back to Part 6 and run it again, and this time write down the two thresholds side by side before you read anything. Then read the answer, write it out again tomorrow from memory, and go on. That counts.

Step 10 depends on you knowing this one properly, because it shows you a model where the opposite is true and the contrast is the whole lesson.

StretchOptionalHarder

If you want more

Bronze

Train trees at depth 1 through 8, and write two columns: the score, and the number of people caught.

One of those columns rises at every single step. The other does not. Say which is which, and say what a number that only ever rises can and cannot tell you.

Silver

Take said_yes_before back out of the list and train a depth 2 tree on the six number columns alone.

It will say no to all 4,521 people and score 0.8848. Explain, in writing, why removing one column of ones and zeros did that.

Gold

Add duration to your list of columns. It has never been in it, so this is putting it in for the first time. Pass the same list to export_text or the tree will be labelled wrongly.

Train a depth 2 tree on all eight columns. Write down the score, what it splits on first, and how many of the 521 it catches. Compare that catch count with your 83.

Then write the sentence you would say to a manager who saw those catches and asked why you are not using it.

Before you move on

All six, or it is a not yet.

  1. The thing you built. A hand-drawn depth 2 tree of your own, with real cuts and four row counts that add up to your table.
  2. The comparison. Three lines in your README: lazy rule, hand rule, tree. Score and catches for each.
  3. The build log. The cell that printed 0.0, the two things it was comparing, and what it printed once you fixed it. Plus one prediction you wrote down that turned out wrong.
  4. Check yourself. Five of six right, including the marked one.
  5. Ninety seconds, out loud. Record yourself following the branches of your own tree with your finger and saying what it would do with a new person. Every step ends this way.
  6. It still runs. Kernel, then Restart Kernel and Run All Cells. Nothing should turn red at all. Two cells misbehave quietly on purpose: the one you were asked to find prints 0.0, and the last cell of Part 6 prints a ValueError as ordinary black text. Quiet is the point of this step.

And keep the note you wrote in Part 5, about why you do not believe 0.9998. Step 3 opens by asking you to read it back to yourself.