You hand it seven columns and let it find its own rule. What it finds first is the one you scored by hand in Step 1.
By the end you will have built a model, drawn it on paper, and been able to say exactly what it does. Not roughly. Exactly, question by question, the way you would explain a form.
That matters more than it sounds. Most people who use models cannot say what theirs does. They can say what it scored. You are going to be able to point at yours and read it out loud.
In Step 1 you read a rule, scored it yourself, and got 0.8929, with 83 people caught and 46 calls wasted. Step 1 said at the time that it was a rule you could have thought of.
In this step you give a computer seven columns and no rule to follow, and let it work one out.
It picks yours. The same rule, the same 83 people, the same 46 wasted calls. Not close to it. Identical.
Then Part 3 asks you to notice the catch in that, because one of the seven columns is your Step 1 rule with the words taken out, and you are the one who put it there. What the computer did was confirm your rule was the best of the seven. What it could not do is find it in two questions, and the Silver task at the bottom has you watch that happen.
That is the best possible introduction to a model, and it is why Step 1 came first. A model is not a wiser thing than you. It is a very fast, very patient thing that tries every split of every column and keeps the best one.
Part 4 is one idea and wants half an hour together: the middle of it shows you a group of people the tree says no to for a reason you cannot see yet, and the end of it is the reason. If you have less than half an hour left, stop at the end of Part 3.
In three-quarter-hour evenings: Parts 0 to 2, then Part 3, then Part 4, then Parts 5 and 6, then Part 7 on its own.
Step 1 finished, including your own table. You will use the same file and the same laptop set-up.
One thing you will notice. The files are called band-00.ipynb, band-01.ipynb and so on. The steps were called bands while this was being written and the file names kept it. Same thing, and nothing you have done changes.
Read this whole part first. Then do it on a table, with real objects. It takes about ten minutes and it is the only explanation of a tree you will need.
Anything, as long as they are different from each other. Things off your desk. Things out of a kitchen drawer. Coins, keys, pens, fruit, cups, tins.
Now pick something you want to sort them into. Two piles. Things I would put in a bag versus things I would not. Things that are mine versus things that are not. It does not matter, as long as you know the answer for all twenty.
Push them into one heap. Look at the heap and think of a single yes or no question you could ask about any one object. Is it made of metal. Is it bigger than my hand. Does it have writing on it.
Write your question on a piece of paper.
Ask your question of every object and push the heap into two smaller heaps, yes and no.
Now look at your two heaps. You want each one to be as pure as possible. A heap that is all one answer is perfect. A heap that is half and half taught you nothing.
Now do each heap again, separately. Pick a new question for the left heap, and a different new question for the right heap. Split each of them into two.
Write both questions on your paper underneath the first, joined to it by lines.
Stop there. Three questions in total: one at the top and two underneath it. Do not go further, even though you will want to.
Look at what you have drawn. One question at the top, splitting into more questions, ending in heaps. Turn the paper so it hangs downwards.
A drawing that branches: one question, then two, then four. And at the bottom, four heaps of objects.
Count each heap. Write two numbers on each: how many things are in it, and how many of those were the answer you were looking for.
That drawing is a decision tree. It is upside down, with the root at the top, which is how everybody draws them and nobody explains why.
You chose your questions by looking at the heap and using judgement. In Part 3 a computer will choose its questions by trying every possible one and measuring which leaves the tidiest heaps. That is the main difference between you and it, and there are others worth knowing: it can only ask whether a number is below a cut, it never goes back and revises its first question after seeing where the second one led, and it cannot invent a question you did not give it a column for. Steps 3 and 5 are largely about those three.
Keep the paper. Part 4 asks you to draw a second tree next to it.
Six words. Read them once, then write each in your own words in your notebook.
balance <= 10537.50 the split is the whole question, and 10537.50 is the cut. Some people say threshold. Same thing.Open band-02.ipynb, in the course folder. Run the cells from the top.
Run every cell from the top before you start. Nothing is remembered overnight, and the cells later in this step need names that were made earlier in it.
If you see NameError: name 'X' is not defined, that is all this is. Nothing is broken.
Check that every random_state=0 on the page is in your cell too. Without it the rows are shuffled differently every run, so your numbers change every time you press Shift and Enter and no two people in the world can compare notes.
If the numbers still differ, run every cell from the top in a fresh kernel: Kernel, then Restart Kernel and Run All Cells.
It does not stay running. Open a terminal, move into the course folder, type jupyter lab and press Enter. Parts 6 and 7 of the setup page show you how.
The setup page installed this for you, so most likely you already have it and the command below will just say so. Open a second terminal, move into the course folder the way setup Part 6 shows, and run it. Use python instead of python3 on Windows.
python3 -m pip install scikit-learn
You should see a line starting Successfully installed, or Requirement already satisfied, which is just as good. Then in JupyterLab choose Kernel, then Restart Kernel.
The name you install is scikit-learn. The name you type to use it is sklearn. Two names, one thing, and nobody ever warns you.
The install did not take, or your kernel is still the old one. Put %pip install scikit-learn in a cell and run it. The % matters. Then Kernel, Restart Kernel, and run from the top.
Load the file, the same way as before.
import pandas as pd
df = pd.read_csv("data/bank.csv", sep=";")
print(df.shape)
(4521, 17)
Turn your Step 1 rule into a column.
In Step 1 the rule you scored was df["poutcome"] == "success", which gives a column of True and False. A tree wants numbers, so put .astype(int) on the end. That turns every True into 1 and every False into 0.
df["said_yes_before"] = (df["poutcome"] == "success").astype(int)
print(df["said_yes_before"].value_counts())
The left hand side is new. df["said_yes_before"] = ... makes a brand new column in the table and fills it.
0 4392 and 1 129
129 ones. Those are the same 129 people your hand rule said yes to in Step 1.
Choose the columns the tree is allowed to see.
cols = ["age", "balance", "day", "campaign",
"pdays", "previous", "said_yes_before"]
X = df[cols]
y = (df["y"] == "yes").astype(int)
print(X.shape, y.sum())
Three things here.
df[cols], with a list inside the brackets, gives you a smaller table: just those seven columns, all 4,521 rows. One name in brackets gives one column. A list of names gives a table.X is what the model may look at. y is the answer you want it to guess. Everybody uses those two letters, so you may as well start now.y holds 1 for yes and 0 for no, made with the same .astype(int) you used a moment ago.(4521, 7) 521
Seven columns to look at, and 521 real yes answers to find.
pdays still has its 3,705 minus ones in it, from Step 0. It does not choke a tree the way it choked an average, because all a tree ever asks is whether a number is below some cut. That is not the same as harmless, and Step 4 goes back and deals with it properly. Step 4 deals with them properly.
duration is missing on purpose. It is the length of the call, so you only know it once the call has ended. Step 1's Silver task, if you did it, asked you to write down why the bank could never use it. Step 3 puts numbers on it, and Step 6 comes back to it.
You are going to allow the tree exactly one question. One split, then it must answer. It has seven columns to choose from and it may cut any of them anywhere.
Which of the seven columns will it pick, and where will it cut?
Then write down what you think it will score. Your hand rule got 0.8929. The lazy rule got 0.8848.
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(max_depth=1, random_state=0)
tree.fit(X, y)
print(round(tree.score(X, y), 4))
Four new things, one line each.
from sklearn.tree import DecisionTreeClassifier brings in somebody else's work. Scikit-learn is a free library of models, and this is the tree one.max_depth=1 means one question and then answer. random_state=0 means do it the same way every time, so your numbers match mine.tree.fit(X, y) is the training. That one line is the whole of it. Everything else in this course is about whether to trust what comes out.tree.score(X, y) asks the tree to guess every row, then reports the share it got right. It is exactly the sum you did by hand in Step 1, right answers divided by rows, done for you in one word. That is why it sits straight beside your 0.8929.0.8929
Look at that number. Then look at Step 1's table, where you wrote your hand rule down.
Now find out exactly what it built.
from sklearn.tree import export_text
print(export_text(tree, feature_names=cols))
export_text prints the tree as text. feature_names=cols hands it your column names, and without it the tree would call them feature_0, feature_1 and so on. It draws the tree lying on its side, using |--- for a branch and indenting one level per question.
A line saying class: 0 is a leaf, and the number is the answer given there. 0 means no and 1 means yes, because that is what your y column holds.
Two branches. said_yes_before <= 0.50 gives class 0, and above it gives class 1.
In plain words: if they said yes last time, say yes. Otherwise say no.
Now count what it caught. tree.predict(X) runs all 4,521 rows down the tree and hands back one answer each, a 1 or a 0, in the same order as the table. So guess.sum() is how many it said yes to, exactly the way rule.sum() worked in Step 1.
guess = tree.predict(X)
print(guess.sum())
print(((guess == 1) & (y == 1)).sum())
print(((guess == 1) & (y == 0)).sum())
129, then 83, then 46
Says yes to 129. Catches 83. Wastes 46 calls. Scores 0.8929.
Those are your four numbers from Step 1. Not close to them. The same four.
It tried every cut of every one of the seven columns and landed exactly where you landed by thinking. A model is not cleverer than you. It is faster and more patient, and it never gets bored halfway through.
Write one line in your notebook: how does that make you feel about the rule you wrote in Step 1?
Go back and look at what you handed over in Part 2. One of those seven columns is said_yes_before, and you built it yourself, out of poutcome, using your Step 1 rule.
So the tree did not go and find your rule in the raw data. It was given your rule, as a column, along with six others, and it worked out that yours was the best of the seven. That is a real and useful thing to have learned, and it is not the same thing as discovering it.
Take that column away and a tree with two questions finds nothing at all. Not a worse rule: nothing. It says no to all 4,521 people and scores 0.8848, which is the lazy rule from Step 1. The Silver task at the bottom has you run it.
Two questions is the limit that matters, not the tree. Give the same six columns a third question and it starts finding people the hard way: 66 of them, at 0.8887. It got there without you, and it needed a bigger tree to do it. That trade is the whole of Step 3.
The reason is the whole of Step 4. The column it needed was poutcome, and poutcome holds words. You did the one piece of work by hand that a tree cannot do at all, in one line, in Part 2, and it looked like nothing.
Somewhere below, a cell works out a score and prints 0.0. It does not crash. A score of zero would mean the model got every single row wrong, which cannot be true of a model that just scored 0.8929.
Find it. It is comparing two things that are not the same kind of thing. Write one line on what those two things are, then fix it so it agrees with the 0.8929 above.
The fix is to change one name. The mended cell still ends in .mean() and prints 0.8929. Replacing the whole thing with tree.score(X, y) does not count, because the point is seeing what the two sides were.
One thing you have not met: .mean() on a column of True and False gives the share that are True, because True counts as 1. So (a == b).mean() asks what fraction of the time a and b agreed. That is another way of writing a score, and you will use it constantly.
You are about to allow it two questions instead of one. What will the score become, and how many of the 521 will it catch?
tree2 = DecisionTreeClassifier(max_depth=2, random_state=0)
tree2.fit(X, y)
print(round(tree2.score(X, y), 4))
print(export_text(tree2, feature_names=cols))
0.8938, then a tree with four endings:
said_yes_before <= 0.50 then age <= 60.50 gives class 0said_yes_before <= 0.50 then age > 60.50 gives class 0said_yes_before > 0.50 then balance <= 10537.50 gives class 1said_yes_before > 0.50 then balance > 10537.50 gives class 0The second question on the yes side is about money. Among the 129 people who said yes last time, it has pulled out the ones with more than 10,537 euros in the bank and decided they will say no.
Check whether that was worth doing.
guess2 = tree2.predict(X)
print(guess2.sum())
print(((guess2 == 1) & (y == 1)).sum())
print(((guess2 == 1) & (y == 0)).sum())
125, then 83, then 42
It still catches all 83. It has stopped ringing 4 people who were never going to say yes. It kept every win and dropped four losses.
That is a real improvement and it is a small one. Write both numbers down. Nine ten-thousandths of a score, and four fewer wasted phone calls. Whether that is worth anything depends on what a phone call costs, which is the argument you had with yourself in Step 1.
Look at the other side of the tree again. On the no side it asked whether age is over 60.5, and then answered no on both branches.
So why did it bother? Find out how many people are in each of those two heaps.
older = (df["said_yes_before"] == 0) & (df["age"] > 60.5)
print(older.sum())
print(y[older].sum())
older is a column of True and False, one per row, built the same way as your rules in Step 1. y[older] means take the answers column and keep only the rows where older is True. It lines up because y and df are still in the same row order.
110, then 38
110 people, of whom 38 said yes. That is more than one in three, in a file where the overall rate is about one in nine.
That was one box out of four. Here is the whole set. You will need this exact shape again on your own table in Part 7, so read it rather than just running it.
yes_before = df["said_yes_before"] > 0.5
old = df["age"] > 60.5
rich = df["balance"] > 10537.5
print((~yes_before & ~old).sum(), (~yes_before & old).sum())
print((yes_before & ~rich).sum(), (yes_before & rich).sum())
Each line of the tree becomes one True and False column, and then & and ~ combine them into the four boxes, exactly the way you counted catches and wasted calls in Step 1.
4282 110, then 125 4
Four boxes. Add them up: they must come to 4,521, because every person is in exactly one box.
The tree found a group that is three times more likely than average to say yes, and then said no to all of them anyway.
It did that because 38 out of 110 is still fewer than half, and this tree gives each leaf whichever answer is commonest in it. It is being asked for one answer per heap, so it gives the safe one.
That is not a fault in the tree. It is a fault in the question we asked it. Step 9 is where you learn to ask for a chance instead of an answer, and Step 7 is where you learn to move the line. Write this leaf down. You will come back to it.
On the same paper as your object tree from Part 0, draw this one by hand.
said_yes_before <= 0.5.age <= 60.5 on the left, balance <= 10537.5 on the right.The two boxes under said_yes_before <= 0.5 holding 4,282 and 110, so one of them is enormous and one is small. The two on the other side holding about a hundred between them, and the box under balance > 10537.5 holding fewer than ten.
All four adding to 4,521. And only one of the four boxes saying yes.
If your two big heaps are on the right hand side, your branches are the wrong way round. <= 0.5 is the people who did not say yes before.
Now put it beside the tree you drew with your hands in Part 0. Same shape: one question, two questions, four boxes.
Two questions gave 0.8938. Now take the limit off entirely and let it ask as many questions as it likes.
What will it score with no limit at all?
big = DecisionTreeClassifier(random_state=0)
big.fit(X, y)
print(round(big.score(X, y), 4))
print(big.get_n_leaves())
print(big.tree_.max_depth)
Look at what is not in that first line. Last time you wrote max_depth=2. Leave it out and there is no limit on the depth: it keeps asking questions until every ending holds people who all gave the same answer, or holds a single person. Leaving a setting out does not mean off, it means use the usual, and here the usual is as deep as you like.
Two more new things. get_n_leaves() has brackets because it is a job you are asking the tree to do, counting its endings. tree_.max_depth has none, because it is not a job, it is a fact the tree wrote down while it trained. This is the mirror of Step 0's planted bug, where brackets were missing from a job and you got a <bound method line instead of a number.
Do not turn that into a rule about which things are jobs. Whoever wrote the library decided, and not always consistently: this same tree gives you its number of endings as tree_.n_leaves with no brackets, and its depth as get_depth() with them. The rule worth keeping is narrower and never fails: if what comes back starts with <bound method, you left the brackets off.
0.9998, then 709, then 30
guess3 = big.predict(X)
print(((guess3 == 1) & (y == 1)).sum())
print(((guess3 == 1) & (y == 0)).sum())
520, then 0
It found 520 of the 521 people who said yes, and it did not waste a single phone call.
Sit with that for a moment. It is nearly perfect. Every hard thing this course has told you is difficult, it has just done, in under a second, on a laptop.
Write down, in your own words, why you do not believe that 0.9998.
Do not look anything up. Do not ask anybody. Just write what bothers you. One or two sentences.
Here is one clue, and it is the only one you get. The tree has 709 endings for 4,521 people. Work out roughly how many people are sitting in each ending.
Keep what you wrote. The first page of Step 3 asks you to read it back.
Your two-question tree cuts the money column at 10,537.50 euros. Suppose the bank had recorded every balance in cents instead of euros, so every number is a hundred times bigger. Nothing about any person has changed. Only the units.
Train the same tree on that table. Write down what happens to the tree, and what happens to its score.
Most people get this wrong, and getting it wrong here costs you nothing.
X_cents = X.copy()
X_cents["balance"] = X_cents["balance"] * 100
cents_tree = DecisionTreeClassifier(max_depth=2, random_state=0)
cents_tree.fit(X_cents, y)
print(round(cents_tree.score(X_cents, y), 4))
print(export_text(cents_tree, feature_names=cols))
Two new pieces in there. X.copy() makes a separate table holding the same things, so that changing the copy leaves X alone. Without it you would quietly damage the table Parts 3 to 5 were built on, and nothing would tell you. And X_cents["balance"] * 100 multiplies all 4,521 balances at once. You do not have to go round them one at a time.
0.8938, and exactly the same tree, except that the money cut now reads 1053750.00 instead of 10537.50.
Same questions, same order, same answers, same score. The cut moved to the same place in the new units.
A tree never adds, multiplies or compares one column against another. The only thing it ever asks is whether one number is below some cut.
Multiplying a column by a hundred does not change which people are below which other people. The order is untouched, so every possible cut is untouched, so the tree is untouched.
Remember this, because it is not true of every model. In Step 10 you meet one that judges how alike two people are by adding their columns up, and for that one the units are everything: a column measured in thousands drowns a column measured in ones. Scaling changes what that model does. Scaling changes nothing here.
You gave the tree seven columns. The file has seventeen, one of which is y, the answer, so sixteen were on the table.
Count how many it could actually have used, though. Nine of the sixteen hold words, and you are about to see what happens with those. Of the seven that hold numbers, six are already in your list. There was exactly one column the tree could have taken and did not, and the Gold task at the bottom is about it.
Now try handing it one of the nine.
try:
DecisionTreeClassifier().fit(df[["age", "job"]], y)
except ValueError as error:
print("ValueError:", error)
ValueError: could not convert string to float: 'unemployed'
You are only running that one, not writing it, so do not worry about its shape. try and except mean: attempt this, and if it goes wrong, catch the complaint and print it as ordinary text instead of turning the screen red. It is written that way so the rest of your notebook still runs afterwards. The words after ValueError: are exactly what you would have seen in red.
There it is. The tree will not take a column of words. Not a single one.
You got around it once, by hand, for poutcome, using .astype(int) in Part 2. There are eight more columns of words in that file, and most of them have more than two answers, so that trick will not stretch.
Step 4 is nothing but this problem, solved properly.
No code here and no answer at the end. Use the table you signed up in Step 0.
mine, not df. my_X and my_y, not X and y. If you reuse the old names the cells above will still run and print numbers that look fine and are wrong.
Train a tree of depth 2 on your own table, draw it by hand, and put its numbers next to the rule you wrote in Step 1.
.astype(int) on the words themselves. You get another ValueError, this one saying invalid literal for int() with base 10, which is pandas making the same complaint in its own words. Compare to one of the two answers first, the way you did in Part 2: mine["new_col"] = (mine["old_col"] == "yes").astype(int). If it has more than two answers, leave it out and put it on a list for Step 4.That happens, and it is a real result, not a failure. You may well see four boxes rather than one, because the tree splits anyway and then every box gives the same answer, which comes to the same thing.
It means none of the columns you gave it separates your two answers well enough to change the commonest answer in any box. The Silver task at the bottom does exactly this on the bank file.
Keep it and write it down. Then try adding one more column and see if the tree wakes up. If it does not, say so in your README. Step 5 is where you learn to build the column it needed.
Put both drawings away for a day. The object tree from Part 0 and your own tree from this part.
Come back to one of them cold and follow the branches with your finger, saying out loud what happens to a new person at each one. If you can get from the top to an answer without stopping, your drawing is right. If you stop, the place you stopped is the thing to fix.
Explain one of the drawings to a person who has never seen it, the same way. If they can tell you what would happen to a new person, that is the stronger version of the same test.
Add to the README you started in Step 0 and extended in Step 1.
Show your hand-drawn tree to somebody who does not code. Do not explain it first.
Ask them one question: what would this do with somebody like you? Write down what they say, word for word.
A tree you can draw is the easiest kind of model for a stranger to check, and drawing it is the cheapest way there is. Bigger models can be checked too, by other means, and Step 19 is the step about doing that when nobody can draw the thing at all.
Six questions. Answer them in writing before you open the answers.
One counts more than the other five. You can get five right and still not pass this step.
.astype(int) do to a column of True and False, and why did the tree need it?Your tree's second question is balance <= 10537.50.
You change every balance in the file from euros to cents, so 10537.50 becomes 1053750. Nothing about any person changes. You train the same tree again, with the same settings.
What happens to the tree, and what happens to its score? Say why.
Go back to Part 6 and run it again, and this time write down the two thresholds side by side before you read anything. Then read the answer, write it out again tomorrow from memory, and go on. That counts.
Step 10 depends on you knowing this one properly, because it shows you a model where the opposite is true and the contrast is the whole lesson.
Train trees at depth 1 through 8, and write two columns: the score, and the number of people caught.
One of those columns rises at every single step. The other does not. Say which is which, and say what a number that only ever rises can and cannot tell you.
Take said_yes_before back out of the list and train a depth 2 tree on the six number columns alone.
It will say no to all 4,521 people and score 0.8848. Explain, in writing, why removing one column of ones and zeros did that.
Add duration to your list of columns. It has never been in it, so this is putting it in for the first time. Pass the same list to export_text or the tree will be labelled wrongly.
Train a depth 2 tree on all eight columns. Write down the score, what it splits on first, and how many of the 521 it catches. Compare that catch count with your 83.
Then write the sentence you would say to a manager who saw those catches and asked why you are not using it.
All six, or it is a not yet.
0.0, the two things it was comparing, and what it printed once you fixed it. Plus one prediction you wrote down that turned out wrong.0.0, and the last cell of Part 6 prints a ValueError as ordinary black text. Quiet is the point of this step.And keep the note you wrote in Part 5, about why you do not believe 0.9998. Step 3 opens by asking you to read it back to yourself.