Step 1

A rule you wrote yourself

You will write a rule by hand. Then you will find out if it is any good.

About 3 hours 30 minutes10 partsBest over two sittingsNo model yet
Read this first3 min

What this step gives you

By the end of this step you will be able to hear anybody say my model is 94% accurate and know, immediately, that they have told you nothing. You will know the questions that turn that sentence into a fact, and you will have asked them of your own work first.

You will also have made a prediction machine. Not a model. A rule, written by you, in words, that takes a row of a table and guesses an answer. It is the simplest possible version of the thing this whole course is about, and you will build it in an afternoon.

Why you write the rule by hand

Almost every course starts by handing you a model. You type one line, a number appears, and you have no way of knowing whether that number is good, bad or meaningless.

If you write the rule yourself, you know exactly what it does, because you decided. So when you score it, the score is the only new thing in the room. That is the only way to learn what a score actually is.

In Step 2 a model does this for you, and by then you will know what it is doing for you and why it is not magic.

What you actually do, in order

  1. Guess ten things on paper, twice, two different ways, and count.
  2. Open the same table you met in Step 0 and count the answers in it.
  3. Build the laziest rule there is and score it. This number is the one everything else has to beat.
  4. Read a rule somebody else wrote, and work out what it believes about people.
  5. Score it two ways, because one way is not enough.
  6. Change one thing in it, and measure what that did two different ways. This is the part people remember.
  7. Write your own rule on your own table, with no help and no answer given.
How the time actually goes

Part 6 is one idea and wants half an hour in one go, because the first half of it looks like you broke the rule and the second half is why you did not. If you have less than half an hour left, stop at the end of Part 5.

In three-quarter-hour evenings: Parts 0 to 2, then Parts 3 and 4, then Part 5, then Part 6, then Part 7 on its own.

What you need before you start

Step 0 finished, including the table of your own that you signed up in Part 7 of it. If you skipped that, go back. This step asks you to use it.

You will write code in this step, but not much, and never before you have read and run the same idea first.

Part 0Unplug15 minAway from the screen

Guess ten times, two ways

Read this whole part first. Then go and do it with a pen, away from the laptop. It takes about ten minutes.

You are about to build a prediction machine on paper, score it, and then discover something about the score that most people never find out.

Set it up

Draw four columns on a page: name, clue, my guess, truth.

Now think of ten people you know well. Write their names down the first column.

Pick something about them that has only two answers. Do they drink tea, yes or no. Do they own a car, yes or no. Anything with two answers where you know the truth for all ten.

Fill in the truth column for all ten. Then fold the page over so you cannot see that column, or cover it with a book.

Round one: use a clue

Pick one clue. Just one. Their age, or where they live, or their job. Write it in the clue column for each person.

Now make a rule out of it, in words, and write the rule at the top of the page. Something like: if they are over 40, I say yes.

Follow your own rule for all ten and fill in my guess. Follow it even when you think it is about to be wrong. That is the point of having a rule.

Uncover the truth column. Count how many you got right, out of ten. That is your clue score.

Round two: do not think at all

Count how many of your ten drink tea. Count how many do not. Take whichever group is bigger, and write that same answer for all ten people.

Count how many that got right, out of ten. That is your lazy score.

A worked example, so you know what this looks like

Say seven of my ten drink tea. My clue is age, and my rule is: over 40 means tea.

My rule gets 6 out of 10. I feel quite good about that, because I invented it.

Then I write tea for all ten people. That gets 7 out of 10, because seven of them drink tea.

My rule, which I thought about, lost to a rule that thought about nothing.

You should see

Two scores on your paper, both out of ten, and a rule written in words at the top of the page.

Your lazy score must equal the size of the bigger group. If six of your ten drink tea and you wrote tea ten times, your lazy score is exactly 6. It can never be below 5, because the bigger group is always at least half.

If your lazy score came out below 5

You wrote down the smaller group. Cross it out and do that round again with the other answer.

Now write one line, and keep the page

Which score was higher, and by how many?

If the lazy score won, you have just found the thing this step is about, on paper, before touching a computer.

If your clue won, note the gap. With only ten people a gap of one or two is luck, not skill. Part 5 shows you the same contest with 4,521 people, where luck runs out.

Either way, keep the page. The last line of this step asks you to put it next to what your laptop produced.

Part 1Words5 min

Words to know

Five words. Read them once. Then write each one in your own words in your notebook.

Row
One line in the table. One person, one sale, one day.
Column
One thing you know about every row. Age is a column.
Target
The column you want to guess. In this file it is called y.
Rule
A test you can run on a row that gives you a guess.
Baseline
The score you get by not thinking at all. Your rule has to beat it.
Do this now

Write your five lines. Do not copy mine. If you cannot write one in your own words, you do not know it yet, and Part 2 will feel harder than it is.

Part 2RunInvestigate25 min

Look at the table first

Open band-01.ipynb, in the course folder. Run the cells from the top. Do not skip any.

If you are coming back on a new evening

Run every cell from the top before you start. Nothing is remembered overnight, and the cells later in this step need names that were made earlier in it.

If you see NameError: name 'truth' is not defined, that is all this is. Nothing is broken.

If your numbers are different from mine

Check that every random_state=0 on the page is in your cell too. Without it the rows are shuffled differently every run, so your numbers change every time you press Shift and Enter and no two people in the world can compare notes.

If the numbers still differ, run every cell from the top in a fresh kernel: Kernel, then Restart Kernel and Run All Cells.

If JupyterLab is not open in your browser

It does not stay running. You start it fresh every time you sit down to work, and after a break it will be gone.

Open a terminal the way Part 4 of the setup page shows. Move into the course folder the way Part 6 shows. Then type jupyter lab and press Enter.

  1. Load the file.

    import pandas as pd
    
    df = pd.read_csv("data/bank.csv", sep=";")
    
    print(df.shape)

    Four small things are happening. Read them once now. You will see all four in every step from here on.

    • import pandas as pd brings in pandas, the library that handles tables, and gives it the short name pd so you do not have to type the whole word. Every notebook in this course brings in pandas in its first cell, so you will type or run this line more than any other.
    • pd.read_csv means ask pandas to read a file of this kind. CSV is the plainest way to save a table.
    • sep=";" tells it that this file separates its columns with a semicolon. Most files use a comma, so most of the time you leave this off. Leave it off here and all 17 columns get squashed into one.
    • df is the name we are giving the table. It is short for data frame, which is what pandas calls a table. Everybody in this field calls it df, so you may as well start now.

    This is the same first cell you ran in Step 0, so you already know what it does.

    What changes in this step: from Part 3 on, every line of code is short enough that you are meant to follow all of it. Where a line is not, the page says so and tells you to run it anyway. If a page has not excused a line, you should be able to read it, and if you cannot, that is a fault in the page rather than in you.

    You should see

    (4521, 17)

    That is 4,521 rows and 17 columns. If you see (4521, 1) you left out the sep.

    If you get ModuleNotFoundError: No module named 'pandas'

    Pandas is not installed where this notebook is running. Put %pip install pandas in a cell and run it. The % matters. Then restart the kernel and run from the top.

    If you get FileNotFoundError

    Python looked for data/bank.csv and it was not there. That means JupyterLab was started somewhere other than the course folder.

    Save your notebook. Close the browser tab. Go to the terminal window you started Jupyter from and press Control and C together to stop it. Then follow Part 6 of the setup page again, so your terminal is standing in the course folder, and start jupyter lab from there.

    Do not move the notebook to a different folder to make the error go away. It will come back in the next step.

  2. Read what the file says about itself.

    print(open("data/bank-names.txt").read()[:1500])

    Somebody wrote notes about this data. Read them before you trust a single number.

    You should see

    Notes from a Portuguese bank. One line says the goal is to guess if a client will subscribe a term deposit.

    So that is what y means. Not a savings account. A term deposit, which is money locked away for a fixed time. You did not have to take my word for it, and you should not have.

  3. Look at three rows.

    df.head(3)

    Read them. Do not skim. Each row is one client in one campaign, and the notes say the call columns describe the last call made to them. Some of these people were rung once. One was rung fifty times. Part 6 charges a euro a row, which is a euro per client rather than a euro per call, and that is a simplification you should know you are making.

    You should see

    Three rows and 17 columns. Row 0 begins 30 unemployed married primary. Row 2 begins 35 management single tertiary.

    The last column, y, says no for all three.

    If your first row is not the 30 year old, you have opened a different file. Check step 1.

  4. Count the answers.

    print(df["y"].value_counts())
    You should see

    no 4000 and yes 521

    So 521 people said yes. The rest said no. That is about 11 in every 100.

Think before you go on

Write the answers in your notebook. One line each.

Part 3PredictRun15 min

The lazy rule

Here is the laziest rule there is. Say no to every single row. Never say yes. Not once.

Predict first. Write it down.

Out of 4,521 rows, how many will the lazy rule get right?

Write a number in your notebook now. Not a range. One number. You cannot get this wrong in a way that costs you anything, and writing it is the whole point.

Now build the lazy rule properly, and score it the same way you will score every rule from here on.

Two new pieces in the next cell, both worth a sentence.

truth = df["y"] == "yes"
lazy = pd.Series(False, index=df.index)

lazy_right = (lazy == truth).sum()
lazy_catches = (lazy & truth).sum()

print(lazy_right)
print(round(lazy_right / len(df), 4))
print(lazy_catches)
You should see

4000, then 0.8848, then 0

The lazy rule gets 4,000 out of 4,521 right. That is 88.48 out of every 100. And it finds nobody at all.

Look at your prediction. Write down the gap between your number and 4,000.

This number has a name

0.8848 is your baseline. Write it at the top of a fresh page and keep it there for the rest of the step.

From now on, any rule that scores less than 0.8848 is worse than not thinking, on this measure. Keep the condition. Part 6 hands you a rule that scores below it and, at the prices you pick there, is worth more money.

Part 4RunInvestigate25 min

Read a rule you did not write

You will write your own rule in Part 7. First you read one that works.

One column is called poutcome. The notes describe it as the outcome of the previous marketing campaign, which means what happened the last time the bank ran a campaign at this person. That line sits past the 1,500 characters you printed in Part 2, so do not take my word for it: change the 1500 in that cell to 4000, run it again, and find the line yourself.

print(df["poutcome"].value_counts())
You should see

Four groups. unknown has 3,705 rows. failure has 490. other has 197. success has 129.

Do not take my word for what unknown means. Check it. There is another column, previous, that counts how many times this person was called before.

print(df.groupby("poutcome")["previous"].mean().round(2))
You should see

unknown is 0.0. The other three are near 3.

So unknown is not a lost record. There was no earlier campaign for those 3,705 people, so there is no earlier outcome to record. Read that carefully: it does not mean the bank had never rung them. Two thirds of them were rung more than once during this campaign. It means there was no last time before this campaign began.

Only 129 people said yes last time. Here is the rule someone wrote:

rule = df["poutcome"] == "success"

print(rule.sum())
print(rule.dtype, len(rule))
You should see

129, then bool 4521

rule is a list of True and False, one for every row. True means the rule says yes.

The second line should read bool 4521. bool means the column holds only True and False. 4521 means there is one for every row.

If you get AttributeError: 'DataFrame' object has no attribute 'dtype', you have typed an extra df[ and ] around it, which gives you a smaller table instead of a column of True and False. Retype the line as exactly rule = df["poutcome"] == "success".

Explain this line

Write these in your notebook before you go on.

The answer to the second one, because you cannot look it up

Python counts True as 1 and False as 0. So adding up a column of True and False gives you the number of Trues.

rule.sum() means: how many rows did the rule say yes to. You will use that fact constantly.

The idea under the rule

People who said yes before will say yes again. That is all it is. You could have thought of it. That is the point.

Part 5Investigate25 min

Score it honestly

Now you find out if the rule is any good. You need two numbers, not one.

  1. Check the true answers. You made truth in Part 3.

    print(truth.sum())
    You should see

    521

    If you get NameError: name 'truth' is not defined

    Your kernel was restarted, or you skipped a cell. Nothing is broken. Run every cell from the top, in order, then come back here.

    One cell in this notebook is wrong on purpose

    Somewhere below, a cell counts the yes rows and prints 0. It should print 521. It does not crash. It prints a lie and moves on.

    Find it and fix it. The fix is one character. When it is right, the cell still reads df["y"] == "..." and prints 521.

    Deleting the cell does not count. Setting it equal to truth does not count. The point is the one character, not the number.

    Then run print(df["y"].unique()) and read what the file actually holds.

  2. Count how many the rule got right.

    rule_right = (rule == truth).sum()
    
    print(rule_right)
    print(round(rule_right / len(df), 4))
    You should see

    4037, then 0.8929

  3. Count the catches and the wasted calls.

    catches = (rule & truth).sum()
    bad_calls = (rule & ~truth).sum()
    missed = (~rule & truth).sum()
    
    print(catches, bad_calls, missed)

    Three symbols to learn here, and you will use them for the rest of the course.

    The & means and. The ~ means not. And in Part 6 you will meet |, which means or. It is the tall straight line, usually above the backslash key on your keyboard, and it is not the letter l.

    You should see

    83 46 438

    The rule found 83 real clients. It rang 46 people who said no. It walked past 438 people who would have said yes.

In Part 4 you worked out the rule must miss at least 392. It missed 438. The gap is the 46 wasted calls, because every wasted call is one of the 129 that was not a real yes.

Put the two rules side by side

Copy this into your notebook by hand. Do not paste it.

The rule beat the baseline by 0.0081. That is less than one point in a hundred. A month of work for that would be a bad month.

But look at the second number. The lazy rule finds nobody. The real rule finds 83. It rings 129 people to get them, and that trade is what Part 6 is about.

The rule of this step

One number cannot tell you if a rule is good. Always write two: how often you were right, and how many real ones you caught.

Part 6PredictModify30 min

Change one thing

Some months are better than others. Run this and look at the counts.

year = ["jan", "feb", "mar", "apr", "may", "jun",
        "jul", "aug", "sep", "oct", "nov", "dec"]

for m in year:
    rows = df[df["month"] == m]
    print(m, len(rows), (rows["y"] == "yes").sum())

The first line is a list of the twelve months, written out in full so you never have to work out which is which. The loop then does the same three things once for each month: take only the rows from that month, count them, and count how many of those said yes.

You should see

Twelve lines, one per month, each with three things on it: the month, how many calls, how many said yes.

oct is 37 yes out of 80 calls. dec is 9 out of 20. mar is 21 out of 49. sep is 17 out of 52.

Those four run from about a third to nearly a half. In the file as a whole it is closer to 12 in 100.

Two other months deserve a look before you get carried away. April is 56 of 293 and February 38 of 222, both about half as good again as the file average. And December's tempting 9 of 20 rests on twenty calls, which is exactly the size of group your Part 0 exercise warned you about. Four months is a defensible choice, and it is a choice rather than a fact.

So make the rule wider. Say yes if they said yes before, or if the call was in one of those four months.

Predict first. Write both numbers.

The old rule scored 0.8929 and caught 83.

Write down what you think the new rule will score. Then write how many it will catch. Two numbers, in your notebook, before you run the cell.

months = ["mar", "sep", "oct", "dec"]
rule2 = (df["poutcome"] == "success") | (df["month"].isin(months))

right2 = (rule2 == truth).sum()
catches2 = (rule2 & truth).sum()
bad2 = (rule2 & ~truth).sum()

print(right2, round(right2 / len(df), 4))
print(catches2, bad2)
You should see

3988 0.8821, then 141 153

Read that twice.

The new rule catches 141 people. The old one caught 83. It found 58 more.

And it got fewer rows right. 3,988 instead of 4,037. That is 49 rows worse, so the score drops from 0.8929 to 0.8821. It is now below the lazy rule, which scores 0.8848.

So which rule is better?

You cannot say yet. Not from these numbers alone. It depends on what a wasted call costs, and what a won client is worth.

So put numbers on it. Pick a price for a call and a price for a client. These are made up. Say so.

call_cost = 1
client_worth = 50

narrow = catches * client_worth - rule.sum() * call_cost
wide = catches2 * client_worth - rule2.sum() * call_cost

print(narrow, wide)
You should see

4021 6756

At those two prices the wide rule wins by 2,735 euros, even though its score is worse.

Now change the two prices. Make a call cost 10 and a client worth 12. Run it again, and this time work out a third number the code does not print: what the lazy rule is worth at these prices.

What just happened

The better rule changed when the prices changed. The data did not move at all.

So the question was never only about the numbers in the file. Your job is to hand the bank both scores and both counts, and let them pick the prices.

Watch me get this wrong

The first time I wrote the wide rule, I wrote this:

rule2 = df["poutcome"] == "success" | df["month"].isin(months)

It crashed. There is a cell in the notebook that runs it so you can read the real message for yourself.

That cell catches the error and prints it as ordinary black text, so the rest of the notebook still works. This happens three times in the course, here and in Steps 2 and 4, and each time the page warns you first. The words after TypeError: are exactly what you would have seen.

Python does the | before the ==. So the first thing it tried was "success" | df["month"].isin(months), the word success ored with a column of True and False, and it gave up there. It never reached the comparison at all. That is why the message mentions ror_, which is Python's internal name for an or that arrived the wrong way round.

Brackets fix it, and it is the left side that needs them. Wrapping only the right side changes nothing, because .isin(...) already binds tighter than |.

Then I ran the fixed version and got 0.8821. I thought I had broken something else, because the score had gone down. I spent ten minutes hunting a bug that was not there. The score really does go down. That is how this part of the step came to be written.

Part 7Make40 minNo answer given

Your own rule, your own table

This part has no code in it and no answer at the end. You have done every piece already.

Use the table you signed up in Step 0. Not the bank file.

The brief

Write one rule by hand that guesses your target column. Then score it the way you scored the bank rule.

Use new names

Call your table mine, not df. Call your rule my_rule, not rule. You will need four more, and they are the same four Part 5 built: my_truth, my_catches, my_wasted and my_missed.

The last one is the only one Part 5 did not spell out on its own line. It is my_missed = (~my_rule & my_truth).sum(), the people who said yes that your rule said no to. The ~ means not.

If you reuse the old names, the cells above will still run and will print numbers that look fine and are wrong. That is the worst kind of wrong.

It is done when

Check your four numbers before you write them down

Run these three lines. They work on any table.

  • print(my_catches + my_missed, my_truth.sum()) must print the same number twice. If not, your catches are counting rows your rule said yes to, not rows it got right.
  • print(my_catches + my_wasted, my_rule.sum()) must print the same number twice.
  • Your baseline must be the share of the bigger group. If your target is yes in 20 rows out of 100, your baseline is 0.80 and not 0.20.

If any of these fails, your table is wrong however good the score looks.

If your rule cannot beat the baseline

That is a real result and you should keep it. Write it down. Then try one more column and stop.

A rule that fails to beat the lazy way is worth more to you than a rule you fiddled with until the number looked good.

Working alone

Put this away for a day. Come back and read your own README cold.

Write what a stranger would ask you about it. Then answer them. That is your review, and it works better than you expect.

Part 8Ship15 min

Write it up

Your README needs five things. No more.

  1. What you are trying to guess, in one line.
  2. Where the table came from, and who made it.
  3. Your baseline number.
  4. Your rule in plain words, then your three numbers.
  5. What is wrong with all of this. Be honest. This part is not optional.

Point five is the one people skip. It is the one that gets read.

Then show it to one person

Print your table of numbers. One page. Show it to somebody who does not code.

Ask them one question: which rule would you pay for? Write down what they say, word for word.

Part 9Check15 min

Check yourself

Six questions. Answer them in writing, in your notebook. Do not open the answers until all six are written down.

One of them counts more than the other five. You can get five right and still not pass this step.

  1. You load the file and forget sep=";". What does df.shape print, and why?
  2. A file has 900 rows. 810 of them say no. What does the lazy rule score?
  3. Your rule says yes to 129 rows. 83 of those really said yes. How many people did you ring for nothing?
  4. Your file is 88.5 out of 100 "no". You write a rule and it scores 89.3.

    Your friend writes no rule at all. They answer "no" to every row.

    What does your friend score? And what has your extra 0.8 actually bought you?

  5. Why does the wide rule in Part 6 catch more people and still score worse?
  6. Someone shows you a rule that is right 97 times in 100. What is the first thing you ask them?
Open the answers, once all six are written down
  1. (4521, 1). Every row lands in one column, because the file uses a semicolon and pandas was looking for a comma. Nothing crashes.
  2. 0.90. The bigger group is the 810 no rows, and 810 out of 900 is 0.90.
  3. 46. That is 129 minus 83.
  4. Your friend scores 0.8848, or 88.5. Your extra 0.8 bought you 83 catches out of 521, and 46 wasted calls. Mark yourself right only if your answer has both parts: a number for your friend, and what the gap bought in catches. If you wrote that your rule is better and did not name the catches, mark it wrong. If you wrote only the friend's score, that is half.
  5. It says yes to 294 rows instead of 129. It catches 58 more real clients and makes 107 more wasted calls, so it gets 49 more rows wrong overall. Catches went up, right answers went down.
  6. Any of these counts: what is the baseline? How many real ones did it catch? What is in the other 3?
If you got the marked one wrong

Go back to Part 3 and Part 5, and do them again with a fresh page. Then read the answer, write it out again tomorrow from memory, and go on. That counts.

Every step after this one leans on it. There is no failing here, only not yet, and this is a not yet.

StretchOptionalHarder

If you want more

Bronze

Find a second rule on the bank file that scores above 0.8848 and catches at least 30 of the 521. Write both numbers down.

You need the second number because the first one is easy to cheat. Try df["age"] > 83. It scores 0.8850, so it beats the baseline. It beats it by getting one more row right than saying no to everybody, out of the three people in the whole file who are over 83. It catches 2 of them. A useless rule with a winning score.

Fair warning, because I went looking: there is no easy answer here among the columns you can honestly use. Cut any of them at any value and only one rule in the whole file clears both bars, and it is the one you were just handed. To find a second you have to join two conditions together, which Part 6 shows you how to do. Read Part 6 first, then come back.

There is one column that clears both bars on its own, at 210 different cutting points, and one of those catches 180 people. It is duration, and the Silver task below is about why you cannot have it. If you find it before you get there, you have found the right thing for the wrong reason.

Silver

Now try df["duration"] > 500. It catches 230 of the 521, which is far more than anything else you have built.

Then read what the notes say duration is. Write one line on why the bank could never use this rule to decide who to ring.

Keep that line. Step 3 puts numbers on it, and Step 6 is built on it.

Gold

Go back to the two prices in Part 6. Find the price of a call, in whole euros, at which the narrow rule starts to beat the wide one. Keep a client worth 50.

Write the sum out in full, and say what that price means in plain words.

Then find the hole in it. Part 6 charges one euro per row, and a row is a client, not a call. The campaign column says the narrow rule's 129 clients took 224 calls and the wide rule's 294 took 524. Redo the sum charging per call and see whether your answer survives. Say which of the two answers you would defend to the person paying.

A warning about Gold

Once you have done Gold, you will find it hard to trust any score again without asking what it costs. Good. That is the whole point of Step 7, and of everything that leads to it.

Before you move on

All four of these, or it is a not yet.

  1. The thing you built. Your notebook runs from top to bottom with nothing remembered from before. In the top menu choose Kernel, then Restart Kernel and Run All Cells. Nothing should turn red. Your README has all five parts, and part five is not empty.
  2. The build log. It names one thing that broke and what fixed it. It names one guess you got wrong.
  3. Check yourself. Five of six right against the answers above, and the marked one is one of them.
  4. Explain it out loud. Ninety seconds on your baseline: why it is there and what would go wrong without it. Recorded if you can, written in one go without editing if the house is asleep. Do not do it twice over.

Then take your paper from Part 0 and put it beside your table of numbers. Write one line: do they say the same thing? If your clue beat the lazy way on ten people, write down why ten people was never enough to know that.