No model in this step. Just a table, and everything that is wrong with it.
By the end of this step you will be able to do one thing, and it is a thing most people who work with data cannot do. You will be able to open a table you have never seen before and work out what is wrong with it, before you trust a single number in it.
That is worth saying plainly, because from the outside this step looks like it is not teaching you machine learning at all. There is no model here. You train nothing. You read code rather than write it.
A model is a machine for finding patterns. It cannot tell whether a pattern is real. Feed it a broken table and it will find the broken pattern, print a good score, and nobody will notice for months.
In Part 3 a number is going to lie to you. It has two decimal places. It looks exactly like every number you have ever trusted. It means nothing at all. Once that has happened to you once, you check. For the rest of your working life, you check.
Four evenings of about three quarters of an hour: Parts 0 to 2, then Part 3, then Parts 4 to 6, then Part 7 on its own.
Part 4 is one idea and wants twenty minutes together, because the middle of it looks like good news and the end of it is not. Part 7 is the one that needs a whole evening, and it does not matter if it takes two.
Python and JupyterLab, on your own laptop. If you have never installed either, or do not know what they are, stop and do the setup page first. It assumes you have never installed anything in your life, and it takes about an hour.
You do not need to know how to code. Every line up to Part 7 is written for you and you only run it. In Part 7 you write four short lines of your own, about a table you choose.
Read this whole part first. Then do it with a pen. The laptop stays open only so you can read the questions; everything you write goes on paper.
Here is why. Later today a number on your screen is going to fool you. This is so that when it does, you have already made one of those numbers yourself, with your own hand, and know exactly how it happened. Ideas about data are slippery. Your own handwriting is not.
You need one sheet of paper, a pen, and your phone face down where you cannot read it. Write the numbers 1 to 12 down the left of the sheet.
You are registering for something. A new job's paperwork, an account, a club. Twelve questions, and this round has one rule: answer only what you actually know, from memory. Nothing looked up, nothing worked out, no phone. If you do not know, leave the line blank and move on.
Give yourself about twenty seconds a line. Do not go back.
Your full name, spelled as it is on your passport or driving licence.
Your date of birth.
The full address where you live, including the postcode.
The day, month and year you moved in there.
Your height in centimetres, to the nearest centimetre.
Your shoe size.
The full name of somebody to contact in an emergency.
That person's phone number. From memory. The phone stays face down.
A second emergency contact, and their number.
The month and year you started your current job, course or main occupation.
The full registered name of that employer or school. Not what everybody calls it. The name on the paperwork.
Its main switchboard number.
Count your blanks. Write the number at the bottom of the sheet.
Now the rule every real form actually has. It will not be accepted with a blank on it. There is no box marked "I do not know" and there is nobody to ask.
Go back to every line you left empty and put something on it. Not a joke and not a refusal. Something that would get the form accepted: a plausible number, a round figure, a date that looks like a date, a name you can half remember. Whatever you would really write at half past four on a Friday to get the thing handed in.
Do it now, before you read on. It takes two minutes and the whole step turns on it.
A completed form. Twelve lines, twelve answers, no blanks.
And nothing anywhere on that sheet showing which of the twelve you made up.
You are unusual, and there is a version of this that will still bite. Fill the same twelve in for the person you live with, or the colleague you sit nearest, without asking them anything.
You will not get past line four.
The blanks in round one are the easy problem. A space with nothing in it. A computer finds those all day long and it will tell you about every single one.
What you wrote in round two is the whole problem. Those lines are filled in. To anything reading that form afterwards they look exactly like the ones you knew. It will count them, add them up and take averages of them, and it will never once mention it.
And notice what you were not doing. You were not lying. The form demanded a value and you supplied one, because that is what the form was for. That is how almost all bad data gets made. Not by liars. By ordinary people with a box that will not let them past.
Keep the sheet and keep your two numbers: how many you knew, and how many you invented. In Part 3 you are going to meet somebody else's round two, and there will be 3,705 of them.
Five words. Read them once. Then write each one in your own words in your notebook.
Write your five lines. Stand-in is the one that matters. If you cannot write that one in your own words, read Part 0 again.
Open band-00.ipynb. It sits in the course folder, in the file list down the left of JupyterLab.
Run every cell from the top before you start. Nothing is remembered overnight, and the cells later in this step need names that were made earlier in it.
If you see NameError: name 'df' is not defined, that is all this is. Nothing is broken.
Check that every random_state=0 on the page is in your cell too. Without it the rows are shuffled differently every run, so your numbers change every time you press Shift and Enter and no two people in the world can compare notes.
If the numbers still differ, run every cell from the top in a fresh kernel: Kernel, then Restart Kernel and Run All Cells.
If none of that means anything yet, stop and do the setup page first. It takes about an hour and twenty minutes and it assumes you have never installed anything in your life.
Then run the cells from the top. Do not skip any.
It does not stay running. You start it fresh every time you sit down to work, and after a break it will be gone.
Open a terminal the way Part 4 of the setup page shows. Move into the course folder the way Part 6 shows. Then type jupyter lab and press Enter.
Load the file.
import pandas as pd
df = pd.read_csv("data/bank.csv", sep=";")
print(df.shape)
Five small things are happening. Read them once now. You will see them in every step from here on.
import pandas as pd brings in pandas, the library that handles tables, and gives it the short name pd so you do not have to type the whole word. Every notebook in this course brings in pandas in its first cell, so you will type or run this line more than any other.pd.read_csv means ask pandas to read a file of this kind. CSV is the plainest way to save a table.sep=";" tells it that this file separates its columns with a semicolon. Most files use a comma, so most of the time you leave this off. Leave it off here and all 17 columns get squashed into one.df is the name we are giving the table. It is short for data frame, which is what pandas calls a table. Everybody in this field calls it df, so you may as well start now.df["age"]. That means the column called age, out of the table called df. The square brackets mean reach inside. The quotes mean this is a name I am typing, not something for Python to go and look up. The name has to match the column exactly, capital letters and all.Nothing here is yours to invent yet. You are reading working code and running it, which is the whole of Step 0. You start writing in Step 1.
(4521, 17)
That is 4,521 rows and 17 columns. If you see (4521, 1) you left out the sep.
Pandas is not installed where this notebook is running. Put %pip install pandas in a cell and run it. The % matters. Then restart the kernel and run from the top.
Python looked for data/bank.csv and it was not there. That means JupyterLab was started somewhere other than the course folder.
Save your notebook. Close the browser tab. Go to the terminal window you started Jupyter from and press Control and C together to stop it. Then follow Part 6 of the setup page again, so your terminal is standing in the course folder, and start jupyter lab from there.
Do not move the notebook to a different folder to make the error go away. It will come back in Step 1.
Name the columns.
print(list(df.columns))
Seventeen names, starting age, job, marital, education.
Read all seventeen out loud. You will meet pdays in Part 3 and it is the reason this step exists.
Look at three rows.
df.head(3)
Three rows and 17 columns. Row 0 begins 30 unemployed married primary. Row 2 begins 35 management single tertiary.
If your first row is not the 30 year old, you have opened a different file. Check step 1.
Take an average that is honest.
print(round(df["age"].mean(), 2))
print(df["age"].min(), df["age"].max())
41.17, then 19 87
The youngest is 19 and the oldest is 87. An average age of 41.17 sits between them and makes sense. Hold on to that feeling. Part 3 takes it away.
There is a column called pdays. Going by the name, it holds how many days ago this person was last called. Hold that loosely: Part 3 is about to show you that a column name is a claim, not a fact.
These are bank clients in a campaign that ran over a year or so. Some were called before, some were not.
What do you think the average of pdays is? Write one number in your notebook. A rough guess is fine. Writing it is the point.
print(round(df["pdays"].mean(), 2))
39.77
So the average client was last called about 40 days ago. That sounds fine. It is completely wrong.
Somewhere below, a cell tries to print an average and prints something that starts <bound method instead of a number.
It does not crash. It prints a thing that looks like output.
Find it. The fix is two characters. Then write one line on what you learn from the fact that Python did not complain.
Now look at the column instead of the average.
print(df["pdays"].min(), df["pdays"].max())
print((df["pdays"] == -1).sum())
-1 871, then 3705
The smallest value is minus one. You cannot be called minus one days ago. And 3,705 of the 4,521 rows hold it.
Do not guess what minus one means. Go and read.
notes = open("data/bank-names.txt").read()
for line in notes.split("\n"):
if "pdays" in line:
print(line)
That opens the notes file, cuts it into lines, and prints only the lines with the word pdays in them. You are not expected to be able to write that yet. You are expected to run it and read what comes back.
A line saying that -1 means the client was not previously contacted.
So minus one is not a number of days. It is a stand-in, exactly like the lines you invented in round two of Part 0. Somebody had to put something in the box, and they put minus one.
Read the whole line, not the last half of it. It says from a previous campaign. So these 3,705 people are not people the bank had never rung. They are people the bank had not rung in an earlier campaign, and two thirds of them were rung more than once during this one. The difference will matter in Step 1.
Now take the average again, this time of the people it actually applies to.
called_before = df["pdays"] != -1
print(called_before.sum())
print(round(df.loc[called_before, "pdays"].mean(), 2))
816, then 224.87
The real answer is 224.87 days, not 39.77. It is more than five times bigger.
It was 816 real day counts and 3,705 minus ones, all added up and divided by 4,521.
It is not a wrong average of the right thing. It is an average of two things that have nothing to do with each other. It has no meaning at all.
Nothing warned you. The number printed. It had two decimal places. It looked like every other number you have ever trusted.
First, one new thing you will use constantly. To see what is actually inside a column of words, attach .value_counts() to it.
print(df["job"].value_counts())
A list of jobs with a count beside each, biggest first. management is top with 969, then blue-collar with 946, then technician with 768.
Near the bottom, unknown with 38. Hold that thought.
Now, pandas has a tool for finding missing values. Run it.
print(df.isna().sum().sum())
0
Zero missing values in the whole table. Every box is filled. The data is clean.
It is not clean. You already know that, because you found 3,705 stand-ins in Part 3, and this said zero.
Now count the word unknown in every column that has it. This next one is a loop. It does the same thing once for each column, and only prints when it finds something.
for name in df.columns:
hits = (df[name] == "unknown").sum()
if hits > 0:
print(name, hits)
job 38, education 187, contact 1324, poutcome 3705
That is 5,254 boxes holding the word unknown, in a table that just told you nothing was missing.
has_unknown = (df[["job", "education", "contact", "poutcome"]] == "unknown").any(axis=1)
print(has_unknown.sum())
3757
So 3,757 of your 4,521 rows hold the word unknown somewhere. That is five rows in every six.
3,705 of those 3,757 are the poutcome column, and you already know what that one means: there was no earlier campaign, so there is no outcome to record. Nobody failed to write it down. It could not exist.
Take poutcome out and run the same line on the other three columns. You get 1,457 rows, which is the honest count of rows where something really was not recorded.
Both numbers are true and they mean different things. The word unknown was doing two jobs in one table, and telling them apart is the work.
isna() finds empty boxes. It does not find lies.
It is a little cleverer than that, and it is worth knowing where the cleverness stops. When pandas reads a file it quietly turns a short list of spellings into empty boxes for you, including NA, N/A, NULL, None and nan. Those it does find. The word unknown, a dash, a zero and a minus one are not on the list, and nothing will ever put them on it, because only you know what they mean here.
Before you trust any column, look at what is actually in it. value_counts() for words. min() and max() for numbers. Every time.
One catch with value_counts(), since you are about to lean on it: by default it leaves the empty boxes out of its list entirely. Ask for value_counts(dropna=False) when you want to see them counted alongside everything else.
You are now looking for problems, so you will start seeing them everywhere. Some of them are not problems. This part is how you tell.
print(df["balance"].min(), df["balance"].max())
print((df["balance"] < 0).sum())
-3313 71188, then 366
366 people have less than nothing in the bank.
Is that broken data? No. It is an overdraft. People really do owe the bank money. This one is fine.
print(df["campaign"].max())
print((df["campaign"] > 10).sum())
50, then 130
Somebody was rung 50 times in one campaign. 130 people were rung more than ten times.
Is that broken data? Probably not. It is probably true, and that is worse. Stop and picture the person who answered the phone on the fiftieth call.
You will build things that decide who gets rung. Write that number down now, while it still bothers you.
You cannot tell by looking at the number. Minus 3,313 and minus 1 look equally strange, and one is real.
You tell by asking what the column means, and whether the strange value could happen in the world. Minus 3,313 euros can happen. Minus 1 days cannot.
There is a column for the day of the month and a column for the month.
print(df["day"].min(), df["day"].max())
print(sorted(df["month"].unique()))
1 31, then twelve short month names, in alphabetical order.
Now look for a year.
print([name for name in df.columns if "year" in name])
[]
An empty list. No column has the word year in its name.
That is not quite a proof, and you should notice the difference. A column called yr or when would have slipped past it. What settles it is the list of seventeen names you printed in Part 2: go back and read them. No year, no date, nothing that stands for one.
The notes say the full file runs from May 2008 to November 2010. So two rows saying 5 May could be a year apart, or two years apart, and nothing in the table can tell you which came first.
So you cannot put these calls in order. You cannot ask whether the bank got better over time. You cannot split the old calls from the new ones.
You cannot fix it. There is no clever code that puts the year back.
What you do instead is write it down, in the part of your README that says what is wrong. Then anybody who reads your work knows the question you could not answer.
Half of doing this job well is saying clearly what your data cannot tell you.
This part has no code in it and no answer at the end. It is also the most important thing you will do in this step, because the table you pick here is the table you carry through every step after it.
It must be about a real thing near you. Your work, your town, your hobby, your family's business. Not a famous practice file.
Good places to look: a spreadsheet somebody at your work already keeps, an open data site run by your city or government, a sports league's published results, the export button on an app you use every day.
.csv. If yours is a spreadsheet ending in .xlsx, open it, choose File then Save As, and pick CSV from the list of formats. In Google Sheets it is File, then Download, then CSV. A CSV holds the same table with the colours and the formulas replaced by the values they worked out, which is all you need.These are the six it runs, in the order it prints them, and the page and the script agree because I went and read the script.
--i-may-use-this at the end.It is a terminal command. Your first terminal is busy running JupyterLab, so leave that one alone and open a second one, the same way Part 4 of the setup page shows. Then move it into the course folder the way Part 6 shows.
To fill in the path to your file without typing it, type the command up to the space, then drag your file onto the terminal window and let go. The path appears by itself.
python3 tools/check_my_table.py path/to/your.csv name_of_the_column_to_guess --i-may-use-this
Six results, each one a line beginning [ok] with a line of detail under it, and then PASSED. All 6 checks.
If any line says [no], it tells you what is wrong and why it matters. Fix it or find a different table. Do not carry a failing table into Step 1.
Under the checks it prints a list headed Things to look at. That is not pass or fail. It is a head start on your three faults, and it is only a head start. It will miss some.
The checker cannot tell whether you have permission. It takes your word.
If your table names real people, somebody has to have actually said yes, in writing, before you type that flag. If they have not, this is not the table.
One page. Not two. It says:
You have not looked properly. Every real table has three. Go through the columns one at a time and run value_counts() or min() and max() on every single one.
If you still cannot, one of your three is this: nobody wrote down where this data came from, so you do not know what it leaves out.
Put the one page away for a day. Come back and read it cold.
Then ask it one question: if somebody used this table to make a decision about a person, who would that decision be unfair to? Write the answer at the bottom.
Make a folder for your own work. Put three things in it.
README.md. A README is just a plain text file saying what is in a folder. The .md stands for Markdown, which is plain text with a few symbols in it for headings and lists. The easiest way to make one: in JupyterLab's file list, click the plus button, choose Text File, type your page, then right click the file and rename it to README.md.Keep the notebook even though it is messy. It is proof that you looked, and every step after this one adds to it.
Print the one page. Show it to somebody who has never seen your table.
Ask them one question: after reading this, would you trust the table? Write down what they say, word for word.
Six questions. Answer them in writing, in your notebook. Do not open the answers until all six are written down.
One of them counts more than the other five. You can get five right and still not pass this step.
df.shape tell you, and in what order?df.isna().sum().sum() prints 0. Name two ways the table could still be full of holes.Your file has 4,521 rows. You print the average of pdays and get 39.8.
Then you notice 3,705 of those rows hold -1, and the file's notes say -1 means the client was never called before.
What was your 39.8? Say what it is an average of.
(4521, 17) is 4,521 rows and 17 columns.-1; the word unknown or none or n/a written in as text; a zero standing in for nothing; a made up default date. isna() only finds empty boxes.Go back to Part 3 and do it again with a fresh page. Then read the answer, write it out again tomorrow from memory, and go on. That counts.
Step 1 asks you to trust a score you worked out yourself. You cannot do that until you know how a number can look right and mean nothing.
Go through all 17 columns of the bank file, one at a time. Use value_counts() on the word columns and min() and max() on the number ones.
Write a list of every column you would not trust, with the reason. You should find more than the four in Part 4.
The 3,705 rows with pdays of minus one are the same 3,705 rows where poutcome is unknown. Prove it in one line of code.
Then say what that tells you about how the file was built.
Write the sentence you would put at the top of this table if you were handing it to somebody else. One sentence, under 30 words, that stops them making the mistake you made in Part 3.
Then go and find a real table at your work or your school and write the same sentence for that one.
All six of these, or it is a not yet.
check_my_table.py prints PASSED. This is the gate into Step 1 and there is no way around it.Then look at the form you filled in at the start of Part 0. What you wrote in round two had a name all along. It is pdays.