The Clubhouse Rule
Kodi will try not to use a technical word before you have experienced the idea that the word names.
Why Would a Machine Need to Learn?
Before “machine learning”, understand the problem that ordinary rules cannot always solve.
🍊 The Fruit Shop Problem
You own a fruit shop. Your new worker must separate ripe oranges from unripe oranges.
You could write a rule: “If orange colour is deep enough, call it ripe.” But what about a ripe fruit that is still partly green? What about softness? Smell? Variety? Damage?
🎮 Playground · Rule Writer vs Changing World
Choose a rule for spam. Then Kodi changes the message.
🤖 Now We Can Name the Idea
Normal programming: a human writes the rule and the computer applies it.
One important form of machine learning: we provide examples and known answers, and an algorithm estimates a model that can predict answers for new examples.
Teach Kodi Without Writing the Rule
Experience supervised learning before learning its vocabulary.
🥭 “Kodi Has Never Seen a Ripe Mango”
You are the teacher. Kodi sees four mangoes and you tell Kodi the correct answer for each one.
| Colour | Firmness | Smell | Your answer |
|---|---|---|---|
| Yellow | Soft | Sweet | Ripe |
| Green | Hard | Little smell | Not ripe |
| Yellow-green | Slightly soft | Sweet | Ripe |
| Green | Very hard | Little smell | Not ripe |
🎮 Playground · New Mango
A new mango is yellow-green, soft and smells sweet. Kodi has never seen this exact mango. What should a useful learner do?
🧠 Vocabulary Reveal
What Does “Learning” Actually Mean?
Not magic. Not consciousness. Parameters are adjusted so predictions better match examples.
🎚 Imagine a Scoring Machine
Suppose Kodi is estimating loan risk. Each clue contributes some influence to a score. In a simple model, learning can mean finding useful numerical settings—often called parameters or weights—from data.
🎮 Playground · Adjust the Weights Yourself
How Does the Machine Know It Is Wrong?
Learning needs feedback: compare prediction with truth, measure error, adjust, repeat.
The Learning Loop
The function that measures how wrong a model is is often called a loss function. Different tasks use different losses.
🎯 Playground · Find the Better Line
Imagine actual house price = 50 + 3 × size. Move Kodi's slope closer to the pattern and watch average error fall.
Two Supervised Questions: “Which Kind?” and “How Much?”
📦 Which category?
Spam/not spam, diseased/healthy, churn/stay.
📏 How much?
House price, demand, delivery time, revenue.
🎮 Playground · Sort the Question
What If Nobody Knows the Answers?
Now remove the labels and see what kind of learning remains possible.
🛒 40,000 Shoppers, No Customer Types
A supermarket has purchase histories but nobody has labelled customers “VIP”, “bargain hunter” or “at risk”. You can still search for customers that behave similarly.
🎮 Playground · Let Groups Emerge
What If There Are No Correct Answers, Only Consequences?
🤖 Kodi in a Maze
Kodi can move left, right, up or down. Nobody provides the correct move for every square.
Reach the apple: +10. Hit a wall: −2. Every unnecessary move: −1.
🎮 Playground · Experience Shapes Preference
Before Modelling: What Is One Row?
A surprisingly important practitioner question.
💼 “Predict Customer Churn” Is Still Too Vague
What exactly is a customer example? One customer ever? One customer each month? One customer at a particular prediction date?
🧭 Playground · Frame the Problem
Data Is the Raw Material — and It Is Usually Messy
🧺 Imagine Cooking With Unsorted Ingredients
Some tomatoes are rotten, some labels are missing, salt is stored in a sugar bag, and the same onion appears twice on the inventory list. A brilliant recipe cannot rescue ingredients you do not understand.
🕵️ Playground · Data Detective
Click suspicious cells.
| Customer | State | Sales | Date |
|---|---|---|---|
| Amaka | Lagos | ₦5,000 | 12/4/26 |
| AMAKA | lagos | 5000 | 12-04-2026 |
| Emeka | LAG | N/A | April 12 |
| John | Abuja | ₦-900 | 14/04/26 |
| ABUJA | 7300 |
Features: Turn Raw History Into Useful Clues
🧠 Humans Do This Naturally
If you see four purchases spread across months, you may mentally ask: How recently did this person buy? How often? How much do they usually spend?
Those derived clues can become model features.
🏭 Playground · Feature Factory
| Date | Amount |
|---|---|
| Jan 2 | ₦4,300 |
| Jan 18 | ₦12,000 |
| Feb 1 | ₦2,500 |
| Apr 9 | ₦7,600 |
The Exam Analogy: Train, Validate, Test
Overfitting: The Student Who Memorised the Answers
📚 Ada Scores 100% on the Practice Sheet
But the teacher changes the numbers in the final exam and Ada fails. She memorised the sheet instead of learning the underlying method.
🎮 Playground · Make the Model Too Clever
Data Leakage: Kodi Saw the Answer Sheet
📝 The 100% Student
Kodi gets every exam question right. Then you discover the final answers were printed faintly at the bottom of every page.
That is not intelligence. The evaluation has been contaminated.
🚨 Playground · Catch the Answer Sheet
Accuracy Can Be a Beautiful Lie
🚨 The Security Guard Who Says “Everybody Is Innocent”
If only 1 person in 1,000 is a thief, the guard can say “innocent” to everybody and be correct 99.9% of the time—while catching zero thieves.
Confusion Matrix — Four Possible Outcomes
caught fraud
false alarm
missed fraud
correct clear
🎚 Playground · Security Threshold
Models Are Different Tools, Not Magic Rankings
Only now—after you understand the problem, data, target, features, generalisation and evaluation—do model names become useful.
Linear Regression
Fits a numerical relationship. Good baseline for some regression problems.
Logistic Regression
Estimates class probability using a relatively simple decision structure.
Decision Tree
Learns branching questions. Easy to visualise; deep trees can overfit.
Random Forest
Combines many varied trees so their votes can be more robust.
Gradient Boosting
Builds models sequentially to correct previous errors; powerful on many tabular tasks.
Neural Networks
Flexible layered models useful across images, text, audio and many complex tasks.
Reward Hacking: The Machine Did Exactly What You Asked
🧹 The Cleaning Robot
You reward a robot every time it collects dirt. It discovers that dumping dirt from its own bin and collecting it again earns endless points.
The robot did not misunderstand the score. You wrote a score that failed to capture the real goal.
🎮 Playground · Fix the Reward
Bias and Fairness: Learning History Can Repeat History
⚖️ “The Data Says So” Is Not the End of the Conversation
Historical records can contain patterns created by previous human decisions, unequal access, measurement choices or missing populations. A model can scale those patterns.
Responsible ML asks who may be harmed, whether features are legitimate, how error rates differ, what human oversight exists and how decisions can be challenged.
Distribution Shift: Yesterday's World Is Not Today's World
🌍 The Recipe Changed
A demand model learns from years when prices, policy and customer behaviour followed one pattern. Then a major event changes the environment. The code may still run perfectly while the model becomes less useful.
🎮 Playground · Move the World
Now Read the Python — Every Line Has a Meaning
First Read It in English
Then Read the Code
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report
X = df[["days_since_last_order", "orders_90d", "spend_90d"]]
y = df["churned_30d"]
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.20, random_state=42, stratify=y
)
model = RandomForestClassifier(random_state=42)
model.fit(X_train, y_train)
predictions = model.predict(X_test)
print(classification_report(y_test, predictions))
Memory hook
X = clues. y = answer. fit = learn. predict = use what was learned. metrics = check whether it actually worked.
A Notebook Is Not Yet a Working Product
🚀 From Model to Real System
This surrounding discipline—versioning, serving, monitoring, reproducibility, retraining and rollback—is part of MLOps.
What an ML Practitioner Actually Does
💼 Tuesday Morning
Capstone: Save KoMart
Now combine the principles rather than answering isolated questions.
🏢 Mission
KoMart has 150,000 transaction records. Management says, “Customers are disappearing. Help us intervene before they leave.”
Your deliverable is not merely a model. It is a defensible ML system and a portfolio story.
🏆 Practitioner Evidence
🧠 Final Neural Pathway
If the learner can explain that chain in ordinary language, demonstrate it in code, and defend the choices in a real project, the vocabulary is no longer floating knowledge. It has a structure.