CSE-403 Machine Learning

Elements of a Supervised Learning Problem

Dataset notation, learning algorithm, model, loss, optimization and evaluation

Md Atikuzzaman
Lecturer
Department of Computer Science & Engineering
atik@cse.green.edu.bd

Today’s Storyline

1. Start with a problem

We want to predict something unknown for a new case.

Example: risk score, mark, house price, spam label.

2. Learn from evidence

We use past examples where correct answers are already known.

This is the labeled dataset.

3. Build a model

A learning algorithm chooses a function that predicts future answers.

Then we evaluate on unseen data.

What Makes It “Supervised”?

Supervised learning means the computer learns from examples that contain both:

Input: information available before prediction.

Target: the correct answer we want to learn.

Input x + correct output y → training example (x, y)

Simple classroom example

Input: study hours, attendance, previous quiz marks.

Target: final exam mark or pass/fail label.

Because the target is known during training, the algorithm receives supervision.

High-Level Structure

Dataset
labeled examples
(x, y)
Algorithm
minimize loss
learn parameters
Model
ŷ = fθ(x)
Prediction
new x → ŷ
Core equation
Dataset + Algorithm → Predictive Model

What each part does

  • Dataset supplies examples.
  • Algorithm learns from the dataset.
  • Model stores the learned pattern.
  • Prediction applies the model to new data.

Notation: One Training Example

x⁽ⁱ⁾ ∈ X
y⁽ⁱ⁾ ∈ Y
(x⁽ⁱ⁾, y⁽ⁱ⁾) = one labeled example

How to read this

i is the example number.

x⁽ⁱ⁾ is the input of example i.

y⁽ⁱ⁾ is the correct target of example i.

Example: x⁽³⁾ = BMI of patient 3, y⁽³⁾ = risk score of patient 3.

Notation: Full Dataset

D = {(x⁽ⁱ⁾, y⁽ⁱ⁾)}ᵢ₌₁ⁿ

X ∈ Rⁿˣᵈ
y ∈ Yⁿ

Meaning

n = number of examples or rows.

d = number of features or input columns.

X = feature matrix.

y = target vector.

A dataset is not only a file. It is the evidence used by the learning algorithm.

Dataset Anatomy

Rows and columns

Each row is one example. Each input column is one feature. The target column is the answer.

Rows = examplesColumns = featuresTarget = label/value

ExampleageBMIBPglucosetarget y
14527.384106233
25123.8829291
33825.37996111
45723.891121152
54424.088109120

Dataset Example: Diabetes Risk

Prediction task

Use BMI as input and predict a quantitative disease progression score.

x = BMI
y = risk score

This one-feature view is used for teaching. Real datasets may use many features at the same time.

Input Space and Output Space

Regression

The output is a number.

Y = R

Risk score, temperature, price, mark.

Binary classification

The output has two classes.

Y = {0, 1}

Spam/not spam, pass/fail, fraud/not fraud.

Multiclass classification

The output has many classes.

Y = {1, 2, ..., K}

Digit 0-9, disease type, document category.

Model as a Function

fθ : X → Y

ŷ = fθ(x)

Key idea

A model is a function that maps input features to a predicted output.

θ represents model parameters learned from data.

ŷ is the prediction, not necessarily the true answer.

Linear Model for Regression

ŷ = θ₀ + θ₁x

θ₀ = intercept
θ₁ = slope

Student-friendly meaning

θ₁ controls how much prediction changes when x changes.

θ₀ shifts the prediction up or down.

For multiple features: ŷ = θ₀ + θ₁x₁ + θ₂x₂ + ... + θdxᵈ.

Interactive Model: Change θ₁ and θ₀

Try the line

ŷ = 7.00x + 25.00

Blue points are training examples. Red line is your current model.

Loss Function: Measuring Mistake

For one example

error⁽ⁱ⁾ = y⁽ⁱ⁾ − ŷ⁽ⁱ⁾
L(y, ŷ) = (y − ŷ)²

For the whole training set

J(θ) = 1/n Σᵢ₌₁ⁿ (y⁽ⁱ⁾ − fθ(x⁽ⁱ⁾))²

This is mean squared error for regression.

Learning Means Optimization

θ* = argminθ J(θ)

Find parameters that minimize training loss.

Plain English

The algorithm tries many possible models and chooses the one with small error.

For linear regression, it learns the best slope and intercept.

The exact optimization method depends on the algorithm.

Learning Algorithm Overview

1. Choose model

Example: linear model, decision tree, neural network.

2. Define loss

Choose what “mistake” means for the task.

3. Optimize

Search for parameters that reduce the loss.

4. Evaluate

Check performance on unseen data.

Training, Validation and Test Split

Training
70%
Validation
15%
Test
15%

Typical split shown for teaching. Exact ratios depend on data size and project requirements.

Why not use all data for training?

If we test on the same data used for training, we may overestimate performance.

The test set gives a more honest estimate of performance on new examples.

Model Generalization

Good supervised learning

A model should not only memorize training examples.

It should perform well on new examples from the same problem setting.

This is called generalization.

Regression vs Classification

How to decide?

Regression: target y is numerical and continuous.

Classification: target y is a category/class.

The type of y usually determines the type of supervised learning task.

From Raw Data to ML-Ready Dataset

Raw data

May contain missing values, text categories, outliers and inconsistent formats.

→

ML-ready data

Numerical features, clean labels, proper train/test split and clear task definition.

In practice, dataset preparation often takes more time than fitting the first model.

Practical Dataset Checklist

Before training

  • What is one row?
  • What are the input features?
  • What is the target?
  • Is the target available for all training examples?
  • Are there missing values or leakage?

Before reporting

  • How was data split?
  • Which metric is used?
  • What baseline is compared?
  • Does the test set represent the real use case?
  • What limitations remain?

Evaluation Metrics

MAE

MAE = average(|y − ŷ|)

Easy to interpret in original unit.

MSE

MSE = average((y − ŷ)²)

Punishes large errors strongly.

Accuracy

Accuracy = correct / total

Useful for balanced classification.

Baseline: Always Compare

Regression baseline

Predict the average target value for every case.

ŷ = mean(y_train)

Classification baseline

Predict the most common class for every case.

ŷ = majority class

A model is useful only when it beats a simple reasonable baseline.

When to Use Traditional Coding vs ML

Traditional coding

  • Exact rule is known.
  • Logic is stable and simple.
  • Correct output must follow a formula.

Celsius to Fahrenheit, grading formula, sorting, tax calculation.

Machine learning

  • Rule is hard to write manually.
  • Many examples exist.
  • Pattern depends on many variables.
  • Some error is acceptable and measurable.

Spam detection, fraud detection, risk prediction, image classification.

Mini Activity

A bank wants to predict whether a transaction is fraudulent. It has past transactions with labels: fraud or legitimate. Identify x, y, and task type.

Colab Code: Dataset and Model


from sklearn import datasets, linear_model
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, mean_absolute_error

# Load dataset
X, y = datasets.load_diabetes(return_X_y=True, as_frame=True)
X = X[['bmi']] * 30 + 25       # one feature for visualization

# Split data
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

# Learning algorithm
model = linear_model.LinearRegression()
model.fit(X_train, y_train)

# Evaluation
pred = model.predict(X_test)
print('theta_1:', model.coef_[0])
print('theta_0:', model.intercept_)
print('MAE:', mean_absolute_error(y_test, pred))
print('MSE:', mean_squared_error(y_test, pred))
  

Standard Supervised Learning Template

1. Define task: X → Y
2. Build dataset D = {(x⁽ⁱ⁾, y⁽ⁱ⁾)}ᵢ₌₁ⁿ
3. Choose model family fθ
4. Choose loss L(y, fθ(x))
5. Learn θ* = argminθ J(θ)
6. Evaluate on unseen data

Key Takeaways

A supervised learning problem begins with labeled examples.

The dataset is written as D = {(x⁽ⁱ⁾, y⁽ⁱ⁾)}, where x is input and y is target.

The model predicts ŷ = fθ(x); the loss measures the gap between y and ŷ.

Training means finding parameters θ* that minimize loss.

Evaluation on unseen data checks whether the model generalizes.

Questions?

Next topic: Linear Regression in detail