Dataset notation, learning algorithm, model, loss, optimization and evaluation
We want to predict something unknown for a new case.
Example: risk score, mark, house price, spam label.
We use past examples where correct answers are already known.
This is the labeled dataset.
A learning algorithm chooses a function that predicts future answers.
Then we evaluate on unseen data.
Supervised learning means the computer learns from examples that contain both:
Input: information available before prediction.
Target: the correct answer we want to learn.
Input: study hours, attendance, previous quiz marks.
Target: final exam mark or pass/fail label.
Because the target is known during training, the algorithm receives supervision.
i is the example number.
x⁽ⁱ⁾ is the input of example i.
y⁽ⁱ⁾ is the correct target of example i.
Example: x⁽³⁾ = BMI of patient 3, y⁽³⁾ = risk score of patient 3.
n = number of examples or rows.
d = number of features or input columns.
X = feature matrix.
y = target vector.
A dataset is not only a file. It is the evidence used by the learning algorithm.
Each row is one example. Each input column is one feature. The target column is the answer.
Rows = examplesColumns = featuresTarget = label/value
| Example | age | BMI | BP | glucose | target y |
|---|---|---|---|---|---|
| 1 | 45 | 27.3 | 84 | 106 | 233 |
| 2 | 51 | 23.8 | 82 | 92 | 91 |
| 3 | 38 | 25.3 | 79 | 96 | 111 |
| 4 | 57 | 23.8 | 91 | 121 | 152 |
| 5 | 44 | 24.0 | 88 | 109 | 120 |
Use BMI as input and predict a quantitative disease progression score.
This one-feature view is used for teaching. Real datasets may use many features at the same time.
The output is a number.
Risk score, temperature, price, mark.
The output has two classes.
Spam/not spam, pass/fail, fraud/not fraud.
The output has many classes.
Digit 0-9, disease type, document category.
A model is a function that maps input features to a predicted output.
θ represents model parameters learned from data.
ŷ is the prediction, not necessarily the true answer.
θ₁ controls how much prediction changes when x changes.
θ₀ shifts the prediction up or down.
For multiple features: ŷ = θ₀ + θ₁x₁ + θ₂x₂ + ... + θdxᵈ.
Blue points are training examples. Red line is your current model.
This is mean squared error for regression.
The algorithm tries many possible models and chooses the one with small error.
For linear regression, it learns the best slope and intercept.
The exact optimization method depends on the algorithm.
Example: linear model, decision tree, neural network.
Choose what “mistake” means for the task.
Search for parameters that reduce the loss.
Check performance on unseen data.
Typical split shown for teaching. Exact ratios depend on data size and project requirements.
If we test on the same data used for training, we may overestimate performance.
The test set gives a more honest estimate of performance on new examples.
A model should not only memorize training examples.
It should perform well on new examples from the same problem setting.
This is called generalization.
Regression: target y is numerical and continuous.
Classification: target y is a category/class.
The type of y usually determines the type of supervised learning task.
May contain missing values, text categories, outliers and inconsistent formats.
Numerical features, clean labels, proper train/test split and clear task definition.
In practice, dataset preparation often takes more time than fitting the first model.
Easy to interpret in original unit.
Punishes large errors strongly.
Useful for balanced classification.
Predict the average target value for every case.
Predict the most common class for every case.
A model is useful only when it beats a simple reasonable baseline.
Celsius to Fahrenheit, grading formula, sorting, tax calculation.
Spam detection, fraud detection, risk prediction, image classification.
A bank wants to predict whether a transaction is fraudulent. It has past transactions with labels: fraud or legitimate. Identify x, y, and task type.
from sklearn import datasets, linear_model
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, mean_absolute_error
# Load dataset
X, y = datasets.load_diabetes(return_X_y=True, as_frame=True)
X = X[['bmi']] * 30 + 25 # one feature for visualization
# Split data
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
# Learning algorithm
model = linear_model.LinearRegression()
model.fit(X_train, y_train)
# Evaluation
pred = model.predict(X_test)
print('theta_1:', model.coef_[0])
print('theta_0:', model.intercept_)
print('MAE:', mean_absolute_error(y_test, pred))
print('MSE:', mean_squared_error(y_test, pred))
A supervised learning problem begins with labeled examples.
The dataset is written as D = {(x⁽ⁱ⁾, y⁽ⁱ⁾)}, where x is input and y is target.
The model predicts ŷ = fθ(x); the loss measures the gap between y and ŷ.
Training means finding parameters θ* that minimize loss.
Evaluation on unseen data checks whether the model generalizes.
Next topic: Linear Regression in detail