The following notes come from the course Machine Learning in Finance and Insurance, taught at ETH Zurich by Patrick Cheridito, Professor of Mathematics at ETH Zurich and Director of RiskLab Switzerland.
Feel free to reach out to me with any comments, clarifications, or corrections.
Index
- Basic Notions of Statistical Learning
- Linear Regression
- Gradient Descent
- Logistic Regression
- Kernel Methods
- Neural Networks
- Classification and Regression Trees
- Bagging and Random Forests
- Gradient Boosted Trees
- Dimensionality Reduction and Autoencoders
- Generative Models
The following notes come from the course Machine Learning in Finance and Insurance, taught at ETH Zurich by Patrick Cheridito, Professor of Mathematics at ETH Zurich and Director of RiskLab Switzerland.
Feel free to reach out to me with any comments, clarifications, or corrections.
1 Basic Notions of Statistical Learning
Definition -algebraFirstly, we remember that a -algebra is a non-empty collection of subsets of that contains itself and is closed under complementation and countable unions, i.e. it satisfies the following:
- contains the whole space
- closed under complementation:
- it is closed under countable unions:
From this definition, we provide some examples:
- Trivial -algebra: (the smallest -algebra)
- Power set: (collection of all subsets of , the largest possible -algebra)
- Borel -algebra (smallest -algebra generated by all open sets in , it is the essential foundation for measure theory and real analysis)
Intuition: in , consider all the intervals and perform all three actions that define the -algebra (complements, countable unions, countable intersections, where the last one nicely derives from the definition of -algebra): the resulting set is the Borel -algebra and represents every subset of that can be obtained using standard mathematical operations.)
We also recall that a function is Borel measurable if:
Intuition: if the preimage of every Borel set in is a Borel set in .
The prediction problem
Consider the following:
- a probability space , i.e.:
- a non-empty set , the sample space
- , a -algebra of subsets of ,
- a probability measure
- a feature space with Borel -algebra (example: )
- a closed or open target space with Borel -algebra (example: )
- some random elements (inputs), (outputs)
We then consider the distributions of and , that is the regular conditional distribution of given , that is:
- is a probability measure on
- , is -measurable
- , -almost surely.
Then, it follows, for all -integrable , we have the disintegration of :
We now go back to the prediction problem: we want to find a measurable function s.t. , i.e. a function that allows us to minimize the expected loss (also called risk), defined as:
for a measurable loss function .
Then, in machine learning, it is common to solve the empirical loss minimization problem!
Regression Functions and Irreducible Error Lemma (theoretically best predictor with respect to a given loss function )
Let be a measurable function such that, for -almost every , minimizes
Then
for every measurable function .
Note that is the ideal predictor we would use if we knew the true conditional distribution of given . For each , if we predict a value , the conditional expected loss is
The function chooses the prediction that minimizes this conditional risk:
From now on, we suppose these assumptions from the lemma hold and we call:
- regression function (or, more generally, the Bayesian-optimal predictor)
- irreducible error (even with the exact conditional distribution, randomness in given can prevent perfect prediction)
Example: Square loss
We consider now the following example with , and a random variable with .
Then, :
so the mapping is strictly convex with:
- a unique minimizer (under square loss: the best constant prediction of is its mean)
- a minimal value of (the minimal possible error is its variance)
We now assume , by disintegration we obtain:
Therefore we have
for -almost every . By the Cauchy–Schwarz inequality, this also implies
for -almost every .
The conditional risk under squared loss is then:
The second term does not depend on , and the first one is minimized when:
Therefore, the unique minimizer is:
Then, we easily obtain the least-squares regression function and its minimum loss (risk):
and
Example: pinball loss
We now consider and a loss function , with and the pinball loss function defined as:
In this formula, is the quantile we want to predict and it determines the asymmetry of the loss. Indeed, in the pinball loss function, the penalty is proportional to , therefore:
- for : underprediction is penalized more
- for : overprediction is penalized more
- for : both directional errors have the same cost
We firstly note that:
We now consider a random variable with . In this case, is convex in and:
So is a minimizer of if and only if and , if and only if is an -quantile of .
Now let , for , be an -quantile of the conditional distribution . Then we have:
- (quantile regression)
For the special case , is a median of , and is the minimal MAE.
Empirical loss (empirical risk)
Let with be independent copies of , called training data.
The empirical loss/risk of a measurable function is . In the case of , we can write:
and we can call this weak law of large numbers.
Intuition: with the training data growing, variance approaches zero.
Hypothesis Classes
A hypothesis class is a family of measurable functions , e.g.:
- all affine functions (i.e. weighted sums of features )
- 2nd order polynomials
- splines (local polynomial pieces joined together, producing a smooth flexible curve)
- trees (sequential if-then rules with piecewise constant regions)
- support vector machines (separate classes with a maximum-margin boundary)
- neural networks (several nonlinear transformations)
Ideally, we would like to minimize over all measurable functions .
However, in practice, we try to find a numerical solution to the empirical loss minimization problem:
Note that can be produced by either:
- a deterministic algorithm, in the form , where is a measurable function of the training data;
- a stochastic optimization algorithm, in the form , where is a random element independent of the training data and is measurable (note that a stochastic algorithm can produce different results with the same training data due to the presence of the noisy term )
Given the numerical solution , we now formalize the expected loss decomposition for the square loss function.
We first define the conditional expected loss given that :
If we define:
then we can write the conditional expected loss as:
The first term is the conditional bias squared, the second term is the conditional variance, and the last term is the conditional irreducible error.
Furthermore, the expected loss is (the calculations are omitted for economy of space; readers are encouraged to compute them by hand):
The first term is the average bias squared, the second term is the average variance, and the last term is the irreducible error.
For a given size of the training set, we can then define:
- the bias as the distance of the average prediction function from the true regression function
- the variance as the variation of the data-dependent (noisy) prediction function
and... of course, we should aim for a good balance between them!

Goodness of fit
If we consider the same square loss function, a training data set , with , and a test dataset , with , and a prediction function trained on the training data, we can define the in-sample goodness of fit as:
where
- SSR = (sum of squared residuals)
- SST = , with (total sum of squares)
Note that, with the same approach, we can define the out-of-sample , using and from the testing dataset.
As an obvious intuition, measures prediction power (higher is therefore better!).
Approximation and Sampling Error
If is not flexible enough to approximate well (i.e. the hypothesis class does not contain the ideal predictor ), we can incur an approximation error, that measures the limitation of the current model family itself in comparison with the ideal model and it is defined as:
If the empirical loss minimization problem, i.e. , has a solution, we call that solution (note that in this case is random, since it depends on ).
If exists, it is random because it depends on the random training sample. In this case, its sampling error (estimation error) is the difference between its true loss and the population loss available in (it results from minimizing the empirical loss instead of the expected loss; it typically decreases with increasing sample size ), i.e.:
Optimization error
Consider a numerical solution of , with:
- measurable and a -valued random element independent of .
We can then define the direct optimization error as:
Note that the first term is the empirical loss achieved by the numerical solution, the second is the lowest empirical loss achievable by any function in . Note also that this error is different from the sampling error:
- the optimization error measures failure to solve the training problem
- the sampling error measures the gap between training performance and test data performance
Decomposition of the Generalization Error
Once we have defined the previous metrics, we can define the following:
For a given and an increasing (i.e. the set of models we can choose from is increasing):
- approximation error is decreasing
- sampling error tends to increase
- there's a tradeoff between approximation error and sampling error (similar to the bias-variance tradeoff)
- for small data should be simple (and for big data, it can be complex!)
Training Error and Test Error
Once we have defined , we condition on it and so it is no longer random.
If we consider training and test data sets, we can define:
If is significantly smaller than , the model is most likely overfitted.
We can then calculate the sample variance:
which, if , follows the Law of Large Numbers and converges, as , to .
Furthermore, by the Central Limit Theorem, we have:
Therefore, we have the following property:
That is a confidence interval with confidence level , where is the -quantile of the standard normal distribution.
We now come back to the overfitting problem: what can we do about it? Firstly, we can reduce the complexity of the hypothesis class, i.e. by considering simpler models. Alternatively, we can regularize the empirical loss minimization problem by penalizing the complexity of potential prediction functions , which, practically speaking, means:
for hyperparameters (strength of the Lasso penalty), (strength of the ridge penalty).
We call LASSO and ridge regression; together, they compose the elastic net.
In general, LASSO penalizes the absolute value of the coefficients and can set them to zero, and it is defined as:
Therefore, LASSO performs feature selection, i.e. limits the features with lower explanatory power.
On the other hand, ridge regression defines a penalty equal to:
It discourages large coefficients but does not make them exactly zero; it smoothly shrinks them around zero.
Ridge regression can be useful when features are correlated. In this situation, LASSO may select one feature and discard another similar one, while ridge tends to distribute the coefficients across them.
When both penalties are present the method is called elastic net, since it combines LASSO's ability to set coefficients to zero and ridge regression's shrinkage of correlated features.