26  Probability and Statistical Learning

Models are probabilistic statements about data. This chapter develops the continuous probability the rest of deep learning relies on—densities, expectations, and how they transform under a map—catalogues the distributions whose negative log-likelihoods are exactly our loss functions, derives maximum-likelihood and MAP estimation (and the priors that become regularizers), turns the resulting posterior integrals into computations with Monte Carlo and variational approximations, builds the statistics needed to tell a real improvement from noise, proves the concentration inequalities that make finite samples trustworthy—following them to uniform convergence, Rademacher complexity, and double descent—and caps it all with naive Bayes: a working classifier, fit by counting and then audited with the chapter’s own tools.

Resources and Further Reading

A short, opinionated shelf for going deeper into the probability and statistical learning that underpins these chapters—random variables and distributions, maximum-likelihood and MAP estimation, Bayesian inference, estimators, and hypothesis testing. We favor free and official sources.

Books

Courses and video lectures

Tutorials and notes