27  Information Theory and Divergences

Information theory is the language of losses. This chapter builds entropy, cross-entropy, and the Kullback–Leibler divergence, shows that minimizing cross-entropy is maximum likelihood, and then broadens to the wider family of divergences and distances (f-divergences, optimal transport, integral probability metrics) that define modern generative objectives, and to mutual information and the contrastive objectives of representation learning.

Resources and Further Reading

A short, curated reading list for information theory and divergences as they appear in machine and deep learning: entropy and cross-entropy, the KL and broader \(f\)-divergences, mutual information, optimal transport, and the information bottleneck.

Books

Courses and lecture notes

Tutorials, blogs, and surveys

Foundational papers