25  Calculus and Automatic Differentiation

Training a network requires derivatives of a scalar loss with respect to its parameters. This chapter follows two tracks. The first three sections develop the differentiation needed for optimization: scalar derivatives, gradients and chain rules, then Jacobian products and automatic differentiation. The final section begins a second track on integration for continuous probability (Chapter 27), expectations, and the differential equations of Chapter 29. Integration is not a prerequisite for backpropagation.

Resources and Further Reading

The following references range from single-variable calculus reviews to matrix calculus and automatic differentiation.

Books

Courses and video lectures

Tutorials, notes, and visual introductions

Automatic differentiation