Dive into Deep Learning
Slides
Courses
GitHub
Discuss
PDF
PyTorch
JAX
TensorFlow
MXNet
Notebooks
PyTorch
JAX
TensorFlow
MXNet
Advanced
15
Generative Adversarial Networks
index.html
Preface
Installation
Notation
Introduction
Basics
1
Preliminaries
1.1
Data Manipulation
1.2
Data Preprocessing
1.3
Linear Algebra
1.4
Calculus
1.5
Automatic Differentiation
1.6
Probability and Statistics
1.7
Documentation
2
Linear Regression in Neural Networks
2.1
Linear Regression
2.2
Object-Oriented Design for Implementation
2.3
Synthetic Regression Data
2.4
Linear Regression Implementation from Scratch
2.5
Concise Implementation of Linear Regression
2.6
Generalization
2.7
Weight Decay
3
Linear Classification in Neural Networks
3.1
Softmax Regression
3.2
The Image Classification Dataset
3.3
The Base Classification Model
3.4
Softmax Regression Implementation from Scratch
3.5
Concise Implementation of Softmax Regression
3.6
Generalization in Classification
3.7
Environment and Distribution Shift
4
Multilayer Perceptron
4.1
Multilayer Perceptrons
4.2
Implementation of Multilayer Perceptrons
4.3
Forward Propagation, Backward Propagation, and Computational Graphs
4.4
Numerical Stability and Initialization
4.5
Generalization in Deep Learning
4.6
Dropout
4.7
Predicting House Prices on Kaggle
5
Computation
5.1
Modules and Model Construction
5.2
Parameters, State, and Memory
5.3
Initialization
5.4
Custom Layers and Functions
5.5
Numerics: Dtypes and Mixed Precision
5.6
Saving, Loading, and Pretrained Weights
5.7
GPUs, Devices, and Memory
5.8
Reproducibility and Inspection
6
Convolutional Neural Networks
6.1
From Fully Connected Layers to Convolutions
6.2
Convolutions for Images
6.3
Padding and Stride
6.4
Multiple Input and Multiple Output Channels
6.5
Pooling
6.6
Convolutional Neural Networks (LeNet)
7
Modern Convnets
7.1
The ImageNet Moment: AlexNet
7.2
Blocks, Bottlenecks, and Branches: VGG, NiN, GoogLeNet
7.3
Normalization Layers
7.4
Residual Networks: ResNet, ResNeXt, and DenseNet
7.5
Efficient ConvNets: Depthwise Separability, Mobile Architectures, and Re-parameterization
7.6
Training Recipes Matter
7.7
ConvNeXt: A ConvNet for the 2020s
7.8
Design Spaces and the Big Picture
8
Sequence Models
8.1
Working with Sequences
8.2
From Text to Tokens
8.3
Language Models
8.4
Recurrent Neural Networks
8.5
Implementing RNN Language Models
8.6
Backpropagation Through Time
8.7
Decoding and Generation
Advanced
9
Optimization Algorithms
9.1
Landscapes
9.2
Gradient Descent
9.3
Stochastic Gradient Descent
9.4
Minibatches
9.5
Momentum
9.6
Adam
9.7
AdamW
9.8
Schedules
9.9
Muon
9.10
Batch Size
9.11
Scaling Up
9.12
Practice
10
Attention
10.1
Queries, Keys, and Values
10.2
Attention Scoring and Masking
10.3
Multi-Head and Cross-Attention
10.4
Positional Information
10.5
The Cost of Attention
10.6
What Attention Computes
11
Transformers
11.1
The Transformer Block
11.2
A GPT from Scratch
11.3
Generation and the KV Cache
11.4
Encoders, Decoders, and Cross-Attention
11.5
Vision Transformer
11.6
Mixture of Experts
11.7
Scaling Laws and the Modern Recipe
12
State Space Models
12.1
Gated Recurrence
12.2
Linear Recurrence and State Space Models
12.3
Selective State Space Models
12.4
The Matrix State: From Linear Attention to Mamba-2
12.5
DeltaNet: Memory That Edits
12.6
Learning at Test Time
12.7
Hybrid Architectures
13
Computational Performance
13.1
The Performance Model
13.2
Hardware
13.3
Compute Graphs and Compilation
13.4
Memory and Precision
13.5
Multi-GPU from First Principles
13.6
Multi-GPU in Practice
13.7
Case Study: Making a Transformer Fast
14
Reinforcement Learning
14.1
Markov Decision Process (MDP)
14.2
Value Iteration
14.3
Q-Learning
15
Generative Adversarial Networks
15.1
Generative Adversarial Networks
15.2
Deep Convolutional Generative Adversarial Networks
16
Diffusion Models
Language Models
17
Natural Language Processing: Pretraining
17.1
Encoder-Decoder Models for Sequence Transduction
17.2
Word Embedding (word2vec)
17.3
Approximate Training
17.4
The Dataset for Pretraining Word Embeddings
17.5
Pretraining word2vec
17.6
Word Embedding with Global Vectors (GloVe)
17.7
Subword Embedding
17.8
Word Similarity and Analogy
17.9
Bidirectional Encoder Representations from Transformers (BERT)
17.10
The Dataset for Pretraining BERT
17.11
Pretraining BERT
18
Natural Language Processing: Applications
18.1
Sentiment Analysis and the Dataset
18.2
Sentiment Analysis: Using Recurrent Neural Networks
18.3
Sentiment Analysis: Using Convolutional Neural Networks
18.4
Natural Language Inference and the Dataset
18.5
Natural Language Inference: Using Attention
18.6
Fine-Tuning BERT for Sequence-Level and Token-Level Applications
18.7
Natural Language Inference: Fine-Tuning BERT
Image Models
19
Computer Vision
19.1
Image Augmentation
19.2
Fine-Tuning
19.3
Object Detection and Bounding Boxes
19.4
Anchor Boxes
19.5
Multiscale Object Detection
19.6
The Object Detection Dataset
19.7
Single Shot Multibox Detection
19.8
Region-based CNNs (R-CNNs)
19.9
Semantic Segmentation and the Dataset
19.10
Transposed Convolution
19.11
Fully Convolutional Networks
19.12
Neural Style Transfer
19.13
Image Classification (CIFAR-10) on Kaggle
19.14
Dog Breed Identification (ImageNet Dogs) on Kaggle
Attic
20
Gaussian Processes
20.1
Introduction to Gaussian Processes
20.2
Gaussian Process Priors
20.3
Gaussian Process Inference
21
Hyperparameter Optimization
21.1
What Is Hyperparameter Optimization?
21.2
Hyperparameter Optimization API
21.3
Asynchronous Random Search
21.4
Multi-Fidelity Hyperparameter Optimization
21.5
Asynchronous Successive Halving
22
Recommender Systems
22.1
Overview of Recommender Systems
22.2
The MovieLens Dataset
22.3
Matrix Factorization
22.4
AutoRec: Rating Prediction with Autoencoders
22.5
Personalized Ranking for Recommender Systems
22.6
Neural Collaborative Filtering for Personalized Ranking
22.7
Sequence-Aware Recommender Systems
22.8
Feature-Rich Recommender Systems
22.9
Factorization Machines
22.10
Deep Factorization Machines
Mathematics for Deep Learning
23
Linear Algebra
23.1
Geometry and Linear Algebraic Operations
23.2
Eigendecompositions
23.3
Singular Value Decomposition and Low-Rank Approximation
24
Calculus and Automatic Differentiation
24.1
Single Variable Calculus
24.2
Multivariable Calculus
24.3
Matrix Calculus and Automatic Differentiation
24.4
Integral Calculus
25
Optimization
25.1
Gradient-Based Optimization
25.2
Stochastic and Adaptive Methods
25.3
Convex Sets and Convex Functions
25.4
Constrained Optimization and Duality
25.5
Numerical Stability and Conditioning
26
Probability and Statistical Learning
26.1
Random Variables
26.2
Distributions
26.3
Maximum Likelihood
26.4
Bayesian Computation
26.5
Statistics
26.6
Concentration and Generalization
26.7
Naive Bayes
27
Information Theory and Divergences
27.1
Entropy, Cross-Entropy, and KL Divergence
27.2
Divergences and Distances Between Distributions
27.3
Mutual Information and Representation Learning
28
Dynamics: Differential Equations and Generative Flows
28.1
Ordinary Differential Equations and Numerical Solvers
28.2
Stochastic Differential Equations
28.3
The Fokker–Planck Equation and Probability Flow
28.4
Score Matching, Diffusion, and Flow Matching
Tools for Deep Learning
29
Tools for Deep Learning
29.1
Notebooks
29.2
Colab and Kaggle
29.3
Cloud Computing
29.4
Hardware
29.5
Ecosystem
29.6
Model Training
29.7
Model Serving
29.8
Developers Guide
29.9
Utility Functions and Classes
29.10
The
d2l
API Document
References
Advanced
15
Generative Adversarial Networks
15
Generative Adversarial Networks
14.3
Q-Learning
15.1
Generative Adversarial Networks