Deep Learning — Study Companion

SkillDocs & knowledge

This is a study companion for the Deep Learning textbook by Goodfellow, Bengio and Courville (MIT Press, 2016), which you can read free at deeplearningbook.org. Once added, your AI can guide you through the book's 20 chapters and tell you which of its advice still holds today. It also helps you plan a reading path that fits your goals.

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

After adding it, ask your AI a question about deep learning or name a chapter you want to study, and it will pull up the relevant material and how it fits with current practice. You can also ask it to map out a reading plan before you start.

Then ask your AI: use the Deep Learning skill

What your AI can do with it

  • Find the right chapter of the Deep Learning textbook for a topic you are studying
  • Explain concepts from all 20 chapters as you read
  • Tell you which of the book's recommendations still hold and which were superseded by newer methods such as transformers and AdamW
  • Plan a reading path through the book based on what you want to learn
  • Point you to the free full text at deeplearningbook.org

What this skill tells your AI

The instructions your AI receives, as published by alirezarezvani/claude-skills in engineering/deep-learning-book/skills/deep-learning-book/SKILL.md and read by ahel’s review.

Source book: Deep Learning, Ian Goodfellow, Yoshua Bengio & Aaron Courville (MIT Press, 2016) · 20 chapters, 3 parts · read free at deeplearningbook.org · companion compiled 2026-08-25.

This is a companion, not a copy. The book is copyrighted, and its site states that the HTML-only format exists to discourage copying under the authors' MIT Press contract. Nothing here reproduces its text. Every chapter file is original synthesis — what the chapter establishes, how to use it, where it has aged — plus a link to the official chapter. Read the book at the link; use this to navigate it, keep it current, and turn it into decisions. See references/rights_and_use.md.

How to Use This Skill

  • No argument — load the core frameworks below.
  • A topic — ask about regularization, saddle points, partition function; resolved through the Topic Index, then that chapter file is read before answering.
  • chNN — load that chapter's file.
  • "is this still true?" — the 2016→2026 delta layer, in every chapter file and in references/book_to_2026_delta.md.
  • "where do I start?" — run scripts/reading_path_planner.py.

When asked about something outside these 20 chapters, say so and route to the delta reference rather than improvising the book's position on material published after it.


Core Frameworks & Mental Models

The (T, P, E) frame — ch05

Name the task, the performance measure, and the experience in one sentence before any model code. Most failed projects failed at P: an unstated metric, or a proxy whose relationship to the real objective was never checked.

Every loss is a negative log-likelihood — ch03, ch06

Choose the output distribution, then take its negative log. Gaussian → MSE, Bernoulli → binary cross-entropy, categorical → cross-entropy, Laplace → MAE. "Which loss?" is always the question "which distribution?" in disguise. Modern contrastive and preference objectives sit outside this frame — a real limit of the book, not a gap in your understanding.

KL asymmetry decides your failure mode — ch03, ch19, ch20

D(p‖q) ≠ D(q‖p). Forward KL is mode-covering (blurry averages); reverse KL is mode-seeking (sharp but partial). This single fact predicts VAE blur, GAN mode collapse, and the characteristic over-confidence of mean-field variational posteriors.

Train-error-first triage — ch11, ch05

High training error → capacity or optimization is the bottleneck; more data will not help. Low training error with a large validation gap → data or regularization. This is the highest-value heuristic in the book. scripts/training_diagnostics.py runs it.

Capacity, the gap, and the U-curve's caveat — ch05, ch07

Regularization trades variance for bias. But the classical U-shaped capacity curve is incomplete: past the interpolation threshold, test error can fall again (double descent, 2019–2020, post-dating the book). Practical consequence: when a large model overfits, try more data, more regularization or longer training before shrinking it.

Architecture is a prior, not a trick — ch09, ch10, ch15

Convolution asserts translation equivariance and locality. Recurrence asserts that the past compresses into a state. A distributed representation asserts that factors combine combinatorially. When the assertion is false, the architecture cannot be rescued by tuning — and when it is true, it beats capacity. This is also why Vision Transformers need more data than ConvNets: they discard the prior and buy it back with examples.

Depth's real cost is gradient flow and activation memory — ch06, ch08, ch10

Backprop is the chain rule scheduled well: one forward-pass-equivalent of compute, and memory proportional to stored activations. Depth fails through vanishing/exploding gradients and ill-conditioning, which is why residual connections, normalization and clipping exist.

The partition function organizes Part III — ch16, ch17, ch18, ch19

For undirected models, the likelihood gradient needs samples from the model itself. Four escape routes: sample it (CD/PCD), sidestep it algebraically (pseudolikelihood, score matching), learn around it (NCE), or estimate it for evaluation (AIS). Score matching's descendants are today's diffusion models — which is why Part III repays reading even though its models did not survive.

Diagnose before you redesign — ch04, ch08, ch11

Gradient norm exploding → clip. Norm large but loss flat → ill-conditioning. Norm near zero with high loss → saturation or dead units. NaN → numerics first. Change one thing per experiment.


Chapter Index

#TitleKey content
ch01Introductionrepresentation learning, depth as composition, curse of dimensionality
ch02Linear Algebranorms, SVD, eigendecomposition, conditioning, PCA
ch03Probability & Information Theorydistributions, entropy, KL, cross-entropy
ch04Numerical Computationunder/overflow, conditioning, gradient descent, KKT
ch05Machine Learning Basicscapacity, bias–variance, No Free Lunch, MLE, manifolds
ch06Deep Feedforward Networksoutput/hidden units, universal approximation, backprop
ch07Regularizationnorm penalties, augmentation, early stopping, dropout
ch08OptimizationSGD, momentum, init, Adam, batch norm, saddles
ch09Convolutional Networkssparse interactions, sharing, equivariance, pooling
ch10Sequence ModelingBPTT, vanishing gradients, LSTM/GRU, attention
ch11Practical Methodologymetrics, baselines, the data-vs-capacity rule, debugging
ch12Applicationsscaling, compression, vision, speech, NLP (dated)
ch13Linear Factor ModelsPPCA, factor analysis, ICA, sparse coding
ch14Autoencodersundercomplete, sparse, denoising, contractive
ch15Representation Learningtransfer, distributed codes, disentanglement
ch16Structured Probabilistic Modelsdirected/undirected, energy-based, d-separation
ch17Monte Carlo Methodsimportance sampling, MCMC, Gibbs, mixing
ch18Confronting the Partition FunctionCD/PCD, pseudolikelihood, score matching, NCE, AIS
ch19Approximate InferenceELBO, EM, mean field, amortization
ch20Deep Generative ModelsBoltzmann machines, VAE, GAN, autoregressive

Topic Index

  • Activation functions, ReLU, GELU → ch06
  • Adam, AdamW, adaptive optimizers → ch08, ch07
  • Attention, transformers → ch10, ch12
  • Autoencoders, denoising, sparse → ch14, ch13
  • Backpropagation, autodiff → ch06
  • Batch / layer normalization → ch08
  • Bias–variance, double descent → ch05
  • Convolution, pooling, receptive field → ch09
  • Cross-entropy, KL divergence, entropy → ch03
  • Diffusion, score matching → ch18, ch14, ch20
  • Dropout, weight decay, early stopping → ch07
  • ELBO, variational inference, EM → ch19
  • Energy-based models, graphical models → ch16
  • GANs, VAEs, generative taxonomy → ch20
  • Gradient clipping, exploding/vanishing → ch10, ch08
  • Hyperparameter search → ch11
  • Initialization → ch08
  • LSTM, GRU, BPTT, teacher forcing → ch10
  • Maximum likelihood, MAP → ch05, ch03
  • MCMC, Gibbs, importance sampling → ch17
  • Numerical stability, softmax, log-space → ch04
  • Partition function, CD, PCD, NCE → ch18, ch16
  • PCA, ICA, factor analysis → ch13, ch02
  • Representation learning, transfer, probes → ch15, ch01
  • Saddle points, ill-conditioning → ch08, ch04
  • SVD, eigendecomposition, condition number → ch02
  • Training diagnostics, metric choice → ch11
  • Universal approximation → ch06

Supporting Files

Tools

S=engineering/deep-learning-book/skills/deep-learning-book/scripts
python3 $S/reading_path_planner.py --goal "train a transformer" --background applied --hours-per-week 5
python3 $S/training_diagnostics.py --train-loss 0.02 --val-loss 1.9 --grad-norm 0.4 --epochs 30
python3 $S/capacity_planner.py --params 12000000 --train-examples 50000 --train-error 0.01 --val-error 0.22
python3 $S/model_arithmetic.py --spec-sample

Every tool supports --help, --sample and --output json, uses the standard library only, and returns typed exit codes.


Scope & Limits

This companion covers the 2016 edition's 20 chapters and the delta between them and 2026 practice. It does not cover: reinforcement learning beyond passing mention, LLM training infrastructure, RLHF/DPO alignment, agentic systems, MLOps tooling, or fairness and safety evaluation — none of which the book treats. For production ML engineering use engineering-team/senior-ml-engineer; for LLM cost work use engineering/llm-cost-optimizer.

When a question lands outside the book, say the book does not cover it and cite the delta reference for what replaced its position. A companion that quietly extrapolates is worse than one that names its boundary.

Signals

GitHub stars
27k
Forks
4k
Last commit
Aug 2026
Advanced
Item type
skill
Key
deep-learning-book
Source
github.com/alirezarezvani/claude-skills