The Mathematical Intuition Behind Deep Learning

From the Dot Product to Multivariate Calculus and the Jacobian, with Python

Alex Punnen
© All Rights Reserved


Contents


Notation

Symbols are defined in each chapter as they are introduced; a few conventions hold throughout:

  • Superscript = layer, subscript = component. \(a^2\) is the activation of layer 2; \(p_i\) is the \(i\)-th element of vector \(p\).
  • Column vectors. A layer computes \(z^l = W^l a^{l-1} + b^l\), and gradients are column vectors, so an error is pushed back through a layer as \((W^l)^T \delta^l\).
  • \(C\) is the cost (loss) and \(\eta\) is the learning rate.
  • Chapter 7’s NumPy code uses the row-vector / batch convention (X @ W); the note at the start of that chapter explains the transposition.