Module 3 — Calculus for Deep Learning
Calculus is one of the most important mathematical foundations of Deep Learning. Every neural network, including Transformer models, relies on derivatives and gradients for optimization through backpropagation.
Topics
- Limits
- Derivatives
- Partial Derivatives
- Chain Rule
- Gradient
- Jacobian
- Hessian
- Taylor Expansion
1. Limits
A limit describes the value that a function approaches as the input approaches a particular value.
Formula
Example
2. Derivatives
The derivative measures the rate of change of a function.
Formula
Definition
Common Derivatives
3. Partial Derivatives
A partial derivative measures the change in a multivariable function with respect to one variable while keeping the others constant.
Formula
Example
If
then
4. Chain Rule
The chain rule computes derivatives of composite functions.
Formula
Multiple Variables
This rule is the mathematical foundation of backpropagation.
5. Gradient
The gradient is a vector containing all partial derivatives of a function.
Formula
Example
6. Jacobian Matrix
The Jacobian represents the first-order partial derivatives of a vector-valued function.
Formula
Expanded Form
7. Hessian Matrix
The Hessian is a matrix of second-order partial derivatives.
Formula
Expanded Form
The Hessian is widely used in optimization and curvature analysis.
8. Taylor Expansion
Taylor expansion approximates a function around a given point.
First-Order Taylor Expansion
Second-Order Taylor Expansion
General Taylor Series
Why Calculus is Important in Transformers
Calculus enables Transformer models to learn by optimizing millions or even billions of parameters.
It is used in:
- Gradient Descent
- Backpropagation
- Adam Optimizer
- Loss Function Optimization
- Neural Network Training
- Attention Weight Updates
- Large Language Models (LLMs)
Summary
| Concept | Formula |
|---|---|
| Limit | |
| Derivative |