Linear Algebra
Linear algebra is the mathematical foundation of Transformer models. Every embedding, attention score, projection layer, and feed-forward network is built using vectors and matrices.
10 minOrder #1Free
Probability and Statistics
Learn probability and statistics formulas used in Transformers, including Bayes' theorem, probability distributions, expectation, variance, covariance, and correlation.
30 minOrder #2Free
Calculus Deep Learning
Learn calculus formulas used in deep learning and Transformers, including limits, derivatives, partial derivatives, chain rule, gradients, Jacobian, Hessian, and Taylor expansion.
50 minOrder #3Free
Optimization Algorithms
Learn optimization algorithms used in deep learning and Transformers, including Gradient Descent, SGD, Momentum, RMSProp, Adam, AdamW, and learning rate scheduling formulas.
50 minOrder #4Free
Neural Network Fundamentals
Learn neural network fundamentals with formulas for linear layers, bias, activation functions, Sigmoid, Tanh, ReLU, GELU, Softmax, Layer Normalization, and Dropout used in Transformers.
59 minOrder #5Free
Embedding Mathematics
Learn Transformer embedding mathematics including vocabulary, token embeddings, positional embeddings, segment embeddings, input embeddings, and learned embeddings with complete formulas.
100 minOrder #6Free
Positional Encoding
Learn positional encoding in Transformer models with formulas for sinusoidal encoding, learned positional embeddings, relative position encoding, and Rotary Position Embedding (RoPE).
101 minOrder #7Free
Self-Attention
Learn the mathematics of Self-Attention in Transformers with complete formulas for Query, Key, Value, attention scores, scaled dot-product attention, attention weights, and output computation.
10 minOrder #8Free
Multi-Head Attention
Learn the complete mathematics of Multi-Head Attention in Transformers, including multiple attention heads, head projections, concatenation, output projection, and scaled dot-product attention formulas.
100 minOrder #9Free
Feed Forward Network
Learn the mathematics of the Feed Forward Network (FFN) in Transformer models, including dense layers, hidden layers, GELU activation, output projection, and complete FFN equations.
100 minOrder #10Free
Residual Connections
Learn the mathematics of Residual Connections in Transformer models, including skip connections, Add operation, Layer Normalization, residual learning, and complete formulas.
90 minOrder #11Free
Transformer Encoder
Learn the complete mathematics of the Transformer Encoder, including encoder blocks, self-attention, feed forward networks, residual learning, layer normalization, and encoder stack formulas.
90 minOrder #12Free
Transformer Decoder
Learn the complete mathematics of the Transformer Decoder, including masked self-attention, cross-attention, feed forward networks, residual connections, layer normalization, and output layer formulas.
91 minOrder #13Free
Encoder–Decoder Architect
Learn the complete mathematics of the Transformer Encoder–Decoder architecture, including cross attention, sequence-to-sequence learning, machine translation, encoder-decoder interaction, and output generation formulas.
10 minOrder #14Free
Output Projection
Learn the complete mathematics of Output Projection in Transformer models, including vocabulary projection, logits computation, Softmax probability distribution, and token prediction formulas.
90 minOrder #15Free
Loss Functions
Learn the complete mathematics of Transformer loss functions, including Cross Entropy, Binary Cross Entropy, Label Smoothing, KL Divergence, and training objective formulas.
10 minOrder #16Free
Transformer Training
Learn the complete mathematics of Transformer training, including backpropagation, gradient clipping, learning rate warmup, weight decay, AdamW optimization, and parameter update formulas.
90 minOrder #17Free
Transformer Decoding
Learn the complete mathematics of Transformer decoding strategies, including Greedy Search, Beam Search, Top-k Sampling, Top-p Sampling, Temperature Sampling
70 minOrder #18Free
BERT Mathematics
bert mathematics, bert formulas, masked language modeling, mlm, next sentence prediction, nsp, bert encoder, transformer encoder, bert architecture, transformer mathematics, deep learning mathematics, language models
90 minOrder #19Free
GPT Mathematics
Learn the complete mathematics of GPT, including causal attention, decoder-only Transformer architecture, next token prediction, autoregressive language modeling, and training objective formulas.
10 minOrder #20Free
T5 Mathematics
Learn the complete mathematics of T5 (Text-to-Text Transfer Transformer), including encoder-decoder architecture, span corruption, text-to-text learning, attention mechanisms, and sequence generation formulas.
50 minOrder #21Free
Vision Transformer (ViT)
Learn the complete mathematics of Vision Transformers (ViT), including patch embedding, image tokens, CLS token, positional embeddings, Transformer encoder, and image classification formulas.
50 minOrder #22Free
Advanced Attention
Learn the complete mathematics of advanced Transformer attention mechanisms, including Rotary Position Embedding (RoPE), ALiBi, FlashAttention, Grouped Query Attention (GQA), and Multi Query Attention (MQA).
90 minOrder #23Free
Mixture of Experts (MoE)
Learn the complete mathematics of Mixture of Experts (MoE), including routers, experts, sparse activation, gating networks, Top-k routing, load balancing, and expert output aggregation formulas.
59 minOrder #24Free
Modern LLM Architecture
Learn the complete mathematics of modern Large Language Model (LLM) architectures, including RMSNorm, SwiGLU, KV Cache, Grouped Query Attention (GQA), and Rotary Position Embedding (RoPE).
59 minOrder #25Free
Transformer Complexity
Learn the complete mathematics of Transformer computational complexity, including time complexity, memory complexity, FLOPs, scaling laws, efficient attention algorithms, and long-context optimization formulas.
90 minOrder #26Free
Complete Transformer Pipe
Learn the complete end-to-end Transformer pipeline, from tokenization and embeddings to attention, encoder, decoder, output projection, softmax, and token generation with mathematical formulas.
70 minOrder #27Free
Important Transformer
Organized from embeddings → attention → FFN → normalization → output → training.
10 minOrder #28Free