🚀 Most Important Transformer Formulas
Organized from embeddings → attention → FFN → normalization → output → training.
1. Token Embedding
Each token ID is converted into a vector:
Where:
- = vocabulary size
- = embedding dimension
Example:
token ID: 25
↓
Embedding table
↓
[0.12, -0.42, 0.73, ...]
2. Positional Encoding
Original sinusoidal positional encoding:
Where:
- = token position
- = embedding dimension index
- = model dimension
3. Query, Key, Value
The most important transformation in self-attention:
Where:
X
│
├── WQ → Query
├── WK → Key
└── WV → Value
4. Attention Score
For individual tokens:
5. Scaled Dot-Product Attention ⭐⭐⭐
Why divide by ?
Without scaling, dot products become large → softmax becomes extremely sharp → gradients suffer.
6. Softmax ⭐⭐⭐
Example:
Result is a probability distribution: .
7. Attention With Mask (Causal)
Prevents looking at future tokens.
8. Multi-Head Attention ⭐⭐⭐
For each head :
9. Head Dimension
Example: , → .
10. Feed-Forward Network ⭐⭐⭐
Original Transformer:
(Modern models usually use GELU or SwiGLU.)
11. GELU
Approximation:
12. SiLU / Swish
13. SwiGLU ⭐⭐⭐
- = gate projection
- = up projection
- = down projection
14. Layer Normalization ⭐⭐⭐
15. RMSNorm ⭐⭐⭐
(Used in most modern LLMs.)
16. Residual Connection ⭐⭐⭐
17. Pre-Norm Transformer
18. RoPE ⭐⭐⭐
Applied to and :
19. Linear Projection to Vocabulary
20. Next-Token Probability
21. Cross-Entropy Loss ⭐⭐⭐
If the correct token is :
22. Gradient Descent
23. Adam Optimizer ⭐⭐⭐
24. AdamW ⭐⭐⭐
25. Dropout
26. Temperature Sampling
- → more deterministic
- → more random
27. Parameter Count
Linear layer:
Attention (approx):
FFN (standard):
One Transformer layer (approx):
28. Attention Complexity
Attention matrix memory:
29. KV Cache ⭐⭐⭐
During autoregressive generation:
For the new token:
🧠 The 10 Formulas You Should Memorize
-
QKV
🔥 Complete Modern Decoder Flow
Each layer:
Inside attention:
Finally:
This is the core mathematical pipeline behind modern Transformer / LLM architectures.
Just assign the whole block above (as a template literal or string) to `category.description` and it will render beautifully with math, headings, lists, and code blocks.