title: Transformer Decoding Strategies Mathematics meta_title: Transformer Decoding Strategies | Complete Mathematical Formulas meta_description: Learn the complete mathematics of Transformer decoding strategies, including Greedy Search, Beam Search, Top-k Sampling, Top-p Sampling, Temperature Sampling, and token generation formulas. meta_keywords: transformer decoding, greedy search, beam search, top-k sampling, top-p sampling, nucleus sampling, temperature sampling, decoding strategies, transformer mathematics, llm inference, token generation
Module 18 — Decoding Strategies
After a Transformer model predicts the probability distribution over the vocabulary, a decoding strategy determines how the next token is selected.
Different decoding algorithms balance accuracy, diversity, and creativity during text generation.
Modern Large Language Models (LLMs) such as GPT, LLaMA, Gemma, Mistral, Qwen, and DeepSeek use these decoding strategies during inference.
Topics
- Greedy Search
- Beam Search
- Top-k Sampling
- Top-p Sampling
- Temperature Sampling
1. Token Probability
The model predicts a probability distribution over the vocabulary.
Formula
where
- = Next token
- = Output projection scores
2. Greedy Search
Greedy Search selects the token with the highest probability.
Formula
Advantages
- Fast
- Deterministic
- Low computational cost
Disadvantages
- Less diverse
- Can produce repetitive text
3. Beam Search
Beam Search keeps the top B candidate sequences instead of only one.
Sequence Score
The decoder selects
where
- = Beam Width
- = Sequence Length
Advantages
- Higher quality output
- Better global optimization
Disadvantages
- Higher computation
- Less diverse than sampling
4. Top-k Sampling
Instead of considering every vocabulary token, Top-k Sampling only samples from the k most probable tokens.
Candidate Set
Sampling
where
- = Number of candidate tokens
Advantages
- More diverse than Greedy Search
- Prevents sampling unlikely tokens
5. Top-p (Nucleus) Sampling
Top-p Sampling dynamically selects the smallest set of tokens whose cumulative probability exceeds a threshold.
Formula
where
- = Probability Threshold
- = Number of selected tokens
Sampling is performed only within this subset.
Advantages
- Adaptive vocabulary size
- Better text diversity
- Widely used in modern LLMs
6. Temperature Sampling
Temperature controls the randomness of the probability distribution.
Formula
where
- = Logit
- = Temperature
Special cases
If
Normal Softmax
If
Sharper distribution (more deterministic)
If
Flatter distribution (more random)
7. Random Sampling
A token is sampled according to its probability.
Formula
Unlike Greedy Search, lower-probability tokens can still be selected.
8. Autoregressive Generation
The next token depends on all previous generated tokens.
Formula
The generated sequence probability is
9. Complete Decoding Pipeline
The Transformer inference process is
10. Decoding Algorithms
Greedy
Beam Search
Top-k
Top-p
Temperature
11. Applications
These decoding strategies are used in
- GPT
- GPT-2
- GPT-3
- GPT-4
- ChatGPT
- LLaMA
- Gemma
- Qwen
- Mistral
- DeepSeek
- Phi
- Falcon
- BLOOM
- T5
- BART
- Machine Translation
- Text Summarization
- Code Generation
- Conversational AI
- Large Language Models (LLMs)
Summary
| Concept | Formula |
|---|---|
| Token Probability | |
| Greedy Search |