title: Transformer Decoder Mathematics meta_title: Transformer Decoder | Complete Mathematical Formulas meta_description: Learn the complete mathematics of the Transformer Decoder, including masked self-attention, cross-attention, feed forward networks, residual connections, layer normalization, and output layer formulas. meta_keywords: transformer decoder, decoder block, masked self attention, cross attention, decoder formulas, transformer mathematics, feed forward network, layer normalization, seq2seq decoder, transformer architecture, llm decoder
Module 13 — Transformer Decoder
The Transformer Decoder is responsible for generating output tokens one at a time. Unlike the encoder, the decoder contains an additional Cross-Attention layer that allows it to attend to the encoder output.
Each decoder block consists of:
- Masked Multi-Head Self Attention
- Residual Connection
- Layer Normalization
- Encoder–Decoder Cross Attention
- Residual Connection
- Layer Normalization
- Feed Forward Network
- Residual Connection
- Layer Normalization
The decoder is used in sequence-to-sequence models such as Transformer, T5, BART, and encoder-decoder machine translation systems.
Topics
- Masked Attention
- Cross Attention
- Feed Forward
- Output Layer
1. Decoder Input
The decoder receives the shifted target sequence.
Formula
where
2. Masked Self-Attention
Unlike encoder attention, decoder attention uses a causal mask so that future tokens cannot be seen.
Formula
where
- = Causal Mask
- Future positions are assigned
3. Residual Connection
The masked attention output is added to the input.
Formula
4. Cross Attention
The decoder attends to the encoder output.
Queries come from the decoder.
Keys and Values come from the encoder.
Query
Key
Value
Cross Attention
5. Cross-Attention Residual
The Cross-Attention output is added to the decoder representation.
Formula
6. Feed Forward Network
The normalized representation passes through a Feed Forward Network.
Formula
7. Final Residual Connection
The FFN output is added back.
Formula
8. Decoder Output Layer
The decoder output is projected into vocabulary space.
Linear Projection
where
Vocabulary Probability
Predicted Token
9. Complete Decoder Block
The complete decoder block can be written as
10. Simplified Decoder Formula
A compact mathematical representation is
11. Decoder Stack
A Transformer decoder contains multiple identical decoder layers.
Formula
where
- = Number of decoder layers
12. Matrix Dimensions
Suppose
- Sequence Length =
- Model Dimension =
Decoder Input
Masked Attention Output
Cross Attention Output
Feed Forward Output
Final Decoder Output
Vocabulary Projection
13. Why the Transformer Decoder Matters
The Transformer Decoder is responsible for autoregressive sequence generation.
It is used in
- Original Transformer
- T5
- BART
- PEGASUS
- MarianMT
- Machine Translation
- Text Summarization
- Image Captioning
- Speech Recognition
- Encoder–Decoder Language Models
Decoder-only models such as GPT replace Cross-Attention with additional Masked Self-Attention layers.
Summary
| Concept | Formula |
|---|---|
| Decoder Input |