Module 14 — Encoder–Decoder Architecture
The Encoder–Decoder Architecture is the original Transformer architecture introduced in "Attention Is All You Need" (2017).
The encoder converts the input sequence into contextual representations, while the decoder generates the output sequence one token at a time using Cross Attention.
This architecture is widely used in
- Machine Translation
- Text Summarization
- Speech Recognition
- Image Captioning
- Question Answering
- Sequence-to-Sequence (Seq2Seq) Tasks
Topics
- Cross Attention
- Sequence-to-Sequence
- Translation
1. Encoder
The encoder converts the input sequence into contextual embeddings.
Input
Encoder Output
where
The encoder output is also called the Memory.
2. Decoder Input
The decoder receives the shifted target sequence.
Formula
3. Masked Self Attention
The decoder first performs causal self-attention.
Formula
where
- = Causal Mask
4. Cross Attention
Cross Attention connects the decoder with the encoder.
The decoder Query attends to the encoder Keys and Values.
Query
Key
Value
Cross Attention Formula
This is the most important equation in the Encoder–Decoder architecture.
5. Decoder Output
The decoder combines
- Masked Self Attention
- Cross Attention
- Feed Forward Network
to generate contextual representations.
Formula
6. Vocabulary Projection
The decoder output is projected into vocabulary space.
Formula
where
7. Probability Distribution
The logits are converted into probabilities.
Formula
8. Predicted Token
The next generated token is
9. Sequence-to-Sequence Learning
The model predicts the probability of the entire output sequence given the input sequence.
Formula
where
- = Input sequence
- = Output sequence
- = Target sequence length
10. Translation Objective
For machine translation
Input
Output
The Transformer learns
or equivalently
11. Complete Encoder–Decoder Pipeline
Encoder
Decoder
Cross Attention
Output Projection
Final Prediction
12. Matrix Dimensions
Suppose
- Input Length =
- Output Length =
- Model Dimension =
Encoder Output
Decoder Input
Query
Key
Value
Cross Attention Output
Vocabulary Logits
13. Applications
The Encoder–Decoder architecture is used in
- Original Transformer
- T5
- BART
- MarianMT
- PEGASUS
- NLLB
- Speech-to-Text
- Machine Translation
- Text Summarization
- OCR
- Image Captioning
- Multimodal Transformers
Summary
| Concept | Formula |
|---|---|
| Encoder Output | |
| Decoder Input |