Module 12 — Transformer Encoder
The Transformer Encoder is the primary building block of encoder-based architectures such as BERT, RoBERTa, ALBERT, DeBERTa, and the encoder part of sequence-to-sequence models like T5.
An encoder consists of multiple identical Encoder Blocks stacked together. Each block contains:
- Multi-Head Self Attention
- Residual Connection
- Layer Normalization
- Feed Forward Network
- Residual Connection
- Layer Normalization
Each encoder layer transforms the input into richer contextual representations while preserving sequence length.
Topics
- Encoder Block
- Self Attention
- Feed Forward
- Residual Learning
- Layer Normalization
1. Encoder Input
The encoder receives the embedded input sequence.
Formula
where
- = Sequence Length
- = Hidden Dimension
2. Multi-Head Self Attention
The first operation inside an encoder block is Multi-Head Self Attention.
Formula
where
3. Residual Connection
The attention output is added to the original input.
Formula
4. Layer Normalization
The residual output is normalized.
Formula
Expanded
5. Feed Forward Network
The normalized output passes through a Feed Forward Network.
Formula
6. Second Residual Connection
The FFN output is added back to its input.
Formula
7. Second Layer Normalization
The second normalization completes the encoder block.
Formula
Expanded
8. Complete Encoder Block
The complete Transformer Encoder Block can be written as
This represents the standard encoder architecture used in the original Transformer paper.
9. Simplified Encoder Formula
A simplified mathematical representation is
This combines the major operations into a single equation.
10. Encoder Stack
A Transformer encoder consists of multiple encoder layers stacked sequentially.
Formula
where
- = Number of encoder layers
Examples
- BERT Base → 12 Encoder Layers
- BERT Large → 24 Encoder Layers
11. Matrix Dimensions
Suppose
- Sequence Length =
- Model Dimension =
Input
Self Attention Output
Feed Forward Output
Encoder Output
12. Why the Transformer Encoder Matters
The Transformer Encoder builds contextual token representations by combining:
- Self-Attention for capturing global relationships
- Feed Forward Networks for non-linear transformations
- Residual Connections for stable training
- Layer Normalization for improved optimization
It is the foundation of many modern AI models, including:
- Transformer
- BERT
- RoBERTa
- ALBERT
- ELECTRA
- DeBERTa
- T5 Encoder
- Vision Transformer (ViT)
- LayoutLM
- BEiT
- CLIP Text Encoder
Summary
| Concept | Formula |
|---|---|
| Input Embedding |