Module 7 — Positional Encoding
Transformers process all tokens simultaneously, unlike RNNs that naturally preserve sequence order. Positional Encoding provides information about the position of each token so that the model understands word order within a sequence.
Topics
- Sinusoidal Encoding
- Learned Positional Encoding
- Relative Position Encoding
- Rotary Position Embedding (RoPE)
1. Why Positional Encoding?
The self-attention mechanism is permutation invariant, meaning it has no inherent understanding of token order.
Without positional information,
I love AI
AI love I
would produce nearly identical attention patterns.
Therefore, positional information is added to token embeddings before entering the Transformer.
Input Representation
where
- = Token Embedding
- = Positional Encoding
2. Sinusoidal Positional Encoding
The original Transformer paper uses sinusoidal functions to encode positions.
Even Dimensions
Odd Dimensions
where
- = Position in the sequence
- = Embedding index
- = Embedding dimension
Complete Positional Encoding Matrix
where
- = Maximum sequence length
Final Input
3. Learned Positional Encoding
Instead of fixed sinusoidal functions, modern Transformers learn positional vectors.
Position Embedding Matrix
The final embedding becomes
Learned positional embeddings are used in
- BERT
- GPT-2
- GPT-3
- GPT-4
- ViT
4. Relative Position Encoding
Instead of encoding absolute positions, Relative Position Encoding represents the distance between tokens.
For two positions
their relative distance is
The attention score becomes
where
- = Relative positional embedding
Relative position encoding is commonly used in
- Transformer-XL
- T5
- DeBERTa
5. Rotary Position Embedding (RoPE)
RoPE encodes position by rotating Query and Key vectors in embedding space.
Instead of adding positional vectors, RoPE applies a rotation matrix.
Rotation Matrix
Query Rotation
Key Rotation
Attention with RoPE
RoPE preserves relative positional relationships naturally.
6. Rotation Angle
The rotation angle depends on token position.
7. Why RoPE Works
RoPE enables the model to
- Preserve relative positions
- Generalize to longer sequences
- Improve extrapolation
- Reduce memory usage
It is widely used in
- LLaMA
- Gemma
- Qwen
- DeepSeek
- Mistral
- Phi
- Yi
8. Comparison
| Method | Position Type |
|---|---|
| Sinusoidal Encoding | Absolute |
| Learned Embedding | Absolute |
| Relative Encoding | Relative |
| RoPE | Relative (Rotation) |
Summary
| Concept | Formula |
|---|---|
| Input Embedding | |
| Sinusoidal (Even) |