Phase 9: Attention Mechanism
Module 23: Attention Mechanism
What You Will Learn
In this module, you will learn:
- Why Attention is Needed
- Attention Score
- Alignment
- Context Vector
- Additive Attention (Bahdanau)
- Dot Product Attention
- Scaled Dot Product Attention
- Attention Visualization
- Best Practices
What is Attention?
The Attention Mechanism allows a neural network to focus on the most relevant parts of the input sequence while generating each output token.
Instead of compressing the entire input into a single context vector (as in basic Seq2Seq models), attention lets the decoder dynamically access different parts of the encoder output.
Why Do We Need Attention?
Consider translating the following sentence:
1The little boy is playing football in the park.
When generating the word:
1football
the model should focus on:
1playing football
instead of:
1The little boy
A fixed context vector cannot capture all important information for long sentences.
Attention solves this limitation.
Seq2Seq Without Attention
1Input Sentence 2 3↓ 4 5Encoder 6 7↓ 8 9One Context Vector 10 11↓ 12 13Decoder 14 15↓ 16 17Output Sentence
Problem:
- Long sentences lose information.
- Decoder has access to only one compressed vector.
Seq2Seq With Attention
1Encoder Outputs 2 3h₁ h₂ h₃ h₄ h₅ 4 5 \ | | | / 6 7 Attention 8 9 │ 10 11 ▼ 12 13Context Vector 14 15 │ 16 17 ▼ 18 19Decoder
The decoder can attend to different encoder outputs at every decoding step.
Key Components of Attention
Attention consists of four main components:
- Query (Q)
- Key (K)
- Value (V)
- Attention Scores
Query, Key, and Value
Suppose we have the sentence:
1I love deep learning
Each word is converted into three vectors:
1Word 2 3↓ 4 5Query 6 7↓ 8 9Key 10 11↓ 12 13Value
Meaning:
| Vector | Purpose |
|---|---|
| Query | What information am I looking for? |
| Key | What information do I contain? |
| Value | Information passed to the next layer |
Attention Score
The Attention Score measures how relevant one word is to another.
Example:
1Sentence 2 3I love deep learning
Attention Scores
| Word | Score |
|---|---|
| I | 0.10 |
| love | 0.15 |
| deep | 0.35 |
| learning | 0.40 |
Higher score means greater importance.
Dot Product Attention
The simplest attention score is computed using a dot product.
Equation
1Score = Q · Kᵀ
PyTorch Example
1import torch 2 3Q = torch.randn(1, 64) 4K = torch.randn(1, 64) 5 6score = torch.matmul(Q, K.T) 7 8print(score)
Softmax
Scores are converted into probabilities.
1import torch 2import torch.nn.functional as F 3 4scores = torch.tensor([ 5 2.5, 6 1.2, 7 0.8 8]) 9 10weights = F.softmax( 11 scores, 12 dim=0 13) 14 15print(weights)
Example Output
1tensor([0.69, 0.20, 0.11])
The weights sum to 1.
Context Vector
The Context Vector is a weighted sum of Value vectors.
Equation
1Context = Attention Weights × Values
PyTorch
1import torch 2 3weights = torch.tensor([ 4 [0.2,0.3,0.5] 5]) 6 7V = torch.randn( 8 3, 9 4 10) 11 12context = torch.matmul( 13 weights, 14 V 15) 16 17print(context.shape)
Output
1torch.Size([1,4])
Alignment
Alignment determines which encoder words are most important for generating the current decoder word.
Example
English
1I love AI
French
1J'aime IA
Alignment Matrix
| Decoder | I | love | AI |
|---|---|---|---|
| J' | 0.80 | 0.15 | 0.05 |
| aime | 0.05 | 0.90 | 0.05 |
| IA | 0.02 | 0.08 | 0.90 |
Additive Attention (Bahdanau)
Introduced in 2014.
Uses a small neural network to compute attention scores.
Equation
1Score = Vᵀ tanh(W₁Q + W₂K)
Advantages
- Better for smaller hidden sizes
- Learns nonlinear relationships
Simple Implementation
1import torch 2import torch.nn as nn 3 4class AdditiveAttention(nn.Module): 5 6 def __init__( 7 self, 8 hidden_size 9 ): 10 super().__init__() 11 12 self.Wq = nn.Linear( 13 hidden_size, 14 hidden_size 15 ) 16 17 self.Wk = nn.Linear( 18 hidden_size, 19 hidden_size 20 ) 21 22 self.V = nn.Linear( 23 hidden_size, 24 1 25 ) 26 27 def forward( 28 self, 29 query, 30 keys 31 ): 32 33 scores = self.V( 34 torch.tanh( 35 self.Wq(query) 36 + 37 self.Wk(keys) 38 ) 39 ) 40 41 weights = torch.softmax( 42 scores, 43 dim=1 44 ) 45 46 context = torch.sum( 47 weights * keys, 48 dim=1 49 ) 50 51 return context, weights
Testing Additive Attention
1query = torch.randn( 2 4, 3 1, 4 64 5) 6 7keys = torch.randn( 8 4, 9 10, 10 64 11) 12 13attention = AdditiveAttention(64) 14 15context, weights = attention( 16 query, 17 keys 18) 19 20print(context.shape) 21print(weights.shape)
Output
1torch.Size([4,64]) 2 3torch.Size([4,10,1])
Dot Product Attention
Equation
1Score = QKᵀ
Implementation
1import torch 2import torch.nn.functional as F 3 4Q = torch.randn( 5 2, 6 5, 7 64 8) 9 10K = torch.randn( 11 2, 12 5, 13 64 14) 15 16V = torch.randn( 17 2, 18 5, 19 64 20) 21 22scores = torch.matmul( 23 Q, 24 K.transpose(1,2) 25) 26 27weights = F.softmax( 28 scores, 29 dim=-1 30) 31 32context = torch.matmul( 33 weights, 34 V 35) 36 37print(context.shape)
Output
1torch.Size([2,5,64])
Scaled Dot Product Attention
Large hidden dimensions produce large dot products.
Transformer solves this using scaling.
Equation
1Attention(Q,K,V) 2 3= 4 5Softmax 6 7( 8 9QKᵀ 10 11────── 12 13√dk 14 15) 16 17V
Why Scaling?
Suppose
1Hidden Size = 512
Without scaling,
scores become very large.
Softmax becomes extremely peaked.
Training becomes unstable.
Scaling fixes this.
Implementing Scaled Dot Product Attention
1import math 2import torch 3import torch.nn.functional as F 4 5def scaled_dot_product_attention( 6 Q, 7 K, 8 V 9): 10 11 dk = Q.size(-1) 12 13 scores = torch.matmul( 14 Q, 15 K.transpose(-2,-1) 16 ) 17 18 scores = scores / math.sqrt(dk) 19 20 weights = F.softmax( 21 scores, 22 dim=-1 23 ) 24 25 output = torch.matmul( 26 weights, 27 V 28 ) 29 30 return output, weights
Testing
1Q = torch.randn( 2 2, 3 8, 4 64 5) 6 7K = torch.randn( 8 2, 9 8, 10 64 11) 12 13V = torch.randn( 14 2, 15 8, 16 64 17) 18 19context, weights = scaled_dot_product_attention( 20 Q, 21 K, 22 V 23) 24 25print(context.shape) 26print(weights.shape)
Output
1torch.Size([2,8,64]) 2 3torch.Size([2,8,8])
Attention Visualization
Attention weights can be visualized as a heatmap.
1import matplotlib.pyplot as plt 2 3weights = weights[0].detach().numpy() 4 5plt.imshow( 6 weights, 7 cmap="Blues" 8) 9 10plt.colorbar() 11 12plt.xlabel("Keys") 13 14plt.ylabel("Queries") 15 16plt.title("Attention Map") 17 18plt.show()
Example
1 Keys 2 3 W1 W2 W3 W4 4 5Q1 ██ ██ ░░ ░░ 6 7Q2 ░░ ███ ██ ░ 8 9Q3 ░░ ░░ ██ ███ 10 11Queries
Darker colors indicate higher attention.
Building an Attention Layer
1import torch 2import torch.nn as nn 3import math 4 5class ScaledDotAttention(nn.Module): 6 7 def forward( 8 self, 9 Q, 10 K, 11 V 12 ): 13 14 dk = K.size(-1) 15 16 scores = torch.matmul( 17 Q, 18 K.transpose(-2,-1) 19 ) 20 21 scores = scores / math.sqrt(dk) 22 23 weights = torch.softmax( 24 scores, 25 dim=-1 26 ) 27 28 context = torch.matmul( 29 weights, 30 V 31 ) 32 33 return context
Applications of Attention
- Machine Translation
- Question Answering
- Chatbots
- Text Summarization
- Image Captioning
- Speech Recognition
- Vision Transformers (ViT)
- Large Language Models (LLMs)
Attention vs Seq2Seq
| Basic Seq2Seq | Attention |
|---|---|
| One Context Vector | Dynamic Context |
| Struggles with long sentences | Handles long sequences better |
| Lower translation quality | Higher translation quality |
| Fixed memory | Learns where to focus |
Best Practices
- Use Scaled Dot Product Attention for modern deep learning models.
- Visualize attention maps to understand model behavior.
- Apply masking when processing padded sequences.
- Normalize scores using Softmax.
- Use Multi-Head Attention (covered in the next module) for richer representations.
- Combine attention with positional information when processing sequences in parallel.
Module Summary
In this module, you learned:
- ✅ Why the Attention Mechanism is essential for sequence modeling.
- ✅ How Attention Scores measure the relevance between sequence elements.
- ✅ The concept of Alignment between encoder and decoder tokens.
- ✅ How to compute a Context Vector using attention weights.
- ✅ The differences between Additive Attention, Dot Product Attention, and Scaled Dot Product Attention.
- ✅ How to implement each attention mechanism in PyTorch.
- ✅ How to visualize attention weights using a heatmap.
- ✅ Why attention forms the foundation of modern Transformers and Large Language Models (LLMs).
In the next module, you'll learn Multi-Head Attention, where multiple attention mechanisms operate in parallel to capture different relationships within the same input sequence, enabling the powerful Transformer architecture introduced in Attention Is All You Need.