Phase 8: Encoder-Decoder Models
Module 22: Sequence-to-Sequence (Seq2Seq)
What You Will Learn
In this module, you will learn:
- What is Sequence-to-Sequence (Seq2Seq)?
- Why Seq2Seq Models are Needed
- Encoder
- Decoder
- Teacher Forcing
- Greedy Decoding
- Beam Search
- Machine Translation
- Building a Seq2Seq Model in PyTorch
- Best Practices
What is Sequence-to-Sequence (Seq2Seq)?
A Sequence-to-Sequence (Seq2Seq) model is a neural network architecture that converts one sequence into another sequence.
Unlike traditional classification models that produce a single output, Seq2Seq models generate an entire sequence of outputs.
Examples:
| Input | Output |
|---|---|
| English Sentence | French Sentence |
| Speech | Text |
| Question | Answer |
| Paragraph | Summary |
| Python Code | Java Code |
Why Do We Need Seq2Seq?
Suppose we want to translate:
1English 2 3I love deep learning.
into
1French 2 3J'aime l'apprentissage profond.
The output contains multiple words.
A normal neural network predicts only one output.
Seq2Seq predicts an entire sequence.
Seq2Seq Architecture
1Input Sentence 2 3 │ 4 5 ▼ 6 7+------------------+ 8| Encoder | 9+------------------+ 10 11 │ 12 13 Context Vector 14 15 │ 16 17 ▼ 18 19+------------------+ 20| Decoder | 21+------------------+ 22 23 │ 24 25 ▼ 26 27Output Sentence
Real Example
1Input 2 3I love AI
↓
Encoder
↓
Context Vector
↓
Decoder
↓
1J'aime l'IA
Components of Seq2Seq
A Seq2Seq model consists of two main parts:
- Encoder
- Decoder
Encoder
The Encoder reads the input sequence one token at a time and converts it into a fixed-size representation called the Context Vector.
Example
1Input 2 3I 4 5↓ 6 7love 8 9↓ 10 11AI 12 13↓ 14 15Encoder 16 17↓ 18 19Context Vector
The Encoder is usually built using:
- RNN
- LSTM
- GRU
- Transformer Encoder
Context Vector
The Context Vector summarizes the entire input sentence.
1Input Sentence 2 3↓ 4 5Encoder 6 7↓ 8 9[0.34, -0.12, 0.91, ...] 10 11↓ 12 13Decoder
In early Seq2Seq models, the Context Vector had a fixed size, making it difficult to handle very long sentences.
This limitation motivated the development of the Attention Mechanism.
Decoder
The Decoder generates the output sequence one token at a time.
Example
1Context Vector 2 3↓ 4 5Decoder 6 7↓ 8 9Bonjour 10 11↓ 12 13le 14 15↓ 16 17monde 18 19↓ 20 21<END>
The Decoder predicts one word, then uses that prediction to generate the next word.
Seq2Seq Workflow
1English Sentence 2 3 │ 4 5Tokenization 6 7 │ 8 9Word IDs 10 11 │ 12 13Embedding Layer 14 15 │ 16 17Encoder 18 19 │ 20 21Context Vector 22 23 │ 24 25Decoder 26 27 │ 28 29Output Words 30 31 │ 32 33Translated Sentence
Teacher Forcing
During training, the Decoder can receive:
- The previous predicted word
- The actual (ground-truth) word
Using the ground-truth word is called Teacher Forcing.
Without Teacher Forcing
1<START> 2 3↓ 4 5I 6 7↓ 8 9love 10 11↓ 12 13deep 14 15↓ 16 17Wrong Prediction 18 19↓ 20 21Error Propagates
With Teacher Forcing
1Ground Truth Word 2 3↓ 4 5Decoder 6 7↓ 8 9Next Prediction
Training becomes faster and more stable.
Teacher Forcing Example
Suppose the correct output is:
1Bonjour le monde
Training
| Step | Decoder Input |
|---|---|
| 1 | <START> |
| 2 | Bonjour |
| 3 | le |
| 4 | monde |
Instead of feeding the predicted word, we feed the correct word.
Greedy Decoding
During inference, the true target sentence is unavailable.
The Decoder selects the most probable word at every step.
Example
1Step 1 2 3Hello 4 50.90 6 7Hi 8 90.08 10 11Hey 12 130.02 14 15↓ 16 17Choose 18 19Hello
Greedy Decoding is simple and fast but may not produce the best overall sentence.
Beam Search
Beam Search improves upon Greedy Decoding by keeping multiple candidate sequences instead of only one.
Example
Beam Width = 3
1Candidate 1 2 3Hello World 4 5Score = 0.92
1Candidate 2 2 3Hi World 4 5Score = 0.89
1Candidate 3 2 3Greetings World 4 5Score = 0.82
The sequence with the highest overall probability is selected.
Greedy vs Beam Search
| Greedy Decoding | Beam Search |
|---|---|
| Keeps one sequence | Keeps multiple sequences |
| Fast | Slower |
| Less accurate | More accurate |
| Lower memory | Higher memory |
Building an Encoder
1import torch 2import torch.nn as nn 3 4class Encoder(nn.Module): 5 6 def __init__( 7 self, 8 vocab_size, 9 embed_dim, 10 hidden_dim 11 ): 12 super().__init__() 13 14 self.embedding = nn.Embedding( 15 vocab_size, 16 embed_dim 17 ) 18 19 self.gru = nn.GRU( 20 embed_dim, 21 hidden_dim, 22 batch_first=True 23 ) 24 25 def forward(self, x): 26 27 embedded = self.embedding(x) 28 29 outputs, hidden = self.gru(embedded) 30 31 return hidden
Building a Decoder
1class Decoder(nn.Module): 2 3 def __init__( 4 self, 5 vocab_size, 6 embed_dim, 7 hidden_dim 8 ): 9 super().__init__() 10 11 self.embedding = nn.Embedding( 12 vocab_size, 13 embed_dim 14 ) 15 16 self.gru = nn.GRU( 17 embed_dim, 18 hidden_dim, 19 batch_first=True 20 ) 21 22 self.fc = nn.Linear( 23 hidden_dim, 24 vocab_size 25 ) 26 27 def forward( 28 self, 29 x, 30 hidden 31 ): 32 33 x = self.embedding(x) 34 35 output, hidden = self.gru( 36 x, 37 hidden 38 ) 39 40 prediction = self.fc( 41 output 42 ) 43 44 return prediction, hidden
Building the Seq2Seq Model
1class Seq2Seq(nn.Module): 2 3 def __init__( 4 self, 5 encoder, 6 decoder 7 ): 8 super().__init__() 9 10 self.encoder = encoder 11 12 self.decoder = decoder 13 14 def forward( 15 self, 16 src, 17 tgt 18 ): 19 20 hidden = self.encoder(src) 21 22 outputs = [] 23 24 decoder_input = tgt[:, 0].unsqueeze(1) 25 26 for t in range( 27 1, 28 tgt.size(1) 29 ): 30 31 output, hidden = self.decoder( 32 decoder_input, 33 hidden 34 ) 35 36 outputs.append(output) 37 38 decoder_input = tgt[:, t].unsqueeze(1) 39 40 return torch.cat(outputs, dim=1)
Creating the Model
1SRC_VOCAB = 1000 2TGT_VOCAB = 1000 3 4encoder = Encoder( 5 SRC_VOCAB, 6 128, 7 256 8) 9 10decoder = Decoder( 11 TGT_VOCAB, 12 128, 13 256 14) 15 16model = Seq2Seq( 17 encoder, 18 decoder 19)
Dummy Input
1src = torch.randint( 2 0, 3 1000, 4 (4,10) 5) 6 7tgt = torch.randint( 8 0, 9 1000, 10 (4,12) 11) 12 13outputs = model( 14 src, 15 tgt 16) 17 18print(outputs.shape)
Output
1torch.Size([4,11,1000])
Training the Model
1criterion = nn.CrossEntropyLoss() 2 3optimizer = torch.optim.Adam( 4 model.parameters(), 5 lr=0.001 6) 7 8optimizer.zero_grad() 9 10outputs = model( 11 src, 12 tgt 13) 14 15loss = criterion( 16 17 outputs.reshape( 18 -1, 19 TGT_VOCAB 20 ), 21 22 tgt[:,1:].reshape(-1) 23) 24 25loss.backward() 26 27optimizer.step() 28 29print(loss.item())
Machine Translation Project
Step 1: Sample Dataset
1english = [ 2 3 "i love ai", 4 5 "good morning", 6 7 "how are you" 8] 9 10french = [ 11 12 "j aime ia", 13 14 "bonjour", 15 16 "comment allez vous" 17]
Step 2: Build Vocabulary
1from collections import Counter 2 3counter = Counter() 4 5for sentence in english: 6 7 counter.update(sentence.split()) 8 9vocab = { 10 11 "<PAD>":0, 12 13 "<SOS>":1, 14 15 "<EOS>":2, 16 17 "<UNK>":3 18} 19 20for word in counter: 21 22 vocab[word] = len(vocab) 23 24print(vocab)
Step 3: Encode Sentences
1def encode(sentence): 2 3 ids = [1] 4 5 ids += [ 6 7 vocab.get(word,3) 8 9 for word in sentence.split() 10 ] 11 12 ids.append(2) 13 14 return ids 15 16encoded = [ 17 18 encode(s) 19 20 for s in english 21] 22 23print(encoded)
Step 4: Inference
1model.eval() 2 3with torch.no_grad(): 4 5 outputs = model( 6 src, 7 tgt 8 ) 9 10print(outputs.argmax(-1))
Applications of Seq2Seq
- Machine Translation
- Text Summarization
- Chatbots
- Speech Recognition
- Image Captioning
- Question Answering
- Code Generation
- Grammar Correction
Seq2Seq vs Traditional Models
| Traditional Model | Seq2Seq |
|---|---|
| Single Output | Sequence Output |
| Classification | Sequence Generation |
| Fixed Output Size | Variable Output Size |
| Simple Architecture | Encoder-Decoder |
Limitations of Basic Seq2Seq
- Fixed-size context vector limits performance on long sentences.
- Long sequences may lose important information.
- Training can be slow.
- Decoder errors can accumulate during inference.
- Performance degrades on complex translation tasks.
These limitations led to the development of Attention Mechanisms, where the decoder can focus on different parts of the input sequence instead of relying on a single context vector.
Best Practices
- Use Embeddings instead of one-hot vectors.
- Add
<SOS>,<EOS>, and<PAD>tokens to your vocabulary. - Use Teacher Forcing during training for faster convergence.
- Prefer Beam Search over Greedy Decoding for higher-quality translations.
- Apply gradient clipping to stabilize training.
- Use padding masks when batching variable-length sequences.
- For modern NLP tasks, prefer Attention-based Seq2Seq or Transformer models over basic RNN/LSTM/GRU Seq2Seq architectures.
Module Summary
In this module, you learned:
- ✅ What Sequence-to-Sequence (Seq2Seq) models are and why they are used for sequence generation tasks.
- ✅ How the Encoder converts an input sequence into a context representation.
- ✅ How the Decoder generates the target sequence one token at a time.
- ✅ The purpose of Teacher Forcing during training.
- ✅ The differences between Greedy Decoding and Beam Search during inference.
- ✅ How to build a complete Encoder-Decoder Seq2Seq model in PyTorch using GRUs.
- ✅ How Seq2Seq models are applied to Machine Translation and other sequence generation tasks.
In the next module, you'll learn the Attention Mechanism, which allows the decoder to focus on the most relevant parts of the input sequence at each decoding step. Attention overcomes the fixed context vector limitation and forms the foundation of modern Transformers and Large Language Models (LLMs).