The Hugging Face Transformers documentation organizes hundreds of models into a few fundamental Transformer architecture families. Almost every model belongs to one of these families, based on the original 2017 Transformer paper (Encoder, Decoder, or Encoder–Decoder). Newer vision, audio, and multimodal models extend these same ideas rather than inventing completely different architectures. (Hugging Face)
You can understand the ecosystem with the following architecture tree.
1 Transformer (2017) 2 │ 3 ┌──────────────────────┼──────────────────────┐ 4 │ │ │ 5 Encoder Only Decoder Only Encoder-Decoder 6 │ │ │ 7 Understanding Generation Seq2Seq Tasks 8 │ │ │ 9 BERT, RoBERTa, GPT, LLaMA, T5, BART, 10 DeBERTa, ViT, Mistral, Gemma, mBART, 11 Swin, CLIP Qwen, Falcon Pegasus
Complete Transformer Family
1Transformers 2│ 3├── 1. Encoder-Only Models 4│ │ 5│ ├── NLP 6│ │ ├── BERT 7│ │ ├── RoBERTa 8│ │ ├── ALBERT 9│ │ ├── ELECTRA 10│ │ ├── DeBERTa 11│ │ ├── DistilBERT 12│ │ ├── Longformer 13│ │ └── BigBird 14│ │ 15│ ├── Vision 16│ │ ├── ViT 17│ │ ├── Swin 18│ │ ├── DINOv2 19│ │ ├── BEiT 20│ │ └── ConvNeXt-V2 21│ │ 22│ ├── Audio 23│ │ ├── Wav2Vec2 24│ │ ├── HuBERT 25│ │ └── SEW 26│ │ 27│ └── Documents 28│ ├── LayoutLM 29│ ├── LayoutLMv2 30│ └── LayoutLMv3 31│ 32├── 2. Decoder-Only Models 33│ │ 34│ ├── GPT-2 35│ ├── GPT-Neo 36│ ├── GPT-J 37│ ├── LLaMA 38│ ├── Mistral 39│ ├── Mixtral 40│ ├── Gemma 41│ ├── Falcon 42│ ├── BLOOM 43│ ├── Phi 44│ ├── Qwen 45│ ├── DeepSeek 46│ └── CodeLlama 47│ 48├── 3. Encoder-Decoder Models 49│ │ 50│ ├── T5 51│ ├── mT5 52│ ├── BART 53│ ├── Pegasus 54│ ├── MarianMT 55│ ├── VisionEncoderDecoder 56│ ├── SpeechEncoderDecoder 57│ └── TrOCR 58│ 59├── 4. Vision-Language Models 60│ │ 61│ ├── BLIP 62│ ├── BLIP-2 63│ ├── CLIP 64│ ├── SigLIP 65│ ├── LLaVA 66│ ├── IDEFICS 67│ ├── Florence-2 68│ ├── Qwen2-VL 69│ ├── PaliGemma 70│ └── Kosmos-2 71│ 72├── 5. OCR & Document AI 73│ │ 74│ ├── TrOCR 75│ ├── Donut 76│ ├── Nougat 77│ ├── PaddleOCR-VL 78│ ├── LayoutLM 79│ └── UDOP 80│ 81├── 6. Object Detection 82│ │ 83│ ├── DETR 84│ ├── Deformable DETR 85│ ├── RT-DETR 86│ ├── YOLOS 87│ ├── OWLv2 88│ └── Grounding DINO 89│ 90├── 7. Image Segmentation 91│ │ 92│ ├── SegFormer 93│ ├── MaskFormer 94│ ├── Mask2Former 95│ └── SAM 96│ 97├── 8. Speech Models 98│ │ 99│ ├── Whisper 100│ ├── SpeechT5 101│ ├── Wav2Vec2 102│ ├── HuBERT 103│ └── Encodec 104│ 105├── 9. Video Transformers 106│ │ 107│ ├── VideoMAE 108│ ├── TimeSformer 109│ └── ViViT 110│ 111└── 10. Diffusion Models 112 │ 113 ├── Stable Diffusion 114 ├── FLUX 115 ├── PixArt 116 └── Kandinsky
Architecture Family Relationship
1Transformer 2│ 3├── Text 4│ ├── Encoder 5│ ├── Decoder 6│ └── Encoder-Decoder 7│ 8├── Vision 9│ ├── ViT 10│ ├── Swin 11│ ├── DINOv2 12│ └── BEiT 13│ 14├── Vision + Text 15│ ├── CLIP 16│ ├── BLIP 17│ ├── LLaVA 18│ ├── Florence 19│ └── Qwen2-VL 20│ 21├── OCR 22│ ├── TrOCR 23│ ├── Donut 24│ ├── PaddleOCR-VL 25│ └── LayoutLM 26│ 27├── Audio 28│ ├── Whisper 29│ ├── SpeechT5 30│ ├── Wav2Vec2 31│ └── HuBERT 32│ 33├── Detection 34│ ├── DETR 35│ ├── RT-DETR 36│ ├── Grounding DINO 37│ └── OWLv2 38│ 39└── Video 40 ├── TimeSformer 41 ├── VideoMAE 42 └── ViViT
Original Transformer Design
1 Input Tokens 2 │ 3 Token Embedding 4 │ 5 Positional Embedding 6 │ 7 Multi-Head Attention 8 │ 9 Feed Forward Network 10 │ 11────────────────────────────────────────── 12 13Encoder Only 14Input → Encoder Stack → Output Features 15 16Examples: 17BERT 18RoBERTa 19ViT 20LayoutLM 21 22────────────────────────────────────────── 23 24Decoder Only 25 26Input 27 │ 28Masked Self-Attention 29 │ 30Decoder Stack 31 │ 32Next Token Prediction 33 34Examples: 35GPT 36LLaMA 37Gemma 38Qwen 39Mistral 40 41────────────────────────────────────────── 42 43Encoder + Decoder 44 45Input 46 │ 47Encoder 48 │ 49Cross Attention 50 │ 51Decoder 52 │ 53Generated Output 54 55Examples: 56T5 57BART 58TrOCR 59VisionEncoderDecoder 60Whisper
Hugging Face Model Categories
The Hugging Face Transformers library groups architectures broadly into:
- Text models (BERT, LLaMA, T5, Gemma, Qwen)
- Vision models (ViT, Swin, DINOv2, BEiT)
- Audio models (Whisper, Wav2Vec2, HuBERT)
- Multimodal models (CLIP, BLIP, LLaVA, Florence-2, Qwen2-VL)
- Document/OCR models (LayoutLM, TrOCR, Donut, PaddleOCR-VL)
- Object detection models (DETR, RT-DETR, Grounding DINO)
- Video models (TimeSformer, VideoMAE, ViViT)
- Diffusion/image generation models (Stable Diffusion, FLUX, PixArt) (Hugging Face)
This hierarchy is the most useful mental model for understanding the entire Hugging Face Transformers ecosystem: every architecture is an encoder-only, decoder-only, encoder–decoder, or multimodal extension of the original Transformer.
Collection Description
Transformer Architecture Collection is a comprehensive learning resource that explores the complete ecosystem of Transformer-based AI models. Starting from the original Transformer architecture, this collection explains Encoder-only, Decoder-only, and Encoder-Decoder models before diving into modern applications across natural language processing, computer vision, speech recognition, OCR, document AI, object detection, video understanding, and multimodal AI.
Learn how popular architectures such as BERT, RoBERTa, GPT, LLaMA, Gemma, Mistral, Qwen, T5, BART, ViT, Swin Transformer, DETR, Whisper, CLIP, BLIP, LLaVA, Florence-2, Qwen2-VL, LayoutLM, TrOCR, PaddleOCR-VL, and many more are designed, how they process different types of data, and which tasks they are best suited for.
Each guide includes architecture diagrams, model family relationships, practical use cases, Hugging Face implementations, and explanations of how encoder, decoder, cross-attention, self-attention, embeddings, and multimodal pipelines work. Whether you're a beginner learning Transformers or a researcher exploring state-of-the-art AI architectures, this collection provides a structured roadmap to understand the modern Transformer ecosystem.
Perfect for: AI Engineers, Machine Learning Engineers, Deep Learning Practitioners, NLP Engineers, Computer Vision Developers, Data Scientists, Researchers, and students learning modern AI.
Transformers Learning Roadmap (Beginner to Advanced)
This roadmap is designed for developers who want to master Transformer architecture, Large Language Models (LLMs), Vision Transformers (ViT), Multimodal AI, and Generative AI. It starts from the mathematical foundations and progresses to building and training modern transformer models.
Phase 1: Prerequisites
Module 1: Mathematics for Transformers
Topics
- Linear Algebra
- Scalars
- Vectors
- Matrices
- Matrix Multiplication
- Dot Product
- Tensor Basics
- Eigenvalues
- Probability
- Statistics
- Softmax
- Logarithm
- Exponential Functions
- Entropy
- Cross Entropy
- KL Divergence
- Calculus
- Derivatives
- Chain Rule
- Gradient
Practice
- Matrix Operations using NumPy
- Implement Softmax from Scratch
- Cross Entropy Calculation
Module 2: Deep Learning Foundations
Topics
- Neural Networks
- Activation Functions
- Forward Propagation
- Backpropagation
- Loss Functions
- Gradient Descent
- Optimizers
- Regularization
- Batch Normalization
- Dropout
- Embeddings
Practice
- Build Feed Forward Network
- Train Text Classifier
Phase 2: NLP Fundamentals
Module 3: Natural Language Processing
Topics
- What is NLP
- Text Processing
- Tokenization
- Normalization
- Stemming
- Lemmatization
- Stop Words
- Vocabulary
- N-Grams
- Bag of Words
- TF-IDF
- Word Embeddings
- Word2Vec
- GloVe
- FastText
Practice
- Spam Classifier
- Sentiment Analysis
Module 4: Sequence Models
Topics
- Sequential Data
- RNN
- LSTM
- GRU
- Encoder-Decoder
- Seq2Seq
- Teacher Forcing
- Attention Motivation
- Problems with RNNs
Practice
- Machine Translation using LSTM
Phase 3: Attention Mechanism
Module 5: Attention
Topics
- Why Attention
- Query
- Key
- Value
- Attention Score
- Dot Product Attention
- Additive Attention
- Scaled Dot Product Attention
- Self Attention
- Cross Attention
- Multi Head Attention
Practice
- Implement Self Attention from Scratch
Module 6: Positional Encoding
Topics
- Why Position Information
- Sinusoidal Encoding
- Learned Positional Embeddings
- Rotary Position Embedding (RoPE)
- ALiBi
- Relative Position Encoding
Practice
- Visualize Positional Encoding
Phase 4: Transformer Architecture
Module 7: Encoder
Topics
- Encoder Block
- Multi Head Attention
- Feed Forward Network
- Residual Connection
- Layer Normalization
- Dropout
- Stacking Encoders
Practice
- Build Transformer Encoder
Module 8: Decoder
Topics
- Decoder Block
- Masked Self Attention
- Cross Attention
- Feed Forward Layer
- Residual Connections
- LayerNorm
- Output Projection
Practice
- Build Transformer Decoder
Module 9: Complete Transformer
Topics
- Encoder-Decoder Architecture
- Input Embeddings
- Output Embeddings
- Attention Masks
- Padding Mask
- Causal Mask
- Training Pipeline
- Inference Pipeline
- Greedy Search
- Beam Search
Practice
- Machine Translation Transformer
Phase 5: Hugging Face Ecosystem
Module 10: Transformers Library
Topics
- Installation
- AutoModel
- AutoTokenizer
- AutoProcessor
- Pipelines
- Config
- Model Loading
- Saving Models
- Model Hub
Practice
- Text Classification
- Translation
- Summarization
Module 11: Tokenizers
Topics
- BPE
- WordPiece
- SentencePiece
- Unigram
- Vocabulary
- Special Tokens
- Padding
- Truncation
- Attention Mask
Practice
- Train Custom Tokenizer
Phase 6: BERT Family
Module 12: Encoder Models
Topics
- BERT
- RoBERTa
- ALBERT
- DistilBERT
- ELECTRA
- DeBERTa
- Masked Language Modeling
- Next Sentence Prediction
Practice
- Question Answering
- Text Classification
Phase 7: GPT Family
Module 13: Decoder Models
Topics
- GPT
- GPT-2
- GPT-3
- GPT-4 Concepts
- GPT-Neo
- GPT-J
- GPT-NeoX
- LLaMA
- Mistral
- Falcon
- Qwen
- Gemma
Practice
- Text Generation
- Chatbot
Phase 8: Encoder-Decoder Models
Module 14: Seq2Seq Transformers
Topics
- T5
- FLAN-T5
- BART
- PEGASUS
- MarianMT
- mT5
Practice
- Translation
- Summarization
Phase 9: Vision Transformers
Module 15: Vision Transformer (ViT)
Topics
- Image Patch
- Patch Embedding
- CLS Token
- Position Embedding
- Vision Encoder
- Image Classification
- DeiT
- Swin Transformer
- ConvNeXt
Practice
- CIFAR-10 Classification
Phase 10: Multimodal Transformers
Module 16: Vision Language Models
Topics
- CLIP
- BLIP
- BLIP-2
- Flamingo
- LLaVA
- Qwen2-VL
- InternVL
- Kosmos
- Florence
- PaddleOCR-VL
Practice
- Image Captioning
- Visual Question Answering
Module 17: Speech Transformers
Topics
- Whisper
- Wav2Vec2
- SpeechT5
- Audio Spectrogram Transformer
- Speech Recognition
- Speech Translation
Practice
- Speech-to-Text
Phase 11: Large Language Models
Module 18: LLM Architecture
Topics
- Decoder-only Models
- Scaling Laws
- Context Window
- KV Cache
- Flash Attention
- Grouped Query Attention
- Multi Query Attention
- Mixture of Experts (MoE)
- Sparse Attention
- Sliding Window Attention
Practice
- Compare LLaMA and Mistral
Module 19: Prompt Engineering
Topics
- Zero-shot
- One-shot
- Few-shot
- Chain of Thought
- Self Consistency
- Tree of Thoughts
- ReAct
- Structured Prompting
- Prompt Templates
- System Prompts
Practice
- Build AI Assistant
Phase 12: Fine-Tuning
Module 20: Fine-Tuning Transformers
Topics
- Full Fine-Tuning
- Transfer Learning
- PEFT
- LoRA
- QLoRA
- Adapters
- Prefix Tuning
- Prompt Tuning
- BitFit
- IA3
Practice
- Fine-tune BERT
- Fine-tune LLaMA using LoRA
Module 21: Training Pipeline
Topics
- Dataset
- Data Collator
- Trainer
- SFTTrainer
- TrainingArguments
- Accelerate
- DeepSpeed
- FSDP
- Mixed Precision
- Gradient Checkpointing
Practice
- Train Custom LLM
Phase 13: Retrieval-Augmented Generation (RAG)
Module 22: Embeddings and Retrieval
Topics
- Embeddings
- Sentence Transformers
- FAISS
- ChromaDB
- Pinecone
- Milvus
- Hybrid Search
- Reranking
- Dense Retrieval
- Sparse Retrieval
Practice
- Document Search Engine
Module 23: RAG Systems
Topics
- Chunking
- Metadata
- Indexing
- Retrieval
- Prompt Construction
- Context Injection
- Evaluation
- Hallucination
- Grounding
Practice
- PDF Chatbot
Phase 14: Reinforcement Learning for LLMs
Module 24: Alignment
Topics
- Supervised Fine-Tuning (SFT)
- Reward Models
- RLHF
- DPO
- PPO
- ORPO
- GRPO
- Constitutional AI
- AI Safety
Practice
- Preference Optimization
Phase 15: Optimization
Module 25: Efficient Transformers
Topics
- Quantization
- GPTQ
- AWQ
- GGUF
- ONNX
- TensorRT
- vLLM
- Tensor Parallelism
- Pipeline Parallelism
- Speculative Decoding
Practice
- Deploy Quantized LLM
Phase 16: Deployment
Module 26: Model Serving
Topics
- Hugging Face Inference
- FastAPI
- TorchServe
- Triton Inference Server
- vLLM
- Ollama
- Docker
- Kubernetes
- API Deployment
- Monitoring
Practice
- Deploy Chat API
Phase 17: Advanced Research
Module 27: Latest Transformer Innovations
Topics
- Longformer
- BigBird
- Performer
- Linformer
- Reformer
- RetNet
- RWKV
- Mamba
- Hyena
- State Space Models (SSMs)
Practice
- Benchmark Efficient Architectures
Phase 18: Capstone Projects
Module 28: Real-World Projects
Beginner
- Sentiment Analysis
- Spam Detection
- News Classification
- Text Summarizer
- Machine Translation
Intermediate
- Question Answering System
- Resume Parser
- AI Chatbot
- Named Entity Recognition
- Semantic Search Engine
Advanced
- Retrieval-Augmented Generation (RAG) Assistant
- Multimodal Image + Text Assistant
- Document Intelligence System
- OCR + Vision-Language Pipeline
- AI Coding Assistant
- Medical LLM
- Legal Document Analyzer
- Multi-Agent AI System
- Fine-Tuned Domain-Specific LLM
- End-to-End Production AI Application
Final Learning Outcomes
By completing this roadmap, you will be able to:
- Understand the mathematical foundations behind Transformers.
- Build attention mechanisms and Transformer blocks from scratch.
- Work with encoder-only, decoder-only, and encoder-decoder architectures.
- Use Hugging Face Transformers effectively.
- Fine-tune pretrained models using PEFT, LoRA, and QLoRA.
- Build RAG pipelines with vector databases.
- Develop multimodal applications using vision-language models.
- Optimize and deploy transformer models for production.
- Read, understand, and implement recent Transformer research papers.
- Build production-ready AI systems powered by modern Transformer architectures.