Module 6 — Embedding Mathematics
Embeddings convert discrete tokens into dense numerical vectors that neural networks can process. In Transformer models, embeddings provide semantic meaning, positional information, and sentence relationships before the attention mechanism begins.
Topics
- Vocabulary
- Token Embedding
- Position Embedding
- Segment Embedding
- Input Embedding
- Learned Embedding
1. Vocabulary
A vocabulary is the set of all unique tokens known by a model.
Vocabulary Size
where
- = Vocabulary
- = Number of unique tokens
2. Token Embedding
Each token is mapped to a dense vector using an embedding matrix.
Embedding Matrix
where
- = Vocabulary size
- = Embedding dimension
Token Lookup
For token index
or
where
- = Embedding matrix
3. Position Embedding
Position embeddings encode the order of tokens.
Position Vector
where
- = Maximum sequence length
Sinusoidal Position Encoding
Even Dimensions
Odd Dimensions
4. Segment Embedding
Segment embeddings distinguish different sentences.
Formula
Example
Sentence A
Sentence B
This embedding is mainly used in BERT.
5. Input Embedding
The final embedding is obtained by combining multiple embeddings.
Formula
If segment embeddings are not used (e.g., GPT)
6. Learned Embedding
Instead of fixed positional encodings, modern models learn embeddings during training.
Formula
where
Embedding Matrix Dimensions
Token Embedding
Position Embedding
Segment Embedding
Input Embedding
Embedding Lookup Example
Suppose
- Vocabulary Size = 50,000
- Sequence Length = 128
- Embedding Dimension = 768
Then
Token Embedding
Position Embedding
Input Embedding
Why Embeddings Matter in Transformers
Embeddings are the first stage of every Transformer architecture.
They are used in
- Transformer
- BERT
- RoBERTa
- ALBERT
- GPT
- GPT-2
- GPT-3
- GPT-4
- LLaMA
- Gemma
- Mistral
- Qwen
- Phi
- DeepSeek
- T5
- Vision Transformer (ViT)
Summary
| Concept | Formula |
|---|---|
| Vocabulary Size | $ |
| Embedding Matrix | $E\in\mathbb{R}^{ |
| Token Embedding | |
| One-Hot Embedding |