Module 19 — BERT Mathematics
BERT (Bidirectional Encoder Representations from Transformers) is an encoder-only Transformer model introduced by Google in 2018. Unlike autoregressive models, BERT learns contextual representations by attending to both the left and right context of every token.
BERT is pre-trained using two objectives:
- Masked Language Modeling (MLM)
- Next Sentence Prediction (NSP)
Topics
- Masked Language Modeling (MLM)
- Next Sentence Prediction (NSP)
- Encoder Stack
1. BERT Input Embedding
The input embedding is the sum of three embeddings.
Formula
where
- = Token Embedding
- = Positional Embedding
2. Encoder Stack
BERT uses only Transformer encoder layers.
Formula
where
- = Number of encoder layers
Examples
- BERT Base → 12 encoder layers
- BERT Large → 24 encoder layers
3. Self-Attention
Each encoder layer computes bidirectional self-attention.
Formula
where
4. Multi-Head Attention
Multiple attention heads are concatenated.
Formula
5. Feed Forward Network
Each encoder block contains a position-wise feed forward network.
Formula
6. Residual Connection
Each sublayer uses residual learning.
Formula
7. Encoder Output
The final contextual representations are
Formula
where
8. Masked Language Modeling (MLM)
A subset of input tokens is randomly masked.
The model predicts only the masked tokens.
MLM Probability
where
- = Masked Token
MLM Loss
9. Next Sentence Prediction (NSP)
BERT predicts whether sentence B follows sentence A.
Probability
NSP Loss
where
- = Ground Truth Label
- = Predicted Probability
10. Total Pretraining Loss
BERT combines MLM and NSP losses.
Formula
This is the primary optimization objective during BERT pretraining.
11. Hidden Representation of [CLS]
The first token is the classification token.
Formula
The vector
is used for classification tasks.
12. Vocabulary Prediction
The MLM head projects hidden states into the vocabulary.
Formula
Probability
13. Matrix Dimensions
Suppose
- Sequence Length =
- Hidden Size =
Input
Encoder Output
Vocabulary Logits
14. Applications
BERT is widely used for
- Text Classification
- Sentiment Analysis
- Named Entity Recognition (NER)
- Question Answering
- Semantic Search
- Document Classification
- Information Retrieval
- Text Similarity
- Feature Extraction
Popular BERT-based models include
- BERT Base
- BERT Large
- RoBERTa
- ALBERT
- DistilBERT
- DeBERTa
- ELECTRA
- SciBERT
- BioBERT
Summary
| Concept | Formula |
|---|---|
| Input Embedding |