Loading topic...
Loading topic...
A curated collection of 10 influential AI research papers covering Transformers, LLMs, NLP, deep learning, optimization, and parameter-efficient fine-tuning.
Explore 10 influential research papers that shaped modern Artificial Intelligence and Machine Learning.
Artificial Intelligence, Machine Learning, Deep Learning, NLP, Transformers, Large Language Models, PyTorch, Python, Generative AI, BERT, GPT, LoRA, CNN, GAN, VAE, LSTM
NIPS-2017-attention-is-all-you-need-Paper.pdf
PDF • 556.1 KB
cvpr2016_deep_residual_learning_kaiminghe.pdf
PDF • 2815.1 KB
BERT_Pre-training_of_Deep_Bidirectional_Transformers_for.pdf
PDF • 757.0 KB
Language_Models_are_Few-Shot_Learners.pdf
PDF • 6609.4 KB
Generative_Adversarial_Nets.pdf
PDF • 527.1 KB
Auto-Encoding_Variational_Bayes.pdf
PDF • 3834.6 KB
ADAM_A_METHOD_FOR_STOCHASTIC_OPTIMIZATION.pdf
PDF • 570.9 KB
NIPS-2012-imagenet-classification-with-deep-convolutional-neural-networks-Paper.pdf
PDF • 1385.6 KB
Long_Short-Term_Memory.pdf
PDF • 237.2 KB
Playing_Atari_with_Deep_Reinforcement_Learning.pdf
PDF • 423.8 KB
3-Figure1-1.png
PNG • 81.6 KB
No additional resources.
No contributors yet.
Transformers changed the direction of modern AI. Starting with “Attention Is All You Need” (2017), Transformer-based architectures have become foundational to NLP, large language models, computer vision, and multimodal AI. Here are 10 important papers to study: Attention Is All You Need — The original Transformer architecture and self-attention mechanism. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding — Introduced powerful bidirectional language representations. Improving Language Understanding by Generative Pre-Training (GPT) — Established the GPT-style decoder-only pretraining approach. Language Models are Unsupervised Multitask Learners (GPT-2) — Demonstrated strong zero-shot and multitask capabilities through scaling. Language Models are Few-Shot Learners (GPT-3) — Showed the remarkable effect of scale and in-context learning. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5) — Unified many NLP tasks under a text-to-text framework. An Image is Worth 16×16 Words (Vision Transformer) — Demonstrated that Transformer architectures can work effectively for image recognition. RoFormer: Enhanced Transformer with Rotary Position Embedding — Introduced RoPE, an important positional representation technique used in modern Transformer models. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity — Advanced sparse Mixture-of-Experts Transformer scaling. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Improved the efficiency of attention computation for large models. 📚 Why read these papers? Together, they provide a progression from the original Transformer architecture to BERT, GPT, T5, Vision Transformers, positional encoding, MoE, and efficient attention. If you're learning LLMs, Generative AI, NLP, or Transformer architecture, this is a strong research-paper reading path. 💬 Which Transformer paper had the biggest impact on your work? #Transformers #AI #MachineLearning #DeepLearning #LLM #GenerativeAI #NLP #ArtificialIntelligence #ResearchPapers #Tech3Space