01⚡Why Transformers
0/10 lessonsParallel sequence modeling and the bottleneck it broke
What problem the Transformer actually solved, why attention beats recurrence, and where the 2017 design fits in modern AI's lineage.
Lesson list (10)
- 01Sequence Modeling Before Transformers: RNNs and LSTMs~22 min · history, rnn, lstm, background
- 02The Parallelization Problem: Why GPUs Hated RNNs~18 min · parallelism, gpu, rnn
- 03Long-Range Dependencies and the Vanishing Gradient~18 min · long-range, vanishing-gradient, rnn
- 04The Attention Insight: Direct Pairwise Access~20 min · attention, intuition, core-idea
- 05The 2017 Paper: 'Attention Is All You Need'~16 min · history, paper, vaswani-2017
- 06Impact: From Translation to GPT, BERT, and the Modern LLM Stack~16 min · history, gpt, bert, llm
- 07Beyond Text: Vision, Audio, Biology~14 min · vision, audio, biology, multimodal
- 08The Scaling Hypothesis and What It Got Right~16 min · scaling, chinchilla, kaplan
- 09Encoder, Decoder, Encoder-Decoder — Three Shapes, Three Roles~14 min · encoder, decoder, architecture-shapes
- 10Roadmap: What Tracks 2-8 Will Build~8 min · roadmap, course-overview