In June 2017 a team at Google Brain (Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin) released Attention Is All You Need on arXiv. The title is the thesis: you do not need recurrence, you do not need convolutions, you only need self-attention plus the right scaffolding.
The paper introduced four ideas that have since become the standard vocabulary of modern AI:
- Scaled dot-product attention. The 1/√d_k normalization that keeps softmax behaving well at high dimensions.
- Multi-head attention. Run h attention operations in parallel on lower-dimensional subspaces, then concatenate. Each head learns different relational patterns.
- Positional encoding. Inject position with sinusoids so that order is preserved without changing the architecture's permutation symmetry.
- Pre-LN style residual blocks. Self-attention sub-layer + feed-forward sub-layer, each wrapped in residual connections and layer norm — the unit cell of every modern Transformer.
Receipts
The original encoder-decoder Transformer hit 28.4 BLEU on WMT14 EN-DE and 41.8 on WMT14 EN-FR — state-of-the-art at the time, with a fraction of the training time of comparable models. Base model: 6 encoder + 6 decoder layers, d_model=512, h=8, d_ff=2048, ~65M parameters. Big model: ~213M parameters. Compared to LLaMA 3.3 (70B, 80 layers, d_model=8192) those numbers look quaint — but the unit cell is identical.