The 1/√d_k factor in the attention formula isn't decorative. Without it, large d_k destroys the softmax.
Here's why. Dot products of two random vectors of dimension d_k have variance proportional to d_k. So if Q and K are roughly unit-norm in expectation, Q · K typically has magnitude on the order of √d_k. For d_k = 64, that's about 8; for d_k = 128, about 11. Plug values of magnitude 8 into a softmax and the largest one almost wins outright — softmax saturates, the other entries become near-zero, and gradients through them die.
Dividing by √d_k keeps the variance of the score O(1) regardless of dimension. Softmax then produces a useful distribution that responds smoothly to changes in Q and K, and gradients flow.
This is one of the few places in modern Transformers where a simple constant correction makes the difference between "trains" and "doesn't train." It's also the reason the architecture is called scaled dot-product attention.