Two failure modes, same root cause
The chain rule multiplies derivatives. If those derivatives are mostly less than 1, the product shrinks toward zero exponentially with depth — vanishing gradients. If they are mostly greater than 1, the product grows toward infinity — exploding gradients. Either way, the deep layers stop learning.
Vanishing gradients dominated the late 1990s. Sigmoid and tanh activations saturate at the tails (derivative ≈ 0), so a 20-layer net with sigmoids barely trained the early layers at all. ReLU, careful initialization, and batch normalization fixed most of that — and the rest was fixed by residual connections.
Exploding gradients are usually fixable with one knob
Exploding gradients show up most often in recurrent networks (LSTMs, GRUs) and in transformers with poor learning-rate scheduling. The fix is gradient clipping: cap the global norm of the gradient at some threshold before stepping. torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0) is the one-liner.
Vanishing gradients need architectural answers
You cannot clip your way out of vanishing gradients — there is nothing to clip, the gradients are zero. The fixes are structural: ReLU/GELU activations, He/Xavier initialization, batch or layer normalization, and residual connections. Modern transformers use all four together by default, which is why they train at depths that would have been impossible in 2014.