The Question Backprop Answers
You did a forward pass. You got a loss. Now you need to know: "For each weight in this network — possibly billions of them — how would the loss change if I nudged that weight a tiny bit?" Those numbers are the gradients. With them, gradient descent can update every weight to reduce loss.
Naively, you'd compute each gradient separately by perturbing one weight at a time. That's extra forward passes for weights. With billions of weights, this would take forever.
Backprop's Trick
Backpropagation computes ALL gradients in essentially one extra pass through the network — backwards. The math behind this is the chain rule (you saw it in the Calculus track). The engineering is: cache intermediate values during the forward pass, then walk backward, multiplying local Jacobians as you go.
For each weight in layer :
That's a chain of derivatives. Backprop computes each once, reuses them across all weights.
The Blame-Game Metaphor
Think of backprop as blame propagation. The loss says "we're off the bullseye by this much." Backprop walks backward through the network, asking each layer: "how much did you contribute to the miss?" Each layer answers in proportion to its weight × the gradient flowing in from above. Then it passes the (chained) blame to the layer behind it.
By the end, every weight knows its share of the error. Adjust each weight against its share, and the next forward pass gets closer to the bullseye.