C.W.K.
Stream
Quiz · 5 questions

🎯 Attention Mechanism

Q, K, V — and the engineering that scales them

Level 0Token
0 XP0/94 lessons0/10 achievements
0/120 XP to next level120 XP to go0% complete

Quiz

01What does the √d_k scaling factor in attention prevent?
Hint
What happens to softmax(x) when the magnitude of x is much larger than 1?
02What is the KV-cache used for?
Hint
What can you cache once it's been computed?
03How many KV heads does Llama 3.3 70B use?
Hint
It's the standard GQA group size used by most modern open-weight models.
04What is Flash Attention's key innovation?
Hint
Same math, different memory strategy.
05Why is causal masking necessary in decoder-only training?
Hint
What's the difference between training in parallel and generating sequentially?
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign inPlease sign in to comment.

No comments yet — be the first.