C.W.K.
Stream
Quiz · 4 questions

📐 Embeddings & Position

Token IDs to dense vectors with order

Level 0Token
0 XP0/94 lessons0/10 achievements
0/120 XP to next level120 XP to go0% complete

Quiz

01What is the shape of the embedding matrix in a Transformer with vocab=128,000 and d_model=4,096?
Hint
Rows index over vocabulary; columns index over hidden dimensions.
02Which positional encoding scheme do Llama 3, Mistral, and Qwen all use?
Hint
It's applied as a rotation inside attention rather than as an addition at the input.
03Why is self-attention permutation-equivariant without positional encoding?
Hint
Look at where position would have to appear in the dot-product formula.
04What technique did Llama 3 use to extend its context window from 8K to 128K?
Hint
It's a clever rescaling of RoPE's frequencies, not a brand-new architecture.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign inPlease sign in to comment.

No comments yet — be the first.