Quiz · 4 questions
📐 Embeddings & Position
Token IDs to dense vectors with order
Level 0Token
0 XP0/94 lessons0/10 achievements
0/120 XP to next level120 XP to go0% complete
Quiz
01What is the shape of the embedding matrix in a Transformer with vocab=128,000 and d_model=4,096?
Hint
Rows index over vocabulary; columns index over hidden dimensions.
02Which positional encoding scheme do Llama 3, Mistral, and Qwen all use?
Hint
It's applied as a rotation inside attention rather than as an addition at the input.
03Why is self-attention permutation-equivariant without positional encoding?
Hint
Look at where position would have to appear in the dot-product formula.
04What technique did Llama 3 use to extend its context window from 8K to 128K?
Hint
It's a clever rescaling of RoPE's frequencies, not a brand-new architecture.
Comments 0
🔔 Reply notifications (sign in)Sign in — Please sign in to comment.
No comments yet — be the first.