C.W.K.
Stream
Quiz · 4 questions

🔥 Training & Generation

Loss, schedules, decoding, and alignment

Level 0Token
0 XP0/94 lessons0/10 achievements
0/120 XP to next level120 XP to go0% complete

Quiz

01What pretraining objective do GPT, Llama, and Mistral use?
Hint
It's the same thing the model does at inference — predict the next token.
02What does DPO eliminate compared to PPO/RLHF?
Hint
DPO's simplicity comes from skipping a major component of RLHF.
03What does temperature=0 mean in decoding?
Hint
Think about what 1/T does as T approaches zero.
04What is the modern default floating-point format for training large LLMs?
Hint
It has the range of FP32 but half the bits, and doesn't need GradScaler.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign inPlease sign in to comment.

No comments yet — be the first.