C.W.K.
Stream
Quiz · 5 questions

⚖️ LLM-as-Judge

Using AI to evaluate AI — power and pitfalls

Level 0Guesser
0 XP0/55 lessons0/10 achievements
0/150 XP to next level150 XP to go0% complete

Quiz

01Why use an LLM as a judge for evaluations?
02What does putting reasoning before verdict in the judge's JSON output do?
03What is the standard defense against position bias in pairwise judging?
04When should you use one multi-criteria judge call vs N per-axis calls?
05Why calibrate a judge against human ratings?
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign inPlease sign in to comment.

No comments yet — be the first.