C.W.K.
Stream
← C.W.K. Quests
📏

Eval Quest

Updated: 2026-05-04

Measure what matters in AI systems

Build evaluation practice for prompts, models, RAG, tools, agents, safety, and production AI workflows — from vibes-based inspection to a measurement discipline that ships.

8 tracks · 55 lessons · ~32h · difficulty: beginner-to-advanced

Level 0Guesser
0 XP0/55 lessons0/10 achievements
0/150 XP to next level150 XP to go0% complete
Eval Quest is the bridge from prompt craft to measurable AI engineering. It walks from the failure of vibes-based inspection through the design of datasets and graders, the choice of frameworks and benchmarks, the evaluation of systems and agents, and the operating practices that keep evaluation honest over time. Whether you ship a chatbot, a RAG pipeline, a coding agent, or an internal tool, Eval Quest gives you the discipline to know — not guess — whether each change makes things better or worse.

Tracks

  1. 01🎯Why Evals Matter

    0/9 lessons

    From vibes-based inspection to repeatable evidence

    The mindset shift from 'looks good to me' to 'here is the measurement that protects this behavior.' Every later track depends on this discipline taking root.

    Lesson list (9)Quiz · 5 questions
  2. 02🗃️Datasets and Golden Cases

    0/7 lessons

    The dataset is the ceiling on every evaluation that follows

    Curate the inputs that represent your real workload. No grader, framework, or judge can rescue an eval suite built on the wrong data.

    Lesson list (7)Quiz · 5 questions
  3. 03📊Deterministic Metrics

    0/6 lessons

    The cheap, fast, reproducible measurements that should run first

    Exact match, BLEU, ROUGE, BERTScore, regex, format checks, composites. The metrics you should always exhaust before reaching for an expensive LLM judge.

    Lesson list (6)Quiz · 5 questions
  4. 04⚖️LLM-as-Judge

    0/7 lessons

    Using AI to evaluate AI — power and pitfalls

    When deterministic graders cannot reach the question, an LLM judge can. But judges have biases, costs, and calibration problems. This track is how to wield them without being wielded by them.

    Lesson list (7)Quiz · 5 questions
  5. 05🔧Frameworks and Platforms

    0/7 lessons

    promptfoo, DeepEval, Braintrust, lm-eval-harness, RAGAS, Inspect AI

    The tools you'll actually use in production. Each has a sweet spot — pick for the job, not for the brand.

    Lesson list (7)Quiz · 5 questions
  6. 06🏆Public Benchmarks

    0/7 lessons

    Reading the leaderboards without believing them

    MMLU, HumanEval, GSM8K, HellaSwag, Chatbot Arena, MTEB, ARC-AGI. What each measures, when it lies, and how to use them for model selection without confusing them with product evals.

    Lesson list (7)Quiz · 5 questions
  7. 07🔬System and Agent Evals

    0/6 lessons

    Evaluating compositions, not single calls

    RAG, agents, chatbots, code assistants, A/B production tests, regression suites. Real systems are pipelines; their evals must be too.

    Lesson list (6)Quiz · 5 questions
  8. 08🛡️Safety and Operations

    0/6 lessons

    Red teaming, calibration, cost-quality, building eval culture

    Where evaluation meets the rest of the company — security, reliability, cost discipline, and the team practices that make all of it stick.

    Lesson list (6)Quiz · 5 questions
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign inPlease sign in to comment.

No comments yet — be the first.