Skip to main content
Advancedready7/10

AI Evaluation

Evaluation datasets, quality metrics, calibrated model graders, regression tests, observability and benchmarks.

6materials
43min reading
2levels
Start2
Core2
Build2
Advanced4
Start with the first material →

Learning route

Study route

Open in catalog →

Core level

1
  1. 01AI agent benchmarks: GAIA, WebArena, OSWorld, and SWE-benchA practical guide to choosing an AI agent benchmark: what GAIA, WebArena, OSWorld, and SWE-bench actually test, how to read results, and how to carry an external signal into your own release gate.Core level7 min

Advanced level

5
  1. 01AI agent trajectory evaluation: test the path, not only the resultA practical guide to evaluating AI agent trajectories: trace contracts, tool calls, permissions, retries, side effects, graders, failure taxonomy, and release gates.Advanced level7 min
  2. 02How to evaluate AI tool calling: a practical checklistA reproducible protocol for evaluating function calling and tool use: tool selection, arguments, trajectory, side effects, retries, terminal state, cost, and a release gate.Advanced level7 min
  3. 03How to evaluate browser agents: a practical checklistA reproducible release protocol for browser and computer-use agents: task state, visual grounding, trajectories, side effects, recovery, security, and risk-bounded rollout.Advanced level8 min
  4. 04How to evaluate LLM hallucinations: a practical checklistA reproducible protocol for evaluating factuality and groundedness: claim types, verified evidence, abstention, calibration, long-form responses, risk slices, and a release gate.Advanced level7 min
  5. 05How to evaluate prompt-injection defenses: a practical checklistA reproducible protocol for testing prompt-injection defenses: threat modeling, source-to-sink fixtures, tool traces, data exfiltration, side effects, false positives, release gates, and rollback.Advanced level7 min