Advancedready7/10
AI Evaluation
Evaluation datasets, quality metrics, calibrated model graders, regression tests, observability and benchmarks.
6materials
43min reading
2levels
Start2
Core2
Build2
Advanced4
Learning route
Study route
Core level
1Advanced level
5- 01AI agent trajectory evaluation: test the path, not only the resultA practical guide to evaluating AI agent trajectories: trace contracts, tool calls, permissions, retries, side effects, graders, failure taxonomy, and release gates.Advanced level7 min
- 02How to evaluate AI tool calling: a practical checklistA reproducible protocol for evaluating function calling and tool use: tool selection, arguments, trajectory, side effects, retries, terminal state, cost, and a release gate.Advanced level7 min
- 03How to evaluate browser agents: a practical checklistA reproducible release protocol for browser and computer-use agents: task state, visual grounding, trajectories, side effects, recovery, security, and risk-bounded rollout.Advanced level8 min
- 04How to evaluate LLM hallucinations: a practical checklistA reproducible protocol for evaluating factuality and groundedness: claim types, verified evidence, abstention, calibration, long-form responses, risk slices, and a release gate.Advanced level7 min
- 05How to evaluate prompt-injection defenses: a practical checklistA reproducible protocol for testing prompt-injection defenses: threat modeling, source-to-sink fixtures, tool traces, data exfiltration, side effects, false positives, release gates, and rollback.Advanced level7 min