Hamel Husain walks through eight agent skills he and Shreya Shankar built from work with 50+ companies and 5,000+ eval-course students. The 12-minute tour focuses on a simple correction: inspect real failures before inventing tests.

Start with the router

  • eval-start is the entry point — it inspects what you already have and routes Claude Code or Codex to the right specialist skill
  • Existing traces but no failure taxonomy → start with error discovery
  • Existing evals that may be unreliable → start with the eval audit
  • Hamel’s priority: master those two before reaching for the rest of the toolkit

Error discovery: turn traces into a failure taxonomy

  1. Understand the data — parse its structure, identify the dimensions that vary most, and ask what “bad” could mean
  2. Design a review interface — use color, spacing, contrast, typography, and containers to make the important differences easy to scan
  3. Build the interface around the data — a multi-turn voice-agent trace needs a different view from a document or search result
  4. Sample for discovery — combine random examples with cluster-based diversity instead of reading logs one by one
  5. Adapt while the human reviews — annotations steer the next sample toward both breadth across the dataset and depth around a newly found failure
  • The goal at this stage is finding kinds of failures, not estimating how frequently each one occurs
  • The loop keeps serving examples until the reviewer stops discovering new failure modes
  • The human supplies judgment; the coding agent handles parsing, interface construction, clustering, and the repetitive search for related traces

The eval audit: catch common methodology mistakes

  • Confirms that proposed checks come from observed, annotated failures rather than brainstormed hypotheticals
  • Prefers focused binary pass/fail criteria over noisy letter grades or Likert-style scales
  • Flags over-scoped LLM judges and recommends deterministic code checks where code can answer the question
  • Warns against treating text-similarity scores as general-purpose quality measures
  • Requires each LLM judge to be tested against human labels
  • Checks train/dev separation so judge prompts are not tuned on the same labels used to report their quality
  • Reviews annotation practice, evaluator maintenance, recurring error analysis, and coverage per failure mode; Hamel offers roughly 100 labeled traces per failure mode as a useful target, not a law
  • Produces a diagnostic report instead of silently changing the eval setup

The other six skills

  • Generate synthetic data — stress-test a pre-product system when real users and traces do not yet exist, with emphasis on diverse edge cases
  • Write a judge prompt — create a narrowly scoped LLM-as-judge rubric using patterns that held up across consulting engagements
  • Validate an evaluator — compare a judge with human labels using clean splits before trusting its scores
  • Evaluate RAG — build checks around retrieval and search quality, not only the final generated answer
  • Build a review interface — generate a purpose-built annotation UI; error discovery calls this skill internally
  • eval-start — tie the workflow together and route the agent based on the project’s current stage

The practical order

  • Start with traces from the actual product
  • Use error discovery to understand the failures and create labels
  • Audit the resulting eval plan for weak criteria, leakage, and unvalidated judges
  • Only then add synthetic cases, judge prompts, retrieval metrics, and production automation

“The goal is discovering failure modes, not estimating how common they are.”