Hamel Husain ranks common AI-evaluation practices after working with more than 50 product teams, in a 13-minute video on his channel.

Start with real work, not a scorecard

  • S: Error analysis and reading actual agent traces — inspect failures before deciding what to measure.
  • S: Build an annotation view tailored to the product: render an email-writing agent’s output like an email, not a spreadsheet row.
  • C: Brainstorming hypothetical failures or writing tests before building can start a conversation, but is a poor way to prioritize an unlimited failure surface.
  • F: Generic helpfulness/coherence scores, vendor-provided judges, or choosing an eval vendor before understanding your own data.

Build checks people can trust

  • A: Binary pass/fail criteria beat vague 1–5 ratings; middle scores obscure the decision you actually need to make.
  • S: Compare LLM-judge decisions with human labels. Prefer cheap code-based checks where they work, and validate those against people too.
  • F: One judge prompt for every kind of failure, or off-the-shelf guardrails not tuned to the product.
  • D: Letting only developers annotate a medical or legal assistant; the relevant domain expert needs to judge the outputs.
  • S: Put domain experts close enough to the product to write prompts, not merely review them at the end.

Manage the evaluation portfolio

  • A suite that always passes (D) isn’t stretching the product. Run persistently passing, costly checks less often (A); not every isolated failure warrants its own eval (F).
  • Public benchmarks may help compare models, but cannot replace product-specific checks (F).
  • A trusted decision-maker for each feature or domain can break annotation deadlocks (A); showing intermediate outputs makes expert verification easier (B).
  • Prompt optimization without evals (F) and letting AI do all the evaluating (F) cannot substitute for people inspecting difficult, contextual failures.

“Look at data first.”