Hamel Husain draws a one-minute boundary between two things that are both called “evals” but answer different questions.
Model benchmarks
- Tests such as SWE-bench and GPQA Diamond compare general model capability across coding or advanced academic tasks.
- They help with model selection and track broad progress.
- A strong score does not establish that a complete product will perform its own workflow correctly.
Product evals
- A product eval measures the behavior of a specific application—including its prompts, tools, retrieval, orchestration, and business rules.
- The test should express what success means for the actual user task.
- For an order-management agent, that might mean checking both that the correct order was selected and that the cancellation really occurred.
The practical distinction
- Use benchmarks to learn what a model can do in general.
- Use product evals to learn whether the system you built does what you intended.
- Shipping decisions require the second kind, because users interact with the product—not an isolated leaderboard score.
“Product evals measure whether your specific AI product does what you want it to do.”