Hamel Husain uses a 39-second example to show how a high agreement score can hide a judge that detects no failures at all.
The accuracy trap
- Suppose a product fails on 10% of examples and passes on the other 90%.
- A judge that always predicts “pass” will agree with the human labels 90% of the time.
- The headline score sounds strong even though the judge misses every failure—the cases the eval was built to find.
What teams should inspect
- Treat overall agreement as a warning sign when one outcome is much more common than the other.
- Break results into false positives and false negatives instead of compressing them into one percentage.
- Check recall on the failure class: of the real failures, how many did the judge actually catch?
- Review the confusion matrix and the underlying labeled examples before presenting the metric to stakeholders.
The product lesson
- A metric is useful only if it measures the behavior that matters operationally.
- Product managers should pause whenever “agreement” appears without a class-by-class error analysis.
- A judge that reproduces the majority label is not necessarily evaluating anything.
“You can agree with it 90% of the time by just saying it never fails.” — Hamel Husain