Hamel Husain uses a 39-second example to show how a high agreement score can hide a judge that detects no failures at all.

The accuracy trap

  • Suppose a product fails on 10% of examples and passes on the other 90%.
  • A judge that always predicts “pass” will agree with the human labels 90% of the time.
  • The headline score sounds strong even though the judge misses every failure—the cases the eval was built to find.

What teams should inspect

  • Treat overall agreement as a warning sign when one outcome is much more common than the other.
  • Break results into false positives and false negatives instead of compressing them into one percentage.
  • Check recall on the failure class: of the real failures, how many did the judge actually catch?
  • Review the confusion matrix and the underlying labeled examples before presenting the metric to stakeholders.

The product lesson

  • A metric is useful only if it measures the behavior that matters operationally.
  • Product managers should pause whenever “agreement” appears without a class-by-class error analysis.
  • A judge that reproduces the majority label is not necessarily evaluating anything.

“You can agree with it 90% of the time by just saying it never fails.” — Hamel Husain