Hamel Husain on his own channel — a 40-second clip on why rating-scale evals (1–5) are the wrong default: he pushes for binary pass/fail evals with failures scoped into specific, actionable criteria.

Binary beats a 1–5 scale

  • He tries “really hard” to make every eval binary — pass or fail
  • Scope the failure: “agent took too long → failed because it took too long”, “too many steps → failed because it took too many steps”
  • He has never gotten stuck converting a 1–5 scale into binary, across all the companies he’s worked with
  • The exception: a well-tested rubric everyone is confident in — then a 1–5 scale might genuinely work
  • In most cases the scale just “kicks the can down the road” — the noise doesn’t disappear, it gets hidden

“In a lot of cases, it just kind of kicks the can down the road, and you’re just hiding a lot of noise in there.” — Hamel Husain