Hamel Husain on his own channel — a 40-second clip on why rating-scale evals (1–5) are the wrong default: he pushes for binary pass/fail evals with failures scoped into specific, actionable criteria.
Binary beats a 1–5 scale
- He tries “really hard” to make every eval binary — pass or fail
- Scope the failure: “agent took too long → failed because it took too long”, “too many steps → failed because it took too many steps”
- He has never gotten stuck converting a 1–5 scale into binary, across all the companies he’s worked with
- The exception: a well-tested rubric everyone is confident in — then a 1–5 scale might genuinely work
- In most cases the scale just “kicks the can down the road” — the noise doesn’t disappear, it gets hidden
“In a lot of cases, it just kind of kicks the can down the road, and you’re just hiding a lot of noise in there.” — Hamel Husain