Hamel Husain answers a question about stale gold datasets in a one-minute clip. YouTube reports this particular video cannot be embedded; use the Watch link above.

Keep the challenge relevant

  • Gold datasets go stale as a product improves; periodically inspect new errors and add fresh examples.
  • LLM-judged evaluations cost money to build and run. An eval that always passes may no longer be worth running as often.
  • Evals are a moving target for improvement, not a fixed scoreboard to optimize forever.
  • For a stable long-term comparison, watch product metrics such as churn, revenue, and monthly active users instead.

“You should always have a new challenge as you get better.”