Hamel Husain answers a question about stale gold datasets in a one-minute clip. YouTube reports this particular video cannot be embedded; use the Watch link above.
Keep the challenge relevant
- Gold datasets go stale as a product improves; periodically inspect new errors and add fresh examples.
- LLM-judged evaluations cost money to build and run. An eval that always passes may no longer be worth running as often.
- Evals are a moving target for improvement, not a fixed scoreboard to optimize forever.
- For a stable long-term comparison, watch product metrics such as churn, revenue, and monthly active users instead.
“You should always have a new challenge as you get better.”