How To Build AI Evals — Lucas Rocha

Lucas Rocha, an engineer at Brazilian edtech Nova Escola and alum of Hamel Husain’s AI Evals course, tells his evals rollout story backwards on Hamel’s channel — 34 minutes on going from a messy spreadsheet to four calibrated judges running in CI and on daily production samples. ...

July 17, 2026 · 5 min

How To Build AI That Earns User Trust — Hamel Husain

Hamel Husain makes the case that “are you building the right thing?” matters more than evals: if users can’t verify an AI answer, your product creates work instead of saving it. Three worked examples and four design principles. ~17 minutes on his own channel. The core problem: verification is the bottleneck Classic AI product: user asks “what was revenue for product A last quarter?”, AI queries a database and returns “$4.21M” — but the user can’t validate the number without redoing all the work Lenny’s tweet: data science teams now spend their time reviewing half-assed AI analysis from PMs and data engineers — 50% of the time it’s wrong Creating output is now easy; verification is the bottleneck — design products with that in mind If a product seems impossible to eval, that’s a product smell: your users have to check it too Example 1: business Q&A → evidence-backed analysis Instead of a bare answer, show supporting information: base the analysis on already-vetted analysis (notebooks others signed off on) State the assumptions — confirm the metric definition (from a semantic layer, linkable to the governed definition) Show intermediate calculations (returns, customer counts) and raise issues the user might care about “Open as notebook” UI idea: an AI-generated notebook (Jupyter/Marimo style) with narrative + queries — step through it, edit it, ask AI questions about it, and have the AI state what it couldn’t verify Put yourself in the expert’s shoes: a data scientist checks prior analyses, definitions, and intermediate calcs — give users those same affordances Real product doing this today: Hex — a chat interface that flows into a notebook; progressive disclosure is an established pattern we forgot in AI apps Example 2: PE lesson plan assistant (K-12) Vanilla version: input grade, class length, location, equipment → AI spits out a lesson plan — no way to judge quality Better: anchor the output in existing curriculum — what are trusted colleagues doing? “Here’s a lesson plan exactly like yours, used at 14 schools, run 30 times, by teachers you trust” Show the edits made for your inputs (“shortened to a 45-minute class, matched to your equipment”) — accept or reject each Cold-start problem: seed with expert-vetted plans up front; it makes the app easier to eval and users more confident Example 3: medical claims report generator The task: read a patient’s chart + claim (thousands of documents) and produce a 52-page report supporting or denying a claim — a doctor can’t eyeball 52 pages Redesign as a research assistant: surface atomic elements first — contradictions (open both pages, resolve), key facts (validate, include, dismiss), open questions (add notes) Do the work the user would normally do and guide them through it; understanding built along the way is arguably more valuable than the final report Four design principles Provenance — show where information comes from (prior analyses, source documents, other lesson plans); the more curated/trustworthy, the better Progressive disclosure — show details at the right time; adapt as your product and users mature Signals and heuristics — mirror how experts sanity-check: other reports, smell tests, social proof, supporting evidence Modularity / small steps — break work into small verifiable pieces, show your work along the way, and sometimes force the human in the loop (e.g. doctors step through findings before the final report) “The goal is to make the human understand, not necessarily to generate a report.” ...

July 7, 2026 · 3 min

How to Automate AI Evals (Correctly) — Shreya Shankar

Shreya Shankar (Stanford CS professor, co-creator of the AI evals course with Hamel) kicks off the 12-part AI product engineering series. 27 minutes on Hamel Husain’s channel. Why this matters Output quality is the biggest barrier to productionizing agents (LangSmith annual report) — and figuring out how to evaluate models is genuinely hard Vendors (LangChain, Braintrust, Arize) are selling end-to-end automated eval tools: point an LLM at your traces, it finds and fixes your bugs The catch is epistemic: what “good” means lives in your head, not in the traces — if a tool could fully fix your product, it could fix everyone’s, and there’d be nothing left to differentiate yours AI’s real job: help you express and apply your judgment faster, not replace it The eval lifecycle (analyze → measure → improve) Error analysis — the hardest step: take traces and find failure modes. No perfect definition of “mistake” (you can’t define slop, but you know it when you see it) Measure — how prevalent is each failure mode? Pareto applies: ~80% of issues come from ~20% of failure modes — prioritize those Improve — fix the product: prompt instructions, model switch, fine-tuning. Iterate forever AI is weak at the front (taste-specific error analysis) and strong at the back (measurement, prompt optimization, hill-climbing). ...

July 3, 2026 · 4 min

How to Eval an AI Product That Seems Impossible to Eval — Hamel Husain

Hamel Husain answers a client question — how to eval an AI PE curriculum planner for K-12 — with a one-minute twist: evals and product trust are the same problem. ~1 minute on his own channel. The eval problem A client building an AI curriculum planner for K-12 physical education couldn’t figure out how to eval it The answer: think creatively — put yourself in the educator’s shoes Questions an educator actually asks Was this plan created by an expert? What other teachers are using this plan? Is this plan based on best practices and expert knowledge? The takeaway Infuse those signals into the product — they give the end user trust in what they’re seeing That same design doubles as your eval: provide users a way to assess what they see, and you can assess it too If you can’t eval your product, your product probably sucks “If you can’t eval your product, your product probably sucks.” ...

June 29, 2026 · 1 min

If You Can't Eval Your AI, Your Users Can't Trust It — Hamel Husain

Hamel Husain on his own channel — a 1-minute clip on the “evals take too long” objection and why your users eval your product whether you do or not. The “evals take too long” objection One of the main reasons people skip evals: reading a trace takes forever — a lot of human hours you “don’t have time to do” His example: an AI data agent that answers business questions (supply, revenue) — “the most popular internal tool I’ve seen so far” The question: how do you eval that? If you can’t eval it, it probably sucks One principle to keep in mind: if you can’t eval your product, your product probably sucks — because your users need to eval the product When you present an AI answer to a user — say, an answer to a business question — you need to figure out a way to convey trust “If you can’t eval your product, your product probably sucks. Why is that? It’s because your users need to eval the product.” — Hamel Husain ...

June 26, 2026 · 1 min

Why There's No LLM Judge You Can Trust 100% of the Time — Hamel Husain

Hamel Husain on his own channel — a 1.5-minute clip on LLM-as-a-judge evals: how to measure the judge’s noise and why a judge you can trust 100% of the time doesn’t exist. Measure the judge against ground truth An LLM judge is a “soft check” — you can scientifically measure how good it is by comparing it to labels Assemble a dataset of human labels as ground truth (is this thing good or bad?) and measure how much noise the judge has Even ~100 labels is enough to measure the noise There are different kinds of noise — you can tune the judge to reduce it, but there will always be some A judge is a black-box classifier There’s no such thing as a judge you can blindly trust with no noise at all It’s a classifier, like a stop-sign detector: get it very reliable and there’s still an edge case somewhere producing false positives and false negatives You have to be okay with that Tolerable noise is a business decision Tune the judge to reduce noise, then make the call: is this level of noise tolerable for your use case? Bake the eval’s known characteristics into your decision making — that’s the best you can do with an LLM judge “Even if you can get that to be very reliable, there’ll always be an edge case somewhere where it’s going to give you a false positive, false negative.” — Hamel Husain ...

June 22, 2026 · 2 min

Why AI Rating Scales Make It Harder to Ship — Hamel Husain

Hamel Husain on his own channel — a 2-minute clip on why traditional scoring metrics (ROUGE/BLEU) and rating scales make evals harder to act on, and why he defaults to binary pass/fail. ROUGE/BLEU: string similarity from a different era These scores come from traditional ML/NLP — at their core they measure string similarity A coarse metric that made sense when LLMs could barely produce coherent language; “we’re way past the point of producing coherent English” now Part of the “eval industrial complex”: we have metrics, trust us because we’re experts — it feels fancy but isn’t what you need Rating scales hide uncertainty Unless you have a very sophisticated setup and put real resources into aligning scores with humans — which ~99% of teams don’t — a 1-to-10 or 1-to-5 scale is a bad idea Nobody knows what 4.2 versus 3.7 means Humans just hide their uncertainty in the middle values, so you don’t get good decision making out of it Binary pass/fail is what you can ship on At the end of the day you have to ship your product: is this good or not? Binary evals are easier to align with humans — which you always have to do with an eval — and easier to action A score of 3.2 → “I don’t know what’s wrong.” A fail → “okay, it failed, now you can action on that” “You see a score of 3.2, you’re like, I don’t know what’s wrong… but fail, it’s like, okay, it failed, now you can action on that.” — Hamel Husain ...

June 21, 2026 · 2 min

Why the 1 to 5 Scale Is Where AI Evals Break Down — Hamel Husain

Hamel Husain on his own channel — a 40-second clip on why rating-scale evals (1–5) are the wrong default: he pushes for binary pass/fail evals with failures scoped into specific, actionable criteria. Binary beats a 1–5 scale He tries “really hard” to make every eval binary — pass or fail Scope the failure: “agent took too long → failed because it took too long”, “too many steps → failed because it took too many steps” He has never gotten stuck converting a 1–5 scale into binary, across all the companies he’s worked with The exception: a well-tested rubric everyone is confident in — then a 1–5 scale might genuinely work In most cases the scale just “kicks the can down the road” — the noise doesn’t disappear, it gets hidden “In a lot of cases, it just kind of kicks the can down the road, and you’re just hiding a lot of noise in there.” — Hamel Husain ...

June 20, 2026 · 1 min

The AI Testing Framework Every Business Needs (But Few Use)

Hamel Husain (founder of Parlance Labs, co-author of an upcoming book on AI evals with Shreya Shankar) on the Applied Intelligence podcast — 39 minutes on why evals are the testing framework most businesses skip, and how to think about them the right way. What an eval actually is An eval is structured data analysis and debugging on your AI application — you know what’s broken, and you know what to prioritize fixing It’s the answer to a real problem: AI outputs are stochastic (text), with no deterministic “right or wrong” to test against Software testing assumes a finite surface of failure; AI has an infinite one, so you have to decide what to measure Generic metrics are a trap Search “how to evaluate my AI app” and you’ll find vendors promising a dashboard of helpfulness, conciseness, toxicity, coherence scores (usually 1–5) Nobody knows what those numbers mean, they don’t correlate with what matters, and they actively burn engineering cycles looking at metrics that don’t matter “Reduce hallucination, increase helpfulness” is the same generic-metrics mindset — you’re going through motions without knowing what’s actually wrong There’s no easy button; you can use coding agents to write the evals, but you still have to be thoughtful about what you measure Start bottom-up, not top-down The process starts with error analysis: a blend of qualitative and quantitative work to find what’s broken in your application Most people are too top-down (worried about failure types they imagine) and get lost in generic metrics; the bottoms-up read of your actual data is the missing half Evals are really a process of eliciting your specification, then measuring against it — you can’t know “good” until you look at outputs and iterate AI can’t do your evals for you AI is great at fixing deterministic bugs; it cannot read your mind — it doesn’t know what “good” feels like for your product The future is AI walking you through the evals process and interrogating you, not doing it alone — a human stays in the loop Eval tooling needs to render your data domain-specifically (images as images, emails as emails, chat as chat), so it’s more than a chatbot Two failure modes Hamel sees most Not using AI deeply yourself — no coding with AI, no building with it → bad intuition, bad specifications Reaching for complexity too fast — day-one orchestration frameworks, graph databases, multi-agent setups before you can reason about what’s happening Deliberate, not slow AI lets you build the wrong thing faster — and many wrong things faster; it also lets you build the right thing faster You need product sense, grounding, and taste, or you just churn through bad ideas at higher speed AI amplifies who you are: if you’re okay with slop, it removes the friction and amplifies the slop Chatbot vs MCP Slapping a chatbot on an existing product is the mediocre default when there’s a “we need AI” mandate from the top Better: expose an MCP or an API on your product — the Google Workspace CLI is the canonical example of how much more useful agent-friendly access is Interface matters: a scheduling flow shouldn’t be a brittle text back-and-forth when a picker widget gives visual confirmation and avoids bugs Guardrails A guardrail is a specific eval sitting in the request/response path that blocks a bad output (competitor talk, profanity) Off-the-shelf guardrails are just someone else’s prompt, tuned to a different domain (shopping, travel) — you still have to do the evals work to know which failures to guard against Prioritize failures you can simulate or actually observe; unobservable ones are lower priority The stack, the vendor, and the team Start with the most powerful model you already know, build an eval harness around the metrics that matter, then back off to smaller/cheaper models and reason about the latency–cost–quality tradeoffs Data science gets more important, not less: the ability to ask the right questions is directly proportional to the quality of output you get AI competency is core to any knowledge-work business, so be careful outsourcing it — use third parties to upskill your team (a deliberate training exercise), never as a crutch The takeaway Parlance’s model is a “driving school” boot camp: they pair-program the whole end-to-end evals process on your data until you don’t need them The million-lines-of-code demos are all backed by harness engineering — metrics, logs, traces, observability — and evals are almost all of that harness You don’t need an R&D budget, maybe just a token budget: a $100 plan and deliberate experimentation gets you a long way “AI cannot read your mind. It doesn’t know what you feel like good is.” ...

April 20, 2026 · 4 min