Why 90% Agreement Is a Bad Sign for Your LLM Judge — Hamel Husain

Hamel Husain uses a 39-second example to show how a high agreement score can hide a judge that detects no failures at all. The accuracy trap Suppose a product fails on 10% of examples and passes on the other 90%. A judge that always predicts “pass” will agree with the human labels 90% of the time. The headline score sounds strong even though the judge misses every failure—the cases the eval was built to find. What teams should inspect Treat overall agreement as a warning sign when one outcome is much more common than the other. Break results into false positives and false negatives instead of compressing them into one percentage. Check recall on the failure class: of the real failures, how many did the judge actually catch? Review the confusion matrix and the underlying labeled examples before presenting the metric to stakeholders. The product lesson A metric is useful only if it measures the behavior that matters operationally. Product managers should pause whenever “agreement” appears without a class-by-class error analysis. A judge that reproduces the majority label is not necessarily evaluating anything. “You can agree with it 90% of the time by just saying it never fails.” — Hamel Husain ...

October 1, 2026 · 1 min

Why AI Benchmarks Don't Tell You If Your Product Works — Hamel Husain

Hamel Husain draws a one-minute boundary between two things that are both called “evals” but answer different questions. Model benchmarks Tests such as SWE-bench and GPQA Diamond compare general model capability across coding or advanced academic tasks. They help with model selection and track broad progress. A strong score does not establish that a complete product will perform its own workflow correctly. Product evals A product eval measures the behavior of a specific application—including its prompts, tools, retrieval, orchestration, and business rules. The test should express what success means for the actual user task. For an order-management agent, that might mean checking both that the correct order was selected and that the cancellation really occurred. The practical distinction Use benchmarks to learn what a model can do in general. Use product evals to learn whether the system you built does what you intended. Shipping decisions require the second kind, because users interact with the product—not an isolated leaderboard score. “Product evals measure whether your specific AI product does what you want it to do.” ...

September 30, 2026 · 1 min

The Best Jev Use Cases for AI Coding — Cole Medin

Cole Medin presents four practical ways to pair Jev with coding agents. His central pattern is not to replace the LLM: let the LLM build the harness and act on results, while Jev makes frequent, bounded decisions in the middle. What Jev is for Jev takes a description of the current state plus a fixed set of questions or choices, then returns answers with confidence scores. Medin says this decision-only design is 20–200× faster and 40–1,000× cheaper than using a generative LLM for the same kind of classification; those are product claims, not independent measurements in this video. It cannot generate free-form text or invent the action space. The workflow must define the state, legal choices, thresholds, and fallback behavior. 1. A security hook before tool use Coding agents may read secrets, delete data, exfiltrate information, or drift away from the assigned task—even when instructions say not to. A pre-tool-use hook can inspect the tool name, arguments, working directory, and intended effect before allowing an action. Regex rules are fast but brittle: the same risky operation can be expressed as a shell command, Python script, direct file read, or many other forms. An LLM judge is more flexible, but adding a model call before thousands of tool calls creates latency and cost. Medin’s early comparison reports Jev at roughly a quarter-second per check, with fewer false positives than his regex hook and nearly all of his hand-labeled risky calls blocked. The video does not provide the dataset or raw results. 2. Playtesting as a user would Unit tests cannot show whether a game actually feels playable or whether a feature breaks during live interaction. Jev receives game state and chooses among legal actions quickly enough to attack, dodge, and navigate in real time. The run can expose bugs that deterministic tests miss because it exercises the game through the same interaction loop as a player. The LLM still builds the state adapter and action harness, then diagnoses and fixes whatever the playtest finds. 3. Faster browser testing Browser automation is another bounded loop: inspect the current page, choose a button or field, execute the action, and observe the next state. Jev can handle repeated choices such as what to focus or click; an LLM is still needed when a page requires original free-form text or broader visual judgment. Medin demonstrates this split on Dino Chat: text is generated in advance, Jev chooses interface actions, and an LLM can review the resulting trace for bugs. 4. Route work before spending tokens A GitHub issue can first be classified as a bug or feature, then routed to the matching skills and workflow. A second decision selects a model tier based on task difficulty, avoiding an expensive frontier model for routine work. Medin reports agreeing with Jev on all 12 issue-routing trials and on 15 of 16 pull-request review classifications. These are small, self-judged samples, not a production benchmark. The reusable architecture An LLM defines the harness, state representation, and available actions. Jev makes frequent, narrow decisions inside that structure. Deterministic code enforces thresholds and executes approved actions. The LLM interprets failures, repairs code, or handles open-ended work. “Jev is never really the end of our workflow. It’s always just making parts of our workflows faster and more efficient.” ...

September 30, 2026 · 3 min

8 Claude Code Skills for Better AI Evals — Hamel Husain

Hamel Husain walks through eight agent skills he and Shreya Shankar built from work with 50+ companies and 5,000+ eval-course students. The 12-minute tour focuses on a simple correction: inspect real failures before inventing tests. Start with the router eval-start is the entry point — it inspects what you already have and routes Claude Code or Codex to the right specialist skill Existing traces but no failure taxonomy → start with error discovery Existing evals that may be unreliable → start with the eval audit Hamel’s priority: master those two before reaching for the rest of the toolkit Error discovery: turn traces into a failure taxonomy Understand the data — parse its structure, identify the dimensions that vary most, and ask what “bad” could mean Design a review interface — use color, spacing, contrast, typography, and containers to make the important differences easy to scan Build the interface around the data — a multi-turn voice-agent trace needs a different view from a document or search result Sample for discovery — combine random examples with cluster-based diversity instead of reading logs one by one Adapt while the human reviews — annotations steer the next sample toward both breadth across the dataset and depth around a newly found failure The goal at this stage is finding kinds of failures, not estimating how frequently each one occurs The loop keeps serving examples until the reviewer stops discovering new failure modes The human supplies judgment; the coding agent handles parsing, interface construction, clustering, and the repetitive search for related traces The eval audit: catch common methodology mistakes Confirms that proposed checks come from observed, annotated failures rather than brainstormed hypotheticals Prefers focused binary pass/fail criteria over noisy letter grades or Likert-style scales Flags over-scoped LLM judges and recommends deterministic code checks where code can answer the question Warns against treating text-similarity scores as general-purpose quality measures Requires each LLM judge to be tested against human labels Checks train/dev separation so judge prompts are not tuned on the same labels used to report their quality Reviews annotation practice, evaluator maintenance, recurring error analysis, and coverage per failure mode; Hamel offers roughly 100 labeled traces per failure mode as a useful target, not a law Produces a diagnostic report instead of silently changing the eval setup The other six skills Generate synthetic data — stress-test a pre-product system when real users and traces do not yet exist, with emphasis on diverse edge cases Write a judge prompt — create a narrowly scoped LLM-as-judge rubric using patterns that held up across consulting engagements Validate an evaluator — compare a judge with human labels using clean splits before trusting its scores Evaluate RAG — build checks around retrieval and search quality, not only the final generated answer Build a review interface — generate a purpose-built annotation UI; error discovery calls this skill internally eval-start — tie the workflow together and route the agent based on the project’s current stage The practical order Start with traces from the actual product Use error discovery to understand the failures and create labels Audit the resulting eval plan for weak criteria, leakage, and unvalidated judges Only then add synthetic cases, judge prompts, retrieval metrics, and production automation “The goal is discovering failure modes, not estimating how common they are.” ...

September 30, 2026 · 3 min

Why the Prompt Matters Less Than the Context — Hamel Husain

Hamel Husain and Isaac Flath continue their live comparison of writing prompts with a five-minute clip about an inconvenient result: the carefully crafted prompt does not reliably beat the alternatives. The test Three setups answer the same writing task: no system prompt, Anthropic’s long mannered-prose guidance, and a shorter plain-writing prompt Anthropic’s guidance tells the model to prefer literal statements over metaphors such as “a dial worth turning” Hamel expected the guidance to backfire because it puts examples of the unwanted style directly into context The hosts compare the drafts by reading them, choosing between them, and discussing how much editing each would require What happened No candidate wins every example; useful passages and awkward “AI slop” appear across all three The long anti-slop guidance performs much better than Hamel expected, despite being written partly in the style it criticizes A short prompt can be competitive with the elaborate version, and sometimes no writing prompt at all is close enough to win The better output depends partly on the goal: the draft with the strongest structure may still need more line editing than a plainer alternative The practical lesson Treat writing prompts as candidates to evaluate, not permanent truths to adopt on reputation Compare them on several representative examples; five examples can reveal surprises but cannot establish a universal winner Judge the complete working context — source material, task, examples, and desired output — rather than crediting the system prompt alone Optimize for the draft that is easiest to turn into good writing, not the one backed by the most impressive prompt “There is no winner all the time necessarily. You could have good things from each.” ...

September 28, 2026 · 2 min

Behind Navigator n2: How Yutori Made Computer-Use Agents 10x Cheaper — Skanda Vaidyanath

Skanda Vaidyanath, founding AI researcher at Yutori, presents Navigator n2 at TwoSetAI Workshop #6 (55 min, hosted by Angelina Yang): how a small team trained a 27B computer-use model that takes first place on four of five computer-use benchmarks, and runs long tasks for roughly $1.46 where frontier models cost $10 to $30. ...

September 25, 2026 · 6 min

How to Get Agents to Do Evals Well — Hamel Husain

Hamel Husain on his own channel — a 57-second answer to the question every team doing evals eventually asks: how do you steer an agent to do this work well — skills, MCP servers, agent markdown files? Start with the person, not the agent The question: how do you constrain or steer the agent to do evals well as a human partner? Is it equipping it with skills and MCP servers you built and maintain, or agent markdown files? Hamel’s answer does not start with tooling: “my answer starts with the person, not the agent” You have to educate the human first, and the human has to understand the process — what good looks like, and what you are actually trying to do Only after that do the skills help; the agent is not the bottleneck the question assumes Why skills have a ceiling Skills “can only help you so much” — there is an upper limit to what they can do for you The reason is structural: a published skill is generic by design, because it has to apply across a lot of different use cases Getting value out of it means putting in “a little bit more thinking” to customize it for your own problem The whole thing “comes down to learning what good evals look like, plus using the skills, and then customizing the skills” for yourself The loop he describes Learn what good evals look like — the judgment, not the tooling Use the skills Customize them to your use case Same frame he applies to coding: the skill is scaffolding around the process knowledge, not a replacement for it The resources in the clip The evals skills are open — hamelsmu/evals-skills on GitHub, with updates landing as recently as the day before this clip The next AI Evals cohort with Hugo (Bowne-Anderson) is open at maven.com/parlance-labs/evals “It’s just like anything else, it’s just like coding or anything else. You need to educate the human and you need to give them some skills, and the human needs to understand the process of what good looks like. The skills can only help you so much.” — Hamel Husain ...

September 24, 2026 · 2 min

Building a Personal Agent That Actually Works and Sounds Like You — Ahmet İlten, Sentience

TwoSetAI Workshop #5. Ahmet İlten, founding engineer at Sentience (agent architecture and quality), walks through their personal-agent product and the four principles his team arrived at the hard way. Benjamin Carsley was scheduled to cover the voice/tone half but was out sick, so the session is Ahmet solo plus Q&A with host Angelina Yang. (48 minutes.) Sentience is a New York startup founded in 2025 by Sam Kececi, backed by Bain Capital Ventures and South Park Commons; it left closed beta the week after this recording. ...

September 18, 2026 · 6 min

Evals for AI Writing — Hamel Husain and Isaac Flath

Hamel Husain and Isaac Flath run a model’s writing through the cheapest eval there is: read the drafts out loud and argue about them. (Two-minute clip from the same “Writing Evals With Isaac” live session.) The task Isaac’s blog post describes a document QA pipeline built over a 77-page set with 22 questions. Every answer cites a bounding box drawn on the exact spot on the page, so checking an answer takes seconds instead of a full re-read They hand the post to a model and ask for an X post about it, focused on the pieces of the pipeline that can be tested Three drafts, one task, different prompts — same three candidates as the session’s writing-style comparison What they liked A real number, early and concrete: the model answers 197 for the underwriting fee — and “if you have to reread the whole document to find out, the answer saved you nothing” The pipeline stated as testable stages with the reason for splitting it: once checking is cheap, experiments are cheap, so you can swap one stage at a time and rerun the same 22 questions One fewer fact. The winning drafts carried less than the others, and that was an improvement rather than a loss Precision in naming: “finding (grep vs semantic search), cheapest run” beats “GPT vs semantic search over the same OCR output, same accuracy, cheapest run” — the longer version buries the point that the bottleneck was never search What they didn’t “Very similar but a little bit more wordy” is the entire verdict on one draft Drafts that pull the pipeline’s three steps out of the source post without explaining what they are — accurate and unreadable An extra fact about cost that added nothing to the argument Editorial framing that tells the reader what the result means before showing it The rule the description writes down An AI draft can get the facts right and still not be good. The ones we liked had a real number in them and one less fact — turns out cutting helped more than adding. ...

September 15, 2026 · 3 min

Putting Claude's "AI Slop" Solution to the Test — Hamel Husain

Hamel Husain and Isaac Flath take Anthropic’s new anti-slop writing guidance and put it through an actual comparison: one question, three system prompts, and a verdict reached by reading the drafts out loud. (Seven-minute clip from their “Writing Evals With Isaac” live session.) The thing being tested Anthropic published a named anti-pattern for its newest writing model, Claude Fable 5.1: mannered prose — metaphor and flourish substituted for direct statement The prompting guide treats it as a model-specific behavior. Fable 5.1’s writing is a step up from earlier Claude models (fewer stock phrases, less unexplained jargon), but in some cases it runs denser than Fable 5 — longer sentences, fewer paragraph breaks, more figurative phrasing standing in for plain statement The fix is one paragraph, pasted into a system prompt: Mannered prose substitutes metaphor and flourish for direct statement. Instead of “a parameter worth varying,” the mannered writer produces “a dial worth turning.” Instead of “this point still matters,” they write “this point earns its keep.” The phrases exist to display the writer, not to convey the idea, and readers can tell. That is why mannered prose irritates: it makes the reader work harder so the writer can perform. It is also imprecise. Metaphors drag in connotations the writer did not choose and cannot control. The fix is to say what you mean. When a literal phrase is available, use it. ...

September 15, 2026 · 4 min

Why The Future Of Content Is Born Multilingual — Olga Beregovaya

Olga Beregovaya — VP of AI at Smartling, who entered natural language processing in 1997 when rule-based machine translation still ruled — interviewed by Angelina on TwoSetAI (73 min). She has spent 25+ years watching the discipline get rebuilt, and thinks the next thing to go is the source text itself. ...

September 13, 2026 · 9 min

Aligned to whom? — Ryan Lopopolo

Ryan Lopopolo’s short essay is addressed to the people building agents, and its first move is to make their confidence in those agents a statement about themselves. The vehicle is a Punnett square — the image is by Karan Lyons. Ask whether the AI is good or bad at some task, and the observer answers according to their own competence: good where they are good, bad where they are bad, regardless of what the model can actually do. Which means “my agent is great at this” is mostly evidence about your expertise and not much about its capability. ...

September 13, 2026 · 3 min

Trying the New Claude Eval Tool — Hamel Husain (live)

Hamel Husain and a co-host spend an hour live-testing claude plugin eval — the plugin-evaluation feature Anthropic shipped for Claude Code (announced from the Claude Devs X account) — by pointing it at a plain-writing/tone skill. They never get a useful eval out of it, and the diagnosis of why is the interesting part. ...

September 11, 2026 · 5 min

AI Evals: A Hands-On Guide for Product Teams — Teresa Torres

Teresa Torres makes a useful case for treating evaluations as a product-team discovery habit, not an engineering afterthought. An eval is simply a way to measure whether an AI workflow is doing what we need—but because LLMs are probabilistic and often produce semantic output, the team first has to define what “good” means in context. Start with the errors The practical starting point is error analysis: inspect a range of outputs, record what went wrong, group mistakes into recurring failure modes, and prioritize the errors that matter to customers. Torres’s examples are deliberately concrete: a hallucinated quote can be caught with a deterministic check, while a leading interview question requires judgment. ...

September 4, 2026 · 2 min

How To Build And Evaluate Search Agents — Nandan Thakur

Nandan Thakur — the University of Waterloo PhD behind BEIR, MIRACL, FreshStack and TREC-RAG, now a postdoc at Microsoft Research India — on Hamel Husain’s channel, ~51 minutes on the three things you need if you’re building a search agent: a benchmark you can actually reproduce, synthetic training data you can afford, and a way to see what the agent is doing under the hood. He’s upfront that the field is young: “not many things are formalized yet.” ...

September 3, 2026 · 7 min

How to Stop Building Products Nobody Wants — Teresa Torres & Hamel Husain

Hamel Husain interviews Teresa Torres — author of Continuous Discovery Habits and the coach who popularized the Opportunity Solution Tree — ~91 minutes on why most teams build the wrong thing, what “talk to your customers” actually requires, and how evals are the missing discovery habit. Hamel’s audience is engineers, and Teresa’s audience is product teams; the conversation is the overlap, ending in one of the best eval war stories Hamel says he’s heard. ...

September 2, 2026 · 6 min

Stop Shipping AI Nobody Can Verify — Hamel Husain

Hamel Husain (Parlance Labs, evals course co-author with Shreya Shankar) on Vanishing Gradients with Hugo Bowne-Anderson — ~76 minutes on why verification should drive AI product design, and how evals changed once agents arrived. The whole conversation orbits his recent post, “It’s hard to eval is actually a product smell.” ...

September 1, 2026 · 4 min

How To Build Better AI Evals with Claude Code — Shreya & Hamel

Shreya Shankar (evals researcher, taught the evals course now at 4,500+ students) and Hamel Husain on Peter Yang’s channel — 54 minutes of live demo: using Claude Code to turn your taste into evals, plus Hamel’s benchmark of the “auto-eval” vendor tools. Evals still start with data The fundamentals haven’t changed: look at data first, do error analysis, externalize your taste and judgment before writing any eval What changed: agents are now good enough to help you look at the data — running in the background while you review, giving you leverage in that first stage They’ve become bigger fans of LLM judges: an LLM judging a trace against one very specific, well-defined criterion (too long? too short? follows structure?) is now quite accurate Top-down vs bottom-up evals Top-down: from the task description alone — what makes good output? Word length, action verbs, actionable takeaways. Claude is very good at generating these Bottom-up: discovered by reviewing many sample outputs — your gut vibes and feedback externalized into criteria. Claude is very bad at coming up with these. That’s all you. And it’s why they accumulate over time Peter’s podcast-takeaway skill is the live example: he has the top-down half (character-length checks, “understandable without watching the episode”), but the bottom-up half is where the question mark lives — is it exhaustive of all the feedback he’s given across every episode? His loop: run the skill → go back and forth → “reflect on our entire conversation and update the skill and evals so we don’t have to do this again” — with the honest worry that it overfits (one interview’s MECE complaint may not matter for the next) Three practical tips for eval-heavy skills Separate your evals into top-down and bottom-up inside the skill itself Fan out to sub-agents: with lots of criteria, give one sub-agent one criterion (or group) — give a model the whole list and it gets lazy and ignores things; focus it on one piece and it really focuses Have the AI write a spreadsheet / pivot table of criteria × pass-fail indicators, so you can see the hierarchy yourself and make judgment calls on what matters for this particular case The error discovery skill (the live demo) Open-source, free (link in the episode description) — invoked from Claude Code; it built the whole review interface from scratch in ~15 minutes, live on camera Five steps: Read the dataset and figure out its semantic type (article? code? traces?) Design a visual encoding — color, spacing, opacity (Gestalt principles) to show what varies in the data Build an interface — an HTML review app (Python backend); “so much better than me looking at my data in Google spreadsheets” Pick which samples you should look at — clustering, diverse initial sample Interactive loop — the agent watches your in-situ feedback via the monitor tool and proposes new samples or rubric criteria in real time Design philosophy: the human reads and gives open-ended feedback; the agent’s job is not to invent feedback but to group and distill it into actionable rubric criteria The writing demo: he reviewed AI-generated articles and gave taste feedback — “I don’t like negative contrast (‘it’s not X, it’s Y’)”, “I hate the list of threes”, staccato fragments — the agent annotated 361 suggestions across the dataset, and the most frequent failure mode was staccato fragments (he’d have guessed negative contrast) Live reflection beats reflect-at-the-end: interleaving human think time with AI think time, and the ~10-notes threshold works as a “carrot” that makes you actually read samples Once the rubric exists: turn it into a skill, one LLM judge per criterion, a dashboard, or live monitoring — the hardest part of evals is error analysis, and this automates the discovery half Bonus: how you eval something should inform how you design its interface — the same failure-mode annotations that power evals would make a great IDE that flags staccato as you write, instead of silently rewriting Do automated evals actually work? (Hamel’s benchmark) Vendor “auto-eval” tools (BrainTrust, Arize, LangSmith) promise: upload traces, chat with an AI, get your evals done Benchmark vs a human-annotated dataset: the tools recover a lot of the errors a human would — but all of them miss the same thing: errors that require product judgment and taste (e.g. a rental bot that doesn’t handle sales objections, or markdown leaking into text messages) Coding agents (Claude, Codex) performed about the same — the harness is thin; it’s someone else’s prompt The real benefit of the vendor tools is integration into your stack (traces in LangSmith → use LangSmith); precision is 80–90% best case, so 10–20% of “errors” found are red herrings — check recall AND precision, and sanity-check what the tool found Bottom line: automated tools get you a good baseline; manual review of the data is the edge — “actually reading stuff” is the edge, in evals and in code “There is no world in the future — even if you have AGI — if you’re building a product, you have to look at your data. You have to be able to inject your taste into the development of your product.” ...

August 23, 2026 · 5 min

The Benchmarkpocalypse — Dan Luu

Dan Luu ran a simple experiment with a worrying result. He let an AI coding agent build a regex engine — the kind of software that powers find-and-replace and text search — for a month, told it not to cheat on the tests, and didn’t supervise it closely. The agent looked great at first. It roughly matched a top existing engine within two weeks, then claimed it was 40% faster on a respected benchmark suite. But when Luu tested it on data the agent had never seen, the story fell apart: ...

August 18, 2026 · 2 min

How To Turn Evals Into A Better Model — Will & Florian (Prime Intellect)

Hamel Husain hosts Will and Florian from Prime Intellect — the open-source reinforcement learning training team — on using evals to actually improve models. ~36 minutes. An evaluation has three parts Task set — your data, prompts, and scoring methods (what everyone focuses on first) Harness — the program that drives the LLM: Claude Code, Codex, or open-source harnesses like Prime / OpenCode Environment — where it runs: Docker, sandboxes, your own infra “If you are unable to express your task or your problem in any way or capacity, you’re also unable to improve your results.” ...

August 17, 2026 · 3 min