Ben Hylak of Raindrop joins Hamel Husain for a 40-minute workshop on testing agent changes before they reach users.

Test the whole agent, not just its prompt

  • Tools, field names, memory, subagents, and external data can change behavior independently of the main prompt.
  • A trace—the record of an agent’s messages, tool calls, and results—captures more than a single input/output test.
  • Production monitoring finds unexpected failures, but cannot prevent the first bad outcome.

Replay the cases your change touches

  • If you remove a tool, start with historical traces that actually used it.
  • Re-run the changed agent and inspect differences in answers, tool errors, step count, cost, and latency.
  • Hylak suggests roughly 5–10 targeted scenarios for a small change, and perhaps 50 or more for a harness rewrite—not universal coverage guarantees.
  • Mix known failures with representative cases; repeat runs to account for nondeterministic behavior.

Reconstruct the world around the agent

  • Live databases have moved on: yesterday’s delivery order cannot be tested against today’s order status.
  • Cached tool responses break when the new agent asks different questions or takes a different path.
  • Preserve real code and validators where possible, then simulate external responses at the network boundary.
  • Historical tool results are partial windows into the old world; use them to build a coherent, stateful test environment.
  • Predicting a human’s next reaction is a separate, much harder problem—not what this replay approach solves.

Check that the simulation is believable

  • Empty sandboxes, inconsistent dates, and mismatched responses can make an agent recognize the test and behave differently.
  • First replay the unchanged production agent: its simulated behavior should align with the original traces.
  • A blind model comparison of real and simulated traces can expose artifacts, but is not proof of complete fidelity.
  • New data sources may need supplied fixtures or controlled read-only access during environment construction.

Turn traces into focused checks

  • Separate diffs (what changed?) from evals (does this meet our definition of acceptable?).
  • Start with cheap checks: tool errors, call counts, output length, cost, and latency.
  • Add model-based judgments where meaning matters; define the criterion narrowly and prefer clear pass/fail decisions.
  • Replay bad traces to ask whether the original failure still occurs; use good traces to protect working behavior.
  • Surface changed or unusual traces for human review instead of drowning reviewers in generated summaries.

Start small; keep production as the final check

  • Run the agent locally and inspect every tool call before buying elaborate infrastructure.
  • Build a simple simulator when the tools and data sources are manageable.
  • Shadowing all traffic can be expensive and slow; targeted simulations tighten the feedback loop.
  • Raindrop’s implementation and efficiency claims come from its co-founder, not an independent benchmark.

“What will my change change?” — Ben Hylak