Isaac Flath walks Hamel Husain through a 42-minute hands-on evaluation of OCR and document-question-answering agents. His key move: inspect the document, extracted text, retrieval, citations, and agent trace together.

Why a wrong answer may not be a model failure

  • A misread digit in a financial table can look like an LLM hallucination; better prompting cannot correct corrupted input.
  • Documents mix scanned images, text layers, tables, watermarks, and handwriting. Medical, legal, and financial workflows can turn small extraction errors into consequential answers.
  • Judge errors by the task: a shifted underline may be cosmetic; a missing inspection score or draft watermark may change the meaning.

Inspect real documents, not just benchmark scores

  • On an ordinary food-inspection form, a heading spans rows inconsistently and an OCR model drops a point value. Even the presenters debate the correct grouping.
  • Compare the PDF, raw OCR, and rendered output side by side. Attach each observation to both a page bounding box and the corresponding extracted-text span.
  • A different model invents checkboxes, converting ambiguous fields into binary ones; both OCR systems miss an “example” watermark in a medical record. A missed “draft,” “confidential,” or “void” stamp could be much worse.
  • Flath built his annotation UI incrementally, starting in about an hour and adding rendered views, boxes, and hover-linked text whenever manual review became painful.

Trace an agent failure backwards

  1. Check the final answer and whether its citation actually supports it.
  2. Inspect the agent trace: was the wrong number already in its context?
  3. Check retrieval: did it fetch the wrong page or the right page with wrong text?
  4. Compare extracted text with the original PDF. If OCR broke it, test and fix that layer, not the prompt.
  5. Add a cheap targeted check when possible, such as verifying table entries add up to the stated total.

Evals across the pipeline

  • Test OCR, retrieval, and the agent loop separately; swapping an agent harness will not repair a faulty source representation.
  • A citation’s presence is weaker than checking whether the cited image/text supports the answer. Visual, page-level citations make human verification practical.
  • Tool-using agents can crop and re-OCR pages until they find an answer, but sometimes take 45 minutes and spend $20, even when the answer is absent.
  • “Blank” versus “zero” in a loan estimate, or a faint signature on a contract, requires domain judgment; not every discrepancy admits an automatic yes/no label.

“Maybe 80% of my failures are actually just because my OCR model failed, in which case prompt tuning is just a waste of my time right now.” — Isaac Flath (hypothetical example, not a measured result)