Isaac Flath walks Hamel Husain through a 42-minute hands-on evaluation of OCR and document-question-answering agents. His key move: inspect the document, extracted text, retrieval, citations, and agent trace together.
Why a wrong answer may not be a model failure
- A misread digit in a financial table can look like an LLM hallucination; better prompting cannot correct corrupted input.
- Documents mix scanned images, text layers, tables, watermarks, and handwriting. Medical, legal, and financial workflows can turn small extraction errors into consequential answers.
- Judge errors by the task: a shifted underline may be cosmetic; a missing inspection score or draft watermark may change the meaning.
Inspect real documents, not just benchmark scores
- On an ordinary food-inspection form, a heading spans rows inconsistently and an OCR model drops a point value. Even the presenters debate the correct grouping.
- Compare the PDF, raw OCR, and rendered output side by side. Attach each observation to both a page bounding box and the corresponding extracted-text span.
- A different model invents checkboxes, converting ambiguous fields into binary ones; both OCR systems miss an “example” watermark in a medical record. A missed “draft,” “confidential,” or “void” stamp could be much worse.
- Flath built his annotation UI incrementally, starting in about an hour and adding rendered views, boxes, and hover-linked text whenever manual review became painful.
Trace an agent failure backwards
- Check the final answer and whether its citation actually supports it.
- Inspect the agent trace: was the wrong number already in its context?
- Check retrieval: did it fetch the wrong page or the right page with wrong text?
- Compare extracted text with the original PDF. If OCR broke it, test and fix that layer, not the prompt.
- Add a cheap targeted check when possible, such as verifying table entries add up to the stated total.
Evals across the pipeline
- Test OCR, retrieval, and the agent loop separately; swapping an agent harness will not repair a faulty source representation.
- A citation’s presence is weaker than checking whether the cited image/text supports the answer. Visual, page-level citations make human verification practical.
- Tool-using agents can crop and re-OCR pages until they find an answer, but sometimes take 45 minutes and spend $20, even when the answer is absent.
- “Blank” versus “zero” in a loan estimate, or a faint signature on a contract, requires domain judgment; not every discrepancy admits an automatic yes/no label.
“Maybe 80% of my failures are actually just because my OCR model failed, in which case prompt tuning is just a waste of my time right now.” — Isaac Flath (hypothetical example, not a measured result)