The text a reasoning model produces before its final answer is often called a “chain of thought.” Subbarao Kambhampati and eight coauthors argue that this language smuggles in an unsupported claim: because the text resembles human scratch work, people assume it faithfully reports what the model is doing.

The paper does not deny that intermediate tokens improve performance. It separates that established result from two much weaker assumptions:

  • A correct-looking trace does not guarantee a correct answer, and a correct answer does not guarantee a valid trace.
  • Models trained on incorrect, truncated or even instance-swapped traces can retain or improve answer accuracy.
  • Longer traces are not necessarily evidence of harder thinking; token length can reflect training incentives or a poor fit with the task distribution.
  • Showing users a trace or a summary can increase trust regardless of whether the answer is correct.

The authors suggest a less anthropomorphic explanation: intermediate tokens may act as a learned prompt augmentation, reshaping the model’s context so the final answer becomes more likely. That scaffold can help the model without being a human-readable derivation. This remains a hypothesis, but it fits the intervention evidence better than treating fluent prose as a transparent window into hidden cognition.

The distinction becomes especially useful for agents. Private intermediate text may lack reliable semantics, while tool calls are externally meaningful commitments that can change files, systems or the physical world. Audit the actions, evidence, verifier results and final output—not the model’s persuasive story about how it got there. The paper’s practical recommendation is verification-first: trust should come from checking the result, not from how thoughtful the preamble sounds.

One citation warning: the current ICML 2026 version is a substantial fork of the original 2025 arXiv submission, which had a different title, author list and broader scope. Cite the version you actually read.