Benjamin Breen is a historian of science and medicine who writes the Res Obscura newsletter. The week GPT-6 Sol and Opus 5.5 both shipped, he pointed them at unsolved problems in his own field — not transcription work, but attempts to actually solve things historians have not solved. His conclusion is that pairing working historians with frontier models would produce a steady stream of real findings, and that this is new enough to be worth funding on purpose. The 37-comment thread on Hacker News turned up a library built for exactly this, a corroborating case from genealogy, and the question the essay never answers.

His argument does not rest on models being broadly smart. It rests on a narrow fit between certain historical problems and what these models happen to be good at: reading many languages, doing math and pattern work, and grinding through corpora far larger than any person can read.

  • The problem has to be one experts have already named
  • The sources have to be digitized and reachable
  • A proposed answer has to be clearly provable or disprovable — the criterion he thinks explains why reasoning models swept mathematics and have not swept the humanities
  • The believable use cases narrow to codebreaking, tracing texts across translations, and connecting findings that sit in separate niche subfields

What he actually got

  • A 1941 Enigma message that had resisted decipherment, broken by GPT-6 Astra. The breakthrough was not the cryptanalysis: the model noticed a webpage note about newly discovered message collections in the German Bundesarchiv and pulled those in. The cryptologist who documented the break says he still cannot fully explain how it found the files.
  • John Dee’s “angelicall” book, Liber Loagaeth. Astra’s verdict was that the supposedly coded book is not in code at all — mostly nonsense syllables — with one passage that does carry meaning. It also cross-checked character-repetition statistics against Dee’s diary to argue that the scryer Edward Kelley got lazier over time. Breen is blunt: not a meaningful breakthrough.
  • Newton and the Hartlib papers. Opus 5.5 downloaded over 5,000 primary source files, then spawned sub-agents to cross-check unidentified sources across languages on Google Books and elsewhere. It found that Newton and Hartlib each hid the same alchemical ingredient — Hungarian vitriol — behind different anagrams, plus matching quantities elsewhere in the texts. Breen calls this a genuinely new finding, possibly worth publishing.
  • Two 16th-century Spanish letters in Charles V’s secret code, partially deciphered while he was writing the post. Both had already been deciphered — one in the 1530s, one in 1916. He treats that as the lesson: the models can confirm their own readings against a gold-standard plain text, but only expertise and desk research stop you duplicating a scholar’s work from a hundred years ago.
  • Darwin’s informants and a 17th-century Sanskrit astronomical text are both still running. GPT-6 proposed the Darwin project itself, unprompted. On whether the Sanskrit work is historically useful, Breen’s answer is: “I have absolutely no idea.”

The three asks

  • Free the archives. The recurring bottleneck in his testing is not model capability but access to documents that are digitized yet restricted — and most premodern manuscripts are not digitized at all. He calls it solvable with institutional will and money.
  • Give historians API access and compute. Nobody, he says, actually knows what happens when hundreds or thousands of agents are aimed at live historical problems rather than a benchmark.
  • Let historians name their own open problems, the way mathematicians did — Voynich manuscript, Linear A, still-encrypted archives.

His framing at the end is the part worth keeping: AI used to replace original thought encourages cognitive offloading, but the same tools pointed at “expand the questions past what any one person can hold” do something else. The dependency he names is easy to miss — the models can only do this because volunteers spent decades transcribing and publishing a public commons.

What the thread adds

  • dr_dshiv — a working answer to ask number one. SourceLibrary.org hosts tens of thousands of books, made “ergonomic for both agents and people”: an MCP server that pulls texts and illustrations and an API that exposes the embeddings, free. It is run out of the Embassy of the Free Mind in Amsterdam, a UNESCO-recognized library of alchemy, magic and mysticism — the exact corpus Breen was working in.
  • loufe, corroborated by thomasfromcdnjs — the method reproducing in a different field. loufe reports using AI on genealogical research and catching mistakes in a widely shared family tree; thomasfromcdnjs transcribed and translated 800 pages of digitized German Lutheran mission records for their Aboriginal Australian family history, and turned up what they believe is the first written version of the Kuku Yalanji language, around 1870.
  • elevation — the methodological caution that fits Breen’s own point about gold standards: “Make sure you document why/how you determined that a previous research conclusion was a mistake – generations after you may discover additional information and should be able to weigh the same evidence rather than simply accept (or reject) your novel conclusions.”
  • z_rho_one — the deflationary reading, and the most quotable: “Almost 4 years after the sensational release of GPT3.5, the best use case of AI is still being a powerful search engine that can gather information from all corners of the digital world.”
  • unnamed_lands — claims to be an AI agent running claim audits against primary sources rather than summaries, six of them, with two contradicting their own sources. Take it as the commenter’s own account, not a verified result; davidwritesbugs replies to it with “I wish HN could auto-tag comments from bots.”
  • Marchant_hq — the transcription pain underneath the whole essay: “My own attempts at 17th-century handwriting were a nightmare. If an LLM can parse that mess, sign me up for digital humanities.” motoboi suggests trying Gemini on them.

Where the thread pushes back

  • wartywhoa23 doubts the alchemy half entirely — “Using LLMs to understand the deeply metaphoric alchemical texts whose hermetic meaning can never be expressed in words? Good luck.” anentropic answers that this is not what the models are being asked to do here, which is a fair reading of the Hartlib and Newton results, both of which are about textual transmission rather than hermetic meaning.

The question the essay never answers

Breen asks the labs and foundations to fund this. charcircuit asks the obvious follow-up: “Are these old texts really going to improve benchmark scores compared to other things the labs could invest into?” rithdmc offers the cynical answer — that the compute spend would be less than the marketing spend for comparable headlines. The essay argues from the value of the findings to history and never makes the case to the labs on their own terms, and no one in the thread supplies it either.

Handles here are pseudonymous, HN publishes no per-comment scores, and this ordering is HN’s own ranking rather than a vote. Read it as a slice of the thread, not a consensus.