Andreas Thom is a mathematician whose work with Gabor Kun supplied one of the techniques behind a recent OpenAI-assisted result on non-sofic groups. He and a colleague had spent months discussing related problems with ChatGPT. So when the result appeared, he emailed two OpenAI researchers to ask a direct question.
The question had two parts:
- Did our conversations end up in the training data — the material used to build later models?
- Were those conversations reachable by the system while it was solving the problem?
The complete answer he received was one sentence: “Regarding your conversations with ChatGPT: that did not happen.” Thom argues it addressed only the second half, categorically, with no evidence and no reference to his account’s settings.
The gap matters because OpenAI’s own language elsewhere draws the same distinction. In a separate case it said no specific user data was accessed, then added that it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.” De-identified means names stripped out — the conversations still helped, just without labels attached.
Other points from the thread:
- Thom turned off model training in his settings on 29 June. That control is prospective and unverifiable from outside: it says nothing about earlier conversations.
- Verifying any of this is not a research problem for outsiders. The evidence — privacy settings, which datasets, which model versions, what “de-identified” actually means — exists only inside OpenAI.
- His suspicion wasn’t procedural. The technique the model used wasn’t the most obvious route to the result, which is why he wondered how it got there in the first place.
- His ask is specific: if OpenAI gives a categorical denial, it should say what the denial rests on.
In a follow-up thread the next day, Thom widens the argument from one company’s conduct to what AI does to mathematics as a human institution. A machine-checked proof certificate confirms that a statement is true. It does not record who understood it, why the argument works, or where it belongs in the field — “comparable to the detection of a new star,” he writes.
Publication used to bundle three things at once: a record of knowledge, a basis for assigning credit, and evidence of a mathematician’s ability. AI-generated results break that bundle, and the credit cannot simply default to whoever ran the model.
His conclusion is a governance claim, not a technical one. Once intelligence becomes basic infrastructure for science, education, and public administration, a handful of companies should not be able to collect users’ ideas, control every piece of evidence about how those ideas were used, and exploit the asymmetry against their own users. It should be regulated like water or electricity.