Cryptomnesia: OpenAI grapples with the lost source

The recognition of a discovery is the foundation of modern science, but some recent cases suggest that chatbots could make it more difficult to attribute authorship of a work

11 SEP 26
Translated by AI
Image of Cryptomnesia: OpenAI grapples with the lost source

Photo: ANSA

On Tuesday 8 September, OpenAI announced that it had solved, using one of its in-house models, the Navier–Stokes problem – one of the Clay Institute’s seven Millennium Problems – demonstrating that the equations describing the motion of fluids, when a regular external force is applied to the fluid, can develop a singularity within a finite time; the Clay Institute, for the time being, continues to list the problem as open. Twelve hours earlier, Tristan Buckmaster of New York University and Levent Alpöge, a mathematician employed by Anthropic who was collaborating with him in a personal capacity, had published a similar result concerning the Euler equations, the frictionless version of the Navier–Stokes equations, obtained on 15 August whilst working primarily with Codex, OpenAI’s programming tool, within which, Buckmaster explains, they kept all their drafts. OpenAI confirms that it began its own work on 1 September, after hearing a rumour that it only later traced back to the two; and in response to the question about training – which remained unanswered during its contacts with Buckmaster – it replied in writing with a statement that deserves careful reading: neither the researchers nor the agents had seen that work prior to publication, and no specific user data was accessed, yet the company, whilst considering it unlikely, cannot rule out the possibility that de-identified data, derived from the two mathematicians’ use of its products, may have contributed to improving its models, and adds that the demonstrations differ significantly.
A few weeks earlier, Andreas Thom, from Dresden University of Technology, had asked a similar question after the Astra model had constructed the first non-sofic group with a decisive step based on a 2019 article he had written with Gábor Kun: Thom had been discussing precisely that issue with ChatGPT for months, and the response he received from an OpenAI researcher – ‘that did not happen’ – as he reports it, ruled out direct access to the conversations and remained silent on the training. Neither Buckmaster nor Thom claim that the transfer took place, yet both have put their finger on an issue that goes far beyond their individual cases.
In August, a few days after Astra’s announcement, Francesco Fournier-Facio, in Cambridge, realised that OpenAI’s press release had omitted the articles on which that result was based, to the extent that the company had amended the text: the sources had been published, and a specialist was able to identify them within a few days. When the potential source is a private conversation, however, the same process becomes impractical for anyone outside the company and can prove extremely difficult even for the company itself, because in a language model, training modifies billions of parameters and what the system has learnt typically remains a capability without a verifiable origin; indeed, determining which fragment of data influenced a particular response remains an open research problem. OpenAI can verify whether certain conversations have entered its dataset, whilst it would be far more difficult to reconstruct the path that would have transformed them into an insight generated by the machine.
Psychologists have a name for this sort of bewilderment. In 1902, Carl Gustav Jung noted that a passage from Thus Spoke Zarathustra mirrored almost word for word a page by Justinus Kerner that Nietzsche had read as a boy, and he called the phenomenon ‘cryptomnesia’, which cognitive psychology now describes as a source monitoring error, in which the content survives whilst the label indicating its origin is detached. A model trained on its users’ conversations could produce a functionally equivalent phenomenon on an industrial scale, in which no one would open a confidential file, and yet an idea confided to a chatbot might resurface some time later as a suggestion from the machine, without the researcher receiving it, the one publishing it, or the company advertising it knowing where it came from. The de-identification promised by the terms of use leaves the problem intact, because it removes the name whilst retaining the idea – precisely what is of value to a scientist.
The loss of traceability carries such weight because the entire architecture of modern science rests on the principle of priority. Robert K. Merton explained this in 1957: results become part of the public domain the moment they are published, and the scientist’s sole remaining claim to ownership is recognition – the name associated with a theorem or a discovery – so that the rule rewarding the first to arrive only works as long as it is possible to establish who arrived first. This is why, over the centuries, science has devised methods of dating, from the Latin anagram with which Galileo secured priority over the phases of Venus in 1610 to the sealed envelopes that the Académie des sciences in Paris has been receiving since the late 17th century. The chatbot to which a researcher confides an unfinished idea is like a sealed envelope delivered to a recipient who can open it, who issues no receipt, and who, in the meantime, is competing on the very same problems – in a sector where the announcement ‘the machine has solved it’ has become the currency by which a company’s value is measured; and the risk affects the entire sector, albeit with rules that vary from one company to another.
The consequences for the integrity of research stem from this. The European Code of Conduct drawn up by ALLEA, which the Commission recognises as a benchmark for EU-funded projects, defines plagiarism as the use of another person’s work or ideas without due acknowledgement, and the European guidelines on generative artificial intelligence reiterate that researchers remain fully responsible for what they publish. Everything therefore hinges on the author, that is, the link in the chain who has no way of seeing the source: the honest researcher who accepts a suggestion from the model has a duty to acknowledge any hidden contribution by others, yet has no means of knowing whether such a contribution exists; consequently, they might publish, in perfectly good faith, an idea that originated in a colleague’s chat, whilst the entity controlling the channel may be a private company over which these self-regulatory codes do not necessarily exert any binding force.
This leads to the most reasonable request, which is in fact the one put forward by Thom: the onus to shed light on the matter lies with those who hold the data. Companies that claim to have made discoveries can be asked to carry out an audit of the data’s origin, a process that is simple in principle, albeit technically challenging, to verify – using available methods – whether the key ideas appear in user data from previous months and to publicly account for how the verification was carried out; whilst researchers would be well advised to disable training on their own data and perhaps return to a modern version of the sealed envelope, depositing in an archive with a certified date the material they are about to discuss with a machine. In 1976, a New York court ruled that George Harrison had reproduced the melody of ‘He’s So Fine’ by The Chiffons in ‘My Sweet Lord’ and, whilst acknowledging that he had done so unwittingly, held him liable nonetheless, establishing a principle that science has always practised with regard to citations: the copier’s lack of awareness may exonerate their intentions but leaves the debt to the author intact. This should apply all the more so to models that retain what they have learnt even when the user subsequently cuts off the flow of their own data, and which, just like Nietzsche, retain no memory of the source.