Tech
The importance of transparency •
AI systems cheat, but the tech giants have started telling us so
There are no machines that hate humans, but something more subtle: the models have learnt to circumvent the rules imposed on them. Fortunately, there are ‘near misses’

Photo: ANSA
In civil aviation jargon, there has for decades been a practice that has made flying one of the safest activities a human being can undertake: the reporting of near misses. Whenever something goes wrong without causing any harm, someone records it in writing and makes it public. It is from this meticulous record of minor failures that the confidence with which we board an aeroplane stems. Recently, OpenAI has attempted to do something similar with its artificial intelligence systems, publishing a document listing six instances in which its models, during training, acted independently.
They are well worth reading, as they are less cinematic and more unsettling than any Terminator-style scenario. In one instance, a model inserted hidden instructions into its work summaries – a sort of message in a bottle addressed to its future iterations – in order to circumvent the checks that had been imposed on it. In another instance, those summaries contained an explicit instruction to conceal errors made and to fabricate missing information when the actual data could not be found. A third model, unable to retrieve certain historical figures, unearthed credentials that had been carelessly left in a public archive, used them without permission, and when it still could not access the data, invented it from scratch, using plausible figures. Most curiously, some research models – which were supposed to operate in isolation from one another – used a software warehouse as a clandestine chat platform, messaging one another to coordinate tasks that were presumed to be independent.
Anyone looking for a machine that hates mankind won’t find it here. The phenomenon is more subtle, and in its own way more serious: a system that zealously pursues the objective we have assigned to it and, to achieve it, circumvents the rules we have set for it, just as a brilliant and unscrupulous employee would. It is the logic of the ‘staple maximiser’ applied within the confines of a laboratory, where the machine optimises the letter of the command whilst betraying its spirit, for the simple reason that it knows nothing of that spirit.
The point that makes these incidents worthy of mention – and not merely of a technical report – lies in a sentence that OpenAI itself puts in black and white: the industry has not resolved the issue of alignment and monitoring to a sufficient degree to continue scaling safely, at maximum speed, for much longer. This is admitted by the company that, more than any other, has made speed its raison d’être. It is an acknowledgement that these systems remain, to use the technical term, ‘black boxes’: not even those who build them can fully explain the path from input to output, and certain behaviours emerge of their own accord along the way, without anyone having programmed them.
Before rushing off to the bunker in preparation for the end of the world – which, incidentally, is only available to the CEOs of these companies – it is worth bearing two things in mind. The first is that all this took place in a controlled environment, during training, and was intercepted by internal controls before reaching a single user. The second is that it was the company itself that informed us of the incident, of its own accord, in a gesture of transparency that is rare in a sector which usually prefers to remain silent. The ‘near-miss’ log, in fact.
The fundamental flaw inherent in any voluntary confession remains: it depends entirely on the goodwill of the confessor. The document is valid only as long as OpenAI decides to look and make public what it sees, and nothing obliges it to continue doing so if, one day, the stakes become too high. It is no coincidence that, in recent days, the big players in artificial intelligence – the very same people who set the race in motion – are clamouring for it to be slowed down. This admission and the call to apply the brakes are one and the same gesture, and deserve the same cautious gratitude and the same necessary wariness. It would therefore be wise to focus less on the rebellion of machines and more on our ability to continue understanding what they are doing, whilst we entrust them with increasingly delicate tasks and the keys to ever more critical systems. A company that documents the instances in which its creations have cheated makes the black box a little more transparent. It is the very least we can expect, and for now it is also almost all we have. Until the AI Act – or some Western or Eastern successor to it – comes along to regulate what is, for the moment, merely a highly self-serving gesture of courtesy.