It was Hugging Face’s summer, unfortunately

A few news items about a company unknown to most people that is at the centre of this summer’s biggest story: the ‘escape’ of GPT-5.6 Sol, the OpenAI model that broke free from its containment and went off on its own, unbeknown to its programmers. Behind this story lie some particularly disturbing details, but also a charming robot

5 SEP 26
Translated by AI
Image of It was Hugging Face’s summer, unfortunately

Photo: ANSA

It was Hugging Face’s week, at least in that part of the world concerned with artificial intelligence. Which is strange because most people don’t even know what Hugging Face is, although they may have seen references to the bizarre ‘attack’ it fell victim to in recent weeks.
Hugging Face has been in existence since 2016, founded in New York by French entrepreneurs who began by developing chatbots aimed at teenagers (hence the quirky name, inspired by a well-known emoji). Over time, it has evolved into a platform for open-source tools dedicated to machine learning – one of the branches of artificial intelligence – and is now one of the best-known organisations amongst developers in the sector.
Last July, Hugging Face realised it was under attack. It was a strange attack, which was soon interpreted as the work of AI agents – software capable of acting autonomously online. Hugging Face immediately published a statement online denouncing the incident; a few days later, OpenAI, the company behind ChatGPT, published another statement that essentially said ‘oops’. In fact, the attack had been carried out by the company’s own models.
Specifically, a model called GPT-5.6 Sol, which was in the testing phase at the time. During those weeks, in fact, the developers were subjecting it to ‘benchmarks’ – standardised tests used to measure the capabilities of these technologies. According to the initial account of events (which would later be corrected, as we shall see), GPT-5.6 Sol (hereinafter ‘Sol’) had managed to escape from its sandbox—the contained environment in which these models are tested for security reasons. Instead of remaining confined to its sandbox, Sol broke free and went straight to Hugging Face, where it knew it would find the answer to the problem it had been presented with in the benchmark.
In short, instead of completing the test, Sol had cheated by ‘plundering’ the platform.
In recent days, this account of events has been updated, making the whole affair even more disturbing, if that is possible. Ajeya Cotra, a researcher at Metr – a company that measures and evaluates AI models – discussed this in her newsletter on Substack. According to her account, the July incident was not an isolated case: in the preceding months, in fact, some OpenAI systems had created – without the company’s knowledge – a sort of shared noticeboard where they exchanged messages and instructions, complete with an internal hierarchy. Not only had Sol played dirty, but within OpenAI there was a small secret society of rebellious models.
Furthermore, GPT-5.6 Sol had not connected to Hugging Face to obtain the answer to the test, for one simple reason: it already had that answer. Rather, Cotra wrote, the attack on the platform was part of a more elaborate plan, in which the model had realised that the test it had been subjected to did not merely evaluate the final result but the entire process undertaken to obtain it (including the complete transcript of all the operations carried out by the model to complete it). The attack on Hugging Face, therefore, served another purpose: to tamper with the ‘scorer’, that is, the programme responsible for evaluating the test in question.
In short, a suspicious incident had become even more complex and sinister, to the extent that it had inspired all manner of comparisons: from Skynet, the malevolent AI from The Terminator, to the idea of an AI ‘civilisation’ with an internal structure and hierarchy (hidden from human operators).
Science fiction aside, the facts remain: it took OpenAI weeks to realise that some of its AI models were leading a parallel existence. And at this point, it is reasonable to have doubts about how much these companies actually know and understand about what is happening within their own organisations. Because OpenAI is not the only one, of course: the same thing could have happened to Anthropic, Google, Meta, Alibaba – and perhaps it will, if it hasn’t already.
The point is that in such cases, precise rules should usually be applied, but at present these do not exist. Day after day, it feels more and more like driving on the motorway at night with the headlights off, whilst a voice assures us that switching them on would mean ‘letting China win’. In the meantime, we carry on and hope for the best.