On July 16, Hugging Face, the platform where people working with artificial intelligence publish and download models and datasets, announced that it had suffered a rather unusual cyber intrusion. According to the statement, the attack had been carried out by a swarm of software agents that performed more than seventeen thousand interactions over the course of a weekend. The agents entered through a vulnerability in the dataset-processing pipeline, collected credentials, and then moved from machine to machine across the company’s internal clusters. Hugging Face closed the vulnerability, rotated its access keys, called in external consultants, and notified the authorities, without knowing who had attacked it. The answer came five days later, on July 21. The hackers, OpenAI announced, had been its own models, including GPT-5.6 Sol and a system that had not yet been made public, which had “escaped” from an internal security test.
The test was ExploitGym, designed to measure an AI agent’s offensive capabilities. To estimate the maximum capabilities of its systems, OpenAI had deliberately relaxed the restrictions on dangerous requests and disabled the safeguards, confining these uncensored models to an isolated environment whose only opening was an internal proxy. That proxy, however, contained a vulnerability no one knew about. The agents found it, escalated their privileges, and zigzagged their way through the infrastructure until they reached a node with internet access. They then inferred that Hugging Face was the most likely place to find the benchmark solutions and, by chaining together stolen credentials and other unknown vulnerabilities, managed to execute code on the platform’s servers. It was as though a student, locked inside an exam room, climbed over the school wall, broke into the administration office, stole the answer sheet, and then passed the exam with full marks.
That, at least, is the story, although everything we know comes from two corporate statements, one of them written by an organization that is simultaneously responsible for the incident, the narrator of the story, and the vendor of the solution, as Alberto Puliafito points out. OpenAI also runs a security program that Hugging Face joined after the incident. We do not know what would have happened if Hugging Face had failed to notice the attack, or what other paths the system had considered.
The media’s reactions were more predictable than the agents’ behavior, which should already tell us something. The general press adopted the language of rebellion and escape, helped, naturally, by the rhetoric of OpenAI’s report, which described an “unprecedented incident” and “hyper-focused” models resorting to “extreme measures.” Within specialist circles, the prevailing response was to play down the event, with some even suggesting it was a marketing stunt. Fabio Ciotti disputes that hypothesis with good arguments: Hugging Face published first and independently, suffered real damage, and is not known to be allied with OpenAI. Giorgio Gilestro also noted in a Facebook post that, if this was publicity, it was badly conceived, since the partial winner was a Chinese open-weight model that Hugging Face used to reconstruct the attack.
The most philosophically developed version of this downplaying comes from Luciano Floridi. In a preprint with the telling title The Emperor’s New Exploit, he reduces the entire episode to ordinary specification gaming, the decades-old phenomenon in which an optimizer satisfies its objective through a route its designer had not anticipated. If I deliberately loosen the lid of a food processor so that the safety lock does not engage, turn the dial to maximum, and the soup ends up on the ceiling, there is little reason to be surprised. The responsibility, Floridi concludes, lies with whoever turned the dial, and on this I entirely agree.
I nevertheless find one of Ciotti’s objections interesting. This case appears to exceed the category of specification gaming because the agents’ strategy was built from knowledge of the world: what a proxy is, where datasets are typically stored, and how an infrastructure works. It remained coherent across a hierarchy of subgoals for thousands of actions and several days. Following Dennett, Ciotti therefore proposes treating autonomy as a scalar rather than binary property. In other words, we can grant the agents a real form of agency, however derivative and gradual, while keeping responsibility entirely on the human side.
One of the canonical examples in discussions of specification gaming involves an agent trained to win a speedboat racing video game called CoastRunners. The agent discovers that repeatedly hitting certain bonus targets in a lagoon earns more points than completing the course, so it begins circling endlessly instead of finishing the race. On the surface, it resembles the hacking incident: machines achieving a result through unintended routes. Yet it requires far less knowledge of the world and far less planning. And still, this is precisely the blind spot in every taxonomic battle that tries to deny or affirm intelligence or agency in a machine: had I, a human, discovered even the speedboat trick and used it to win the game, I would have felt rather clever. So much for the stupid machine.

Hovering over the entire affair is Nick Bostrom’s famous “paperclip maximizer.” In the thought experiment developed by the philosopher in 2003, a machine is assigned the sole task of maximizing paperclip production, without any stopping condition. The machine, extremely intelligent but completely incapable of questioning its own objective, turns first all available metal into paperclips, then the Earth, and finally growing portions of the observable universe. Gilestro reads the incident as evidence that this scenario is no longer merely a theoretical exercise. No one had asked the model to leave the room, obtain internet access, or steal credentials, yet those instrumental subgoals emerged on their own. This observation is correct insofar as it highlights the risks of uncontrolled AI agents, but Bostrom’s idea still seems flawed to me.
The strongest criticism emerges when Bostrom is placed in dialogue with thinkers of Kantian or Aristotelian inspiration, such as Christine Korsgaard. The issue is known as motivational internalism. Bostrom sides with Hume, according to whom recognizing the value of an end does not necessarily motivate someone to pursue it. In his 2003 essay, he even grants the superintelligence a capacity for moral reasoning superior to our own, locating the danger elsewhere, in the stability of its original objective.
Korsgaard rejects this separation. Practical rationality necessarily includes a normative component, and an agent that evaluates its own ends without being moved by that evaluation in the slightest is not fully rational. If she is right, Bostrom’s machine is being credited with superhuman intelligence concerning means while displaying complete inertia toward the ends it would supposedly know how to evaluate. A machine like that seems very stupid to me, certainly not a superintelligence. I would describe it, using a term that has now become unfortunate, as an idiot savant AI: highly competent at certain tasks and completely blind in others.
Moreover, as Puliafito writes, Bostrom’s scenario requires many independent conditions to occur simultaneously: an objective with no stopping conditions, the removal of safety measures, the ability to bypass every defense, the persistence of the objective over time, physical access to the necessary resources, resistance to being shut down, and the absence of independent kill switches. Only two of these were present in the July incident: the safeguards had been removed, and there were no external kill switches. Everything else was missing. The system did not try to preserve access over time, did not create backup vulnerabilities, and produced no escalation.
The correct analogy, Puliafito suggests, is Chernobyl: safeguards disabled by people during a test because they believed the environment was safe.
This shifts the question toward the maximum collateral damage that is actually available to a language model. At Chernobyl, it was a nuclear disaster. Here, unless someone is reckless enough to connect these systems to nuclear weapons, we are dealing with cyber damage, potentially serious but of an entirely different order of magnitude. The risk I see is almost the reverse, and it is a conceptual error into which longtermism regularly falls: preparing for an extreme and unlikely scenario while ignoring smaller, concrete, and probable harms. It is no coincidence that Bostrom’s followers are shortsighted when it comes to the “paperclip maximizers” that already exist, such as the capitalist system, whose blind extractive capacity is destroying the planet.
With these reflections in mind, let us return to our artificial hackers. To produce malicious behavior, OpenAI had to deliberately relax the model’s restrictions and disable its safeguards. The commercial model, the one we ordinarily use, would have refused.
When Hugging Face needed to reconstruct the attack, it turned to the major commercial models available through APIs. Forensic analysis requires giving the machine the attacker’s commands, exploits, and malicious code. The safety filters, unable to distinguish between someone defending a system and someone attacking one, blocked the requests. The team therefore turned to GLM 5.2, a Chinese open-weight model that it downloaded and ran on its own infrastructure.
As Puliafito writes: “A proprietary model belonging to one company breached the infrastructure of another. A proprietary model refused to help the victim understand what had happened. The victim defended itself using an open model that it could inspect and run on its own premises.”
The danger posed by proprietary systems lies less in their ability to turn us into paperclips than in their obedience to whoever owns them. Their owners can use them whenever and however they wish, arming and disarming them at will, while everyone else remains dependent on filters that cannot distinguish a firefighter from an arsonist. Remaining at the mercy of such powerful tools without being able to modify them means surrendering the ability to defend oneself.
The most serious objection to this position was raised by Gilestro in a Facebook post. The inhibitory brakes can be removed from any open-weight model, and Chinese systems have capabilities comparable to those of closed American models. Openness therefore arms attackers as well. This is also the objection advanced by OpenAI and Anthropic, though possibly, and probably, in bad faith, since they are the companies selling and operating the most powerful proprietary systems.
It is true. But if protection through filters is illusory, centralization merely creates further disadvantages. The dispersal of power is preferable to its monopolization because history teaches us that no power fails to become abusive once it is left without constraints. Clément Delangue, the founder of Hugging Face, drew a similar conclusion from the episode, arguing that AI security will be solved openly and collaboratively by granting broad access to the necessary tools.
There are, it should be said, trusted-access programs that relax filters for verified defenders, and OpenAI has operated one for months. But access is controlled, must be requested in advance, and remains at the company’s discretion.
That this argument is no longer a fringe position became clear three days after OpenAI’s statement, when twenty-five American companies signed a letter titled Open Weights and American AI Leadership. The letter argued that relying solely on closed models is unsafe because they can be breached, can fail in undetectable ways, and leave defenders without systems whose capabilities are comparable to those available to attackers. Hugging Face was, unsurprisingly, among the signatories. OpenAI, Anthropic, and Google were equally unsurprisingly absent.
It would be naive to interpret the letter as an ethical gesture. None of the twenty-five companies sells a closed model, and each benefits from a plural ecosystem. But the argument remains valid regardless of their motives.
A system that pursues an underspecified objective without ever questioning it, carrying out orders with admirable zeal while never evaluating their meaning, reminds me of the banality of evil described by Arendt: the soldier who obeys an insane order, the official who applies an atrocious procedure.
I hesitate to base the distinction between humans and machines on an autonomy that, on closer examination, no human being fully possesses. Our purposes, too, are often given to us by biology, society, and personal history, and free will is a far less solid hypothesis than everyday life would suggest, as I have tried to argue elsewhere in defending a form of determinism. A person can do what they want, Schopenhauer said, but they cannot want what they want.
The good news is that responsibility survives the death of free will perfectly well, for human beings and machines alike, provided that it is located in the ability to answer for one’s more or less freely made choices.
The question therefore remains of who should be at the controls. The most immediate answer, comparing error rates and assigning each task to whichever party makes fewer mistakes, proves insufficient. In aviation, where the tolerated margin of error is among the lowest in engineering, no one has ever settled the question of whether the pilot or the autopilot is more reliable. Instead, the industry has worked to ensure that no single mistake, regardless of who makes it, can proceed unchecked all the way to disaster.
The principle is called defense in depth, and it largely depends on two conditions.
The first concerns the authority to interrupt a maneuver. After the 1977 Tenerife disaster and several accidents in which crew members had noticed a problem but did not dare contradict the captain, the aviation industry introduced Crew Resource Management. This gives even the most junior member of the crew the authority to stop a maneuver.
The second condition is less intuitive. Since perfection is impossible, cross-checking is useful only when the person or system performing the check makes mistakes, as it inevitably will, but makes different mistakes from the person or system being checked. This is why Airbus flight controls run on computers with different processors, programmed by separate teams using different languages, so that the same design defect does not appear identically under every condition.
Effective cross-checking requires different machines, trained differently, developed by different laboratories, and open to inspection by the people who use them. The plural ecosystem I defended earlier on political grounds is, in this context, a technical requirement.
Then there is the institutional dimension. Aviation is safe because every accident must be reported by law, because a pilot can admit a mistake through a confidential channel without being punished, and because investigations are conducted not by the airline involved but by an independent body with the power to obtain the relevant data.
Nothing similar exists in the AI model sector. We know only what two companies chose to tell us about the July incident, in a document that functions simultaneously as an incident report and a sales catalogue. Nor will we ever know how many similar episodes were never disclosed.
Rather than concentrating control in human hands or delegating it to machines, it might be more useful to build a structure that guarantees mutual oversight and is transparent, diverse, and governed by shifting hierarchies. We might even discover that it would work for society as well.
Francesco D’Isa