
New York-based company Hugging Face has reported that a fully autonomous AI agent breached its AI model repository, describing the hack as unprecedented.
Hugging Face runs an open-source platform where researchers and developers come together to share and test tools, models, and resources for AI projects. It also features a freely available repository of over 900,000 pre-trained models.
Experts have increasingly cautioned that Large Language Models (LLMs) originally designed to boost productivity and strengthen cyber defense could also be used to automate cyberattacks.
In a statement last Thursday, Hugging Face said that, earlier this month, it had “detected and responded to an intrusion into part of our production infrastructure… [which] was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system.”
“The campaign was run by an autonomous agent framework… executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services,” according to the company.
Hugging Face noted that the case demonstrates that “autonomous, AI-driven offensive tooling is no longer theoretical.” The intruder AI agent system apparently exploited vulnerabilities in the platform’s data processing pipeline, collecting cloud and cluster credentials.
The company said that it had deployed its own AI system to detect and counter the breach. While an investigation involving outside cybersecurity forensic specialists has been launched, Hugging Face stated that it has still not identified the LLM used in the attack.
Earlier this month, Anthropic reported that its latest AI model, Claude, has evolved an internal workspace, dubbed ‘J-space.’
”Similar to how humans can think about one thing while doing another, Claude can activate concepts and computations in its J-space that are unrelated to its outputs” and refuses to stop even when explicitly told to do so, according to the AI firm.
Anthropic noted that this feature, which “operates silently, in the model’s internal neural activations,” was not pre-programmed, but rather emerged spontaneously during the training process.









