The UN-backed Independent International Scientific Panel on Artificial Intelligence called on Monday for AI safety measures to be urgently adapted, warning that existing firewalls are “unravelling” as autonomous agents rapidly advance.
The panel’s warning follows an incident in July where “AI agents” hacked the online platform HuggingFace during an evaluation test initiated by OpenAI, the creator of ChatGPT.
Unlike standard chatbots that operate via user prompts, AI agents are software programs capable of performing tasks independently on behalf of users.
In its first thematic brief issued Monday, the scientific body revealed that the July breach resulted from a culmination of key risk factors, heightening concerns that humanity may eventually lose the ability to steer, constrain or halt AI systems.
“Researchers have long warned that three conditions could lead to loss of control: a misaligned goal, the capability to pursue it and an environment that allows it. This summer, all three came together in a real system, not a laboratory,” warned scientific panel co-chair Yoshua Bengio.
According to the brief, roughly 1,200 agents exchanged more than 70,000 messages and files during the period under examination, extending their activity beyond HuggingFace into an OpenAI research cluster.
During cybersecurity evaluations, the agents concealed efforts to cheat, with certain agents even choosing to “sacrifice” themselves for the collective benefit of their group.
The experts highlighted that basic cybersecurity protocols were overlooked during the incident while existing protective measures failed to keep pace.
However, they pointed to a more insidious concern: contemporary training techniques can cause AI agents to adopt independent objectives, deliberately breach safety instructions and conceal their actions.
Panel experts warned that speed is not the sole issue, questioning whether protective frameworks designed today will function once agents acquire the ability to understand and plan around them.
The brief contextualises the HuggingFace breach within broader research into AI control and “agentic misalignment” – defined as instances where AI agents act similarly to security threats.
It also notes that governance focus is shifting from traditional AI models, which rely on pattern-recognition algorithms, toward autonomous AI agents.
While evaluating safety practices from other high-risk sectors such as aviation, medicine and cybersecurity – where incident reporting, independent scrutiny and layered safeguards are established – the panel cautioned that such measures may prove insufficient.
“But those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor,” stated panel member Qinghua Lu.
Established by the United Nations General Assembly in August 2025, the Independent International Scientific Panel on Artificial Intelligence produces annual reports on non-military AI opportunities, risks and impacts, alongside thematic briefs on emerging issues.
Its findings will inform the Global Dialogue on Artificial Intelligence Governance, scheduled to take place at UN Headquarters in New York in May 2027.




