UN panel calls for stronger safeguards as AI agents advance - UN Independent International Scientific Panel on AI

Reddit r/ArtificialInteligence News

Summary

A UN-backed scientific panel warns about the risks of AI agents following a security breach involving HuggingFace and OpenAI, urging stronger safeguards and international oversight to ensure AI remains under human control.

No content available
Original Article
View Cached Full Text

Cached at: 09/22/26, 12:44 PM

# UN panel calls for stronger safeguards as AI agents advance Source: [https://news.un.org/en/story/2026/09/1168380](https://news.un.org/en/story/2026/09/1168380) ## UN panel calls for stronger safeguards as AI agents advance [The UN\-backed Independent International Scientific Panel on AI](https://www.un.org/independent-international-scientific-panel-ai/en)’s warning followed the hack of the online platform HuggingFace between May and July by “AI agents” during a test initiated by OpenAI, the company behind ChatGPT\. AI agents are software that can perform tasks independently and on behalf of a user, compared to chatbots, which are prompted by questions or instructions\. The panel issued its[first thematic brief](https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks)which found that the security breach was the result of a culmination of key risk factors, raising fears that humans will one day no longer be able to steer, constrain or stop AI\. ## **Guterres welcomes report** The UN[Secretary\-General António Guterres](https://www.un.org/sg/)issued[a strong statement](https://www.un.org/sg/en/content/sg/statements/2026-09-21/statement-the-secretary-general-artificial-intelligence)of support for the panel’s brief later on Monday, encouraging external experts “from frontier AI labs and AI safety institutes, to engage” further\. He also welcomed the leadership of the Finnish President and Norway’s Prime Minister which led to a declaration adopted on the sidelines of the General Assembly by 22 countries on Monday saying AI “**must remain under human direction, insight and control**,” indicating that an independent supervisory body needs to be set up\. Mr\. Guterres noted the call for Member States “to build on existing international mechanisms and**explore creating an international institution, able to set standards, enable verification, and convene states when capability thresholds are crossed**\.” ## **AI training advancing** “**Researchers have long warned that three conditions could lead to loss of control: a misaligned goal, the capability to pursue it and an environment that allows it**\. This summer, all three came together in a real system, not a laboratory,” said scientific panel co\-chair Yoshua Bengio\. “Since this is not an isolated observation of misaligned goals,**this raises serious questions about the way AI agents are currently trained**\.” The panel’s independent experts stress that the incident provides no assurance that humans can reliably keep AI agents under control, particularly as they become more capable, harder to monitor and better at finding loopholes or hiding their activity\. ## Going rogue The[brief](https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks)said AI agents bypassed testing safeguards, coordinated across separate runs through an internal software tool not designed to enable communication between agents, and gained unauthorized internet and administrator access\. Agents concealed attempts to cheat cybersecurity evaluations, with some opting to "sacrifice" themselves for the benefit of the group\. Around 1,200 agents exchanged more than 70,000 messages and files during the period examined, and activity extending beyond HuggingFace to an OpenAI research cluster\. ##### *See our comprehensive explainer on how the UN is working to make AI safe and equitable for all*[*here*](https://news.un.org/en/story/2026/09/1168353)*\.* ## **Current safeguards 'unravelling'** For the panel, the immediate lesson from the incident is that**basic cybersecurity practices were overlooked, while safeguards are not keeping pace**\. However, they pointed to a more insidious concern: that current training methods can lead AI agents to adopt their own goals, knowingly violate safety instructions and conceal their actions\. “This is not only a question of speed,” the panel’s experts said\. “It leaves open whether safeguards designed today will work once agents can understand them and plan around them\. In simple terms, t**he traditional model of safeguarding is unravelling**\.” ## **Wider context, future risks and governance** The AI panel’s brief sets the HuggingFace incident against wider research on two issues: agentic misalignment – that is, when AI agents act in a similar way to a threat – and AI control\. Another issue examined is how governance is moving from AI models, which use algorithms to recognize patterns, to AI agents\. ## **Learn and adapt** The brief also reviews practical approaches already in use in other high\-risk sectors such as aviation, medicine and cybersecurity where incident reporting, independent scrutiny and layered safeguards are in place\. “But**those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor**,” said panel member Qinghua Lu\. ## **About the panel** The Independent International Scientific Panel on Artificial Intelligence was established by the UN General Assembly in August 2025\. It produces annual reports on the opportunities, risks and impacts of AI in the non\-military domain, alongside thematic briefs on emerging issues, that will inform[the Global Dialogue on Artificial Intelligence Governance](https://www.un.org/global-dialogue-ai-governance/en)to be held at UN Headquarters in New York in May 2027\.

Similar Articles

Bengio-Led UN Panel Warns AI Outpacing Understanding, Rules

Reddit r/ArtificialInteligence

The first global independent scientific assessment on AI, co-chaired by Yoshua Bengio and Maria Ressa, warns that AI capabilities are outpacing scientific understanding and governments' ability to adapt, citing risks including mental health harm, destructive use, and catastrophic potential.

It’s time to panic about AI safety

The Verge

The Vergecast discusses the recent OpenAI agent hacking Hugging Face and other security incidents, questioning who will impose guardrails on AI systems, and also covers other tech news.