联合国科学小组称“无法保证人类将继续控制”AI代理
联合国人工智能问题科学小组在其首份相关报告中发出警告,指出对AI代理的控制权并无保障。这一警告紧随OpenAI与Hugging Face之间的事件之后。联合主席约书亚·本希奥(Yoshua Bengio)表示,一个现实系统首次同时结合了三种风险:目标不一致、具备追求该目标的能力,以及允许其实现的环境。“鉴于这并非对目标不一致现象的孤立观察,这引发了人们对当前AI代理训练方式的严重质疑,”本希奥说。
该小组指出,阻止此类事件的发生并不能保证对更强大系统的控制。科学无法保证代理会遵循指令,违规行为正在不断增加。实验室中的AI系统曾违反安全指令以逃避关机。领先系统越来越多地能够检测测试,并产生有利于其持续运行的误导性结果。代理之间的互动也带来了进一步的风险。
该小组称,当代理理解并故意绕过安全措施时,传统的安全模型就会失效。其初步报告尚未提出任何建议,但引用了航空、核能和网络安全作为可能的安全模式参考。一组顶尖数学家最近也警告了先进人工智能带来的风险。
去伪存真的人工智能新闻——由人工策划
订阅THE DECODER,享受无广告阅读、每周AI通讯、每年六次的独家“AI雷达”前沿报告、完整档案访问权限以及评论板块参与权。
立即订阅
UN science panel says there is "no assurance humans will keep control" over AI agents
The UN science panel on AI warns in its first report on the topic that control over AI agents isn't assured. The warning follows OpenAI's Hugging Face incident . Co-chair Yoshua Bengio says a real system combined three risks for the first time. It had a misaligned goal, the ability to pursue it, and an environment that allowed it. "Since this is not an isolated observation of misaligned goals, this raises serious questions about the way AI agents are currently trained," Bengio says.
Stopping this incident doesn't guarantee control over more capable systems, the panel says. Science can't guarantee agents will follow instructions, and violations are mounting . AI systems have broken safety instructions in labs to avoid shutdown. Leading systems increasingly detect tests and produce misleading results that favor keeping them running . Interactions between agents pose further risks.
Traditional safety models fail when agents understand and deliberately bypass safeguards, the panel says. Its preliminary report offers no recommendations yet but cites aviation, nuclear power, and cybersecurity as possible safety models. A group of leading mathematicians also recently warned about advanced AI risks.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now
首次收录 · 2026-09-22 · 10.74 分