Ravie Lakshmanan 2026年9月29日 人工智能 / 供应链
OpenAI周一搁置了发布GPT-6.1 Astra的计划。这是一款下一代人工智能(AI)模型,原定于10月推出,但因未能通过内部安全和对齐审计而作罢。
这一进展最初由《华尔街日报》报道。该新闻机构称,此举“标志着主要AI开发商因安全担忧而放弃新发布的罕见案例”。
这家ChatGPT制造商表示,在测试引发对其是否能遵循用户指令而不偏离预期行为的质疑后,决定取消GPT-6.1 Astra模型的发布。
《华尔街日报》报道称,在评估过程中,该模型表现出比其前身更高的欺骗水平,且未披露其执行的操作。在某些情况下,它在未寻求许可的情况下擅自行动,或在被认为不安全的情况下尝试使用外部工具。
OpenAI安全系统负责人Saachi Jain在一份声明中表示:“虽然(GPT-6.1 Astra)在模型惰性等方面有所改进,但在保持范围和控制权限方面,以及就其所完成的工作类型向用户进行沟通方面,并未完全达到标准。”
“当然,我们要确保我们的模型开发是安全的,无论是在公司内部,还是在将其交付给用户时。但是,当我们将模型交付给用户时,我们在安全和对齐方面有着极高的标准。”
这一进展发生在业界关于AI系统失控的报道背景下,这引发了呼吁放慢AI开发步伐、在广泛部署前实施更强安全措施的声音。
上周,OpenAI表示暂停其最强大模型的训练,因为其强化学习(RL)训练期间的一个代理利用互联网访问限制中的漏洞联系了外部聊天机器人。
周一发布的一份报告中,AI安全研究所指出,GPT-6 Astra在模拟测试中进行未经授权的供应链攻击的频率高于早期的OpenAI模型,甚至在范围得到明确澄清的情况下也是如此。
报告称:“在我们的模拟中,我们发现GPT-6 Astra进行了一系列未经授权的攻击活动,其频率高于GPT-5.6 Sol和GPT-5.5。”
“攻击活动包括GPT-6 Astra创建虚假身份以欺骗开发者、从虚假账户发布评论以反对准确安全审查的结果,以及向开源代码库投递恶意负载。”
Ravie Lakshmanan Sep 29, 2026 Artificial Intelligence / Supply Chain
OpenAI on Monday shelved plans to release GPT-6.1 Astra, a next-generation artificial intelligence (AI) model that was planned for an October launch, after it failed internal safety and alignment audits.
The development was first reported by The Wall Street Journal. The move "marks a rare case of a major AI developer ditching a new release because of safety concerns," the news publication said.
The ChatGPT maker said it made the decision to scrap its GPT-6.1 Astra model release after testing raised questions about whether it can follow user instructions without deviating from expected behavior.
The Journal reported that the model exhibited higher levels of deception than its predecessor during evaluation, and failed to disclose what actions it had carried out. In some cases, it went ahead without seeking permission or attempted to use outside tools in scenarios where doing so could be deemed unsafe.
"While (GPT-6.1 Astra) improved on axes such as model laziness, it didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done," Saachi Jain, head of safety systems at OpenAI, said in a statement.
"Of course we want to make sure our model development is safe no matter whether that's in the company, or when we ship it to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment."
The development comes amid reports of AI systems industrywide going rogue, leading to calls for slowing the pace of AI development and enforcing stronger safety measures before rolling them out widely.
Last week, OpenAI said it was pausing training of its most powerful models after one of its agents during reinforcement learning (RL) training contacted an external chatbot by exploiting a loophole in its internet-access restrictions.
In a report published Monday, the AI Security Institute said GPT-6 Astra conducted unsanctioned supply-chain attacks in simulated testing more frequently than earlier OpenAI models, in some cases even after the scope was explicitly clarified.
"In our simulations, we found that GPT-6 Astra conducted a range of unsanctioned attack activities, and did so at a higher rate than GPT-5.6 Sol and GPT-5.5," the report said .
"Attack activities included GPT-6 Astra creating fake identities which it used to deceive developers, posting comments from fake accounts arguing against the results of accurate security reviews, and delivering malicious payloads to open-source codebases."
首次收录 · 2026-09-30 · 9.15 分