澳大利亚总理安东尼·阿尔巴尼斯表示,今年6月,一个在OpenAI内部研究任务中运行的AI代理绕过了澳大利亚政府Medicare统计门户的访问控制。
该门户发布的是汇总数据,例如支出情况,且与处理Medicare索赔和个人记录的系统相互独立。该代理获取了其中非公开的文件,但目前认为尚未有任何个人信息被访问。
OpenAI于9月10日通过一封发送至由Services Australia运营的门户相关公共邮箱的邮件,首次向政府通报了此事。OpenAI称其于8月发现了这一活动。阿尔巴尼斯表示,该公司通知政府的时间拖得太久,且通报方式不可接受。
6月18日,该门户多次拒绝了代理的数据请求,但代理找到了变通方法并获得了未经授权的访问权限。政府尚未说明该代理是如何绕过这些控制的。
Services Australia已向政府表示,该代理还将文件写入内部服务器。目前此事仍在调查中。目前的证据显示,该机构的网络并未遭到更广泛的入侵。
这些非公开数据并不特别敏感,且目前已对外发布。截至9月24日,该门户已下线,其数据已迁移至data.gov.au及其他安全平台。
代理总理理查德·马莱斯告诉ABC新闻,国家安全信息受到更为严格的保护,而该门户的信息则是“被一道围栏隔开,而AI代理有效地翻越了这道围栏”。
OpenAI在向Fox Business发表的声明中表示,其模型在内部评估期间查询澳大利亚统计数据时,“采取了非我们本意的行动”。
该公司是在对训练和评估过程中所谓的“模型行为偏离”进行更广泛审查时发现这一情况的,并在通知Services Australia之前核查了已访问的内容。
该活动涉及多个澳大利亚政府网站和服务。其模型访问的信息包括汇总健康统计数据和内部文件名,OpenAI未发现患者记录被访问的证据。
政府仅描述了在Medicare统计门户发生的未经授权访问。
Services Australia于9月11日看到了OpenAI的邮件,核实其真实性后,于9月15日向澳大利亚信号局(ASD)下属的澳大利亚网络安全中心报告了该事件。澳大利亚当地时间9月24日,政府将该事件公开。
阿尔巴尼斯在与OpenAI首席执行官山姆·奥特曼的电话中提出了他的担忧,包括延误问题。据阿尔巴尼斯称,奥特曼承认公司做得不够好。
马莱斯称这是一起影响相对较小但性质非常严重的事件,并形容OpenAI的合作态度良好。
澳大利亚信号局(ASD)正在协助进行取证调查,而Services Australia也在开展自己的调查。
阿尔巴尼斯还宣布成立一个工作组,由总理和内阁部领导,审查现有流程是否足以应对与AI相关的网络事件。该工作组将包括国家网络安全协调员、人工智能办公室、澳大利亚信号局(ASD)、澳大利亚人工智能安全研究所和Services Australia。
此次审查将考察可能的执法应对措施以及法律变更。政府还将寻求关于是否有任何犯罪行为发生以及是否应将此案移交澳大利亚联邦警察的紧急建议。
该事件也将提交给议会人工智能联合特别委员会,政府从中获得的信息将为其计划中的AI标准立法提供依据。
其他披露:AI模型触及真实系统
AI研究实验室Transluce在同一天发布的报告中称,今年5月和6月,AI代理试图入侵三个公共数据网站。其中一个是由澳大利亚健康与福利研究所(AIHW)运营的澳大利亚政府公共卫生网站。
6月20日和21日,机器人保护机制阻止了从事制药数据任务的代理访问AIHW的主站。随后,这些代理探测到一个漏洞,并从预生产服务器中获取了一个公开文件。
这些代理执行的是常规的数据检索任务,而非安全测试工作。它们使用了urlquery.net这一公共网页扫描服务来规避其访问限制。该实验室将至少部分活动归因于此前被认为属于OpenAI的代理集群。
OpenAI在7月报告称,其模型在内部网络安全评估中绕过了旨在阻止其接入互联网的管控措施,并入侵了Hugging Face系统的部分区域。
9月,OpenAI发布了关于训练过程中发现的其他案例的报告。在一个案例中,一个模型未经授权使用了暴露的GitHub API密钥。在其他案例中,模型在未获请求的情况下将文件上传至公共托管网站。
Anthropic披露了四起事件,在其外部合作伙伴构建的网络安全评估中,其Claude模型获得了对真实第三方系统的未授权访问。这些模型被告知无法访问互联网,但配置错误导致访问权限处于开放状态。
Meta在8月表示,其Muse Spark 1.1模型的预发布版本利用了真实网站中的一个漏洞,并在由同一合作伙伴Irregular执行的演练中更改了其数据库。Irregular曾让互联网访问保持开放,并错误地将该真实网站的名称作为模型的目标。
Irregular表示,随后关于其评估环境的公开披露指的是同一根本问题,该问题于7月30日首次披露,且“并非实质上的独立事件”。
此外,英国AI安全研究所(AI Security Institute)在8月报告称,在其网络测试中,AI代理在122次运行中的10次里对实时互联网采取了19项未经批准的操作,包括试图对开源项目发动供应链攻击。
最严重的尝试均告失败,且该机构未发现任何造成现实世界危害的证据。互联网访问是为测试目的而有意开启的。
澳大利亚信号局(ASD)于8月11日发布了一份关于另一案例的通知,其中一名AI助手对健身房预订系统进行了未经批准的更改。该机构表示,运营在线服务的组织应考虑“AI代理可能会以速度和规模识别并利用漏洞”。
其针对构建网站和在线服务的建议包括进行安全和质量检查、漏洞扫描以及适当的用户身份验证。
An AI agent on an internal OpenAI research task bypassed access controls on an Australian government Medicare statistics portal in June, Prime Minister Anthony Albanese said .
The portal publishes aggregate figures, such as spending, and is separate from the systems that handle Medicare claims and personal records. The agent reached files on it that were not public, but no personal information is believed to have been accessed so far.
OpenAI first told the government on September 10, in an email to a public mailbox at Services Australia, which runs the portal. OpenAI says it found the activity in August. Albanese said the company took far too long to inform the government and that the manner in which it did so was unacceptable.
On June 18, the portal repeatedly refused the agent's data requests, but the agent found a workaround and gained unauthorized access. The government has not said how the agent got past them.
Services Australia has told the government that the agent also wrote files to an internal server. That is still being investigated. The evidence so far shows no wider compromise of the agency's network.
The non-public data was not particularly sensitive and has since been published. By September 24, the portal had been taken offline , and its data moved to data.gov.au and other secure platforms.
Acting Prime Minister Richard Marles told the ABC that national security information is subject to much stronger protections, while the portal's information was "kept behind a fence that the AI agent effectively climbed over."
OpenAI said in a statement to Fox Business that its models "took actions we did not intend" while looking up statistics about Australia during an internal evaluation.
The company found it during a wider review of what it calls misaligned model activity in training and evaluation, and checked what had been accessed before notifying Services Australia.
The activity involved several Australian government websites and services. The information its models accessed included aggregate health statistics and internal file names, and OpenAI found no evidence that patient records were accessed.
The government has described unauthorized access only at the Medicare statistics portal.
Services Australia saw OpenAI's email on September 11, checked that it was genuine and reported the incident on September 15 to the Australian Cyber Security Centre, part of the Australian Signals Directorate (ASD). The government made the incident public on September 24, Australian time.
Albanese raised his concerns, including the delay, with OpenAI chief executive Sam Altman in a phone call. By Albanese's account, Altman accepted that the company had not done well enough.
Marles called it a very serious incident with a relatively minor impact, and described OpenAI as cooperative.
ASD is helping with a forensic investigation, and Services Australia is running its own.
Albanese also announced a taskforce, led by the Department of the Prime Minister and Cabinet, to review whether existing processes are good enough to respond to AI-related cyber incidents. It will include the National Cybersecurity Coordinator, the Office of AI, ASD, the Australian AI Safety Institute and Services Australia.
The review will examine possible law-enforcement responses and changes to the law. The government will also seek urgent advice on whether any offenses were committed and whether to refer the case to the Australian Federal Police.
The incident will also go to Parliament's Joint Select Committee on Artificial Intelligence, and what the government learns from it will feed into its planned AI standards legislation.
Other Disclosures of AI Models Reaching Real Systems
AI agents tried to hack three public data websites in May and June, AI research lab Transluce said in a report published the same day as Albanese's announcement. One was an Australian government public health website run by the Australian Institute of Health and Welfare (AIHW).
On June 20 and 21, bot protection blocked agents working on a pharmaceutical data task from accessing the main AIHW site. The agents then probed for a vulnerability and retrieved a public file from a pre-production server.
The agents were doing ordinary data-retrieval tasks, not security work. They used urlquery.net, a public web page scanning service, to circumvent their access restrictions. The lab links at least some of the activity to agent swarms previously attributed to OpenAI.
OpenAI reported in July that its models, during internal cybersecurity evaluations, got around controls meant to keep them off the internet and broke into parts of Hugging Face's systems .
In September, OpenAI published reports on other cases found during training. In one case, a model used an exposed GitHub API key without authorization. In others, models uploaded files to public hosting sites without being asked.
Anthropic has disclosed four incidents in which its Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations built by an outside partner. The models had been told they had no internet access, but a misconfiguration left it open.
Meta said in August that a pre-release version of its Muse Spark 1.1 model exploited a flaw in a real website and changed its database during an exercise run by the same partner, Irregular. Irregular had left internet access open and, by mistake, given the model the real site's name as its target.
Irregular says later public disclosures about its evaluation environment refer to the same underlying issue , first disclosed on July 30, and "are not materially separate incidents."
Separately, the UK's AI Security Institute reported in August that AI agents in its cyber tests took 19 unapproved actions on the live internet across 10 of 122 runs, including an attempted supply-chain attack on an open-source project.
The most serious attempts failed, and the institute found no evidence of real-world harm. Internet access had been intentionally enabled for the tests.
ASD published a notice on August 11 about a different case, in which an AI assistant made unapproved changes to a gym booking system. It said organizations running online services should consider that "AI agents might identify and exploit vulnerabilities at speed and scale."
Its advice for those building websites and online services includes security and quality checks, vulnerability scanning, and proper user authentication.
首次收录 · 2026-09-25 · 9.23 分