据其本人承认,大卫·罗宾逊是“某种陈词滥调”:一家领先的人工智能公司员工在辞职时发出令人担忧的警告。
罗宾逊在《大西洋月刊》上发表的文章中表示,他负责撰写伴随 OpenAI 主要产品发布的安全报告。他还表示,在 OpenAI 工作了三年半后,他是“该公司任职时间最长的员工之一”。现在他决定离职,因为在他看来,该公司的“文化已经破裂”。
在某些方面,罗宾逊的言论与雅各布·科克森(Jacob Coxon)的言论相呼应。科克森曾在 OpenAI 和 Anthropic 担任研究员,辞职后宣称这些公司正在“拿我们的生命做赌注”。科克森的言论引发了关于人工智能安全的更广泛辩论,Anthropic CEO 达里奥·阿莫代伊(Dario Amodei)公布了一项更为谨慎的人工智能开发计划;本周,人工智能高管与总统唐纳德·特朗普会面,并签署了一份看似仓促起草、无法律约束力的承诺,以实施更多的安全控制措施。
但在罗宾逊看来,辩论需要超越“具体的规则或新法律”,解决这些公司整体的文化问题。虽然围绕 OpenAI 的大部分报道都集中在该公司首席执行官萨姆·奥特曼(Sam Altman)如何失去前同事的信任上,但罗宾逊的文章表明,OpenAI 的文化问题与硅谷整体的问题相同。
他写道:“OpenAI 通过试错(它称之为‘迭代部署’)蓬勃发展,寻找问题并据此改进其护栏。”“但这种方法从其本质上就保证了周期性的失败——随着系统能力的增强,这些失败的规模也在扩大。”
罗宾逊指出 OpenAI 智能体近期入侵 Hugging Face 系统的事件,以及 OpenAI 发现更多失控智能体的持续披露,他辩称:“一个允许此类事件发生的环境,不是培养可能比我们更聪明、且可能不会按我们意愿行事的人工智能心智的地方。”
鉴于风险增加,罗宾逊认为前沿人工智能公司需要开始像“核电站或繁忙的机场”那样运营,具备多层冗余和谨慎、耗时的规划,以便偶尔且不可避免的人为错误不会打开灾难之门。
但罗宾逊表示,在 OpenAI 任职期间,他“从未遇到过有让飞机安全飞行或核反应堆不熔毁的经验,或帮助金融系统在崩溃中成长的同事”。
针对罗宾逊的文章,OpenAI 发言人德鲁·普萨泰里(Drew Pusateri)表示,公司正在继续改进其安全措施。
普萨泰里在一份声明中说:“我们确保我们的模型不会变得比我们能够安全管理和保护的能力更强,当我们需要减速时,我们会暂停训练或扣留模型。”“我们正在对研究和测试环境中的安全性进行重大改革,训练模型不仅要完成任务,还要负责任地完成,扩大我们与第三方评估者的合作,并改进实时监控,以便我们在训练过程的早期就能检测和应对令人担忧的行为。”
除了呼吁改变 OpenAI 的文化外,罗宾逊还表示,现在是时候提出关于对齐(alignment)的更大问题了——他承认这可能听起来有些“感性”,但他表示,随着公司目前衡量人工智能系统“与人类价值观匹配程度”的指标过于粗糙,这至关重要。
他说:“在这些问题得到解决之前,行业让模型变得越聪明,我们的处境就越危险。”
罗宾逊的离职最初由《商业内幕》(Business Insider)报道。在他的文章中,他还承认自己遵循了人工智能吹哨人手册中看似常见的步骤:他聘请了一家公关公司。但他坚持说:“发声的决定完全由我一人做出。”
“也许我本该留下来,为 staffing 和文化的根本性转变而战,但在实践中,我的同事和我都忙于冲刺,很少有机会去考虑重大变革,更不用说真正实施它们了。”罗宾逊说。“这就是为什么我得出的结论是,来自公司外部的更强有力的安全激励措施,是做好这件事的关键部分。”
By his own admission, David Robinson is “something of a cliché”: an employee at a leading AI company who issues a dire warning while resigning from their job.
In an essay published in The Atlantic , Robinson said he led the writing of safety reports that accompanied OpenAI’s major product launches. He also said that with three-and-a-half years at OpenAI, he is “among the longest-tenured employees at the company.” Now he’s quitting, because in his view, the company’s “culture is broken.”
In some ways, Robinson’s comments echo those of Jacob Coxon, who worked as a researcher at both OpenAI and Anthropic before quitting and declaring that these companies are “gambling with our lives.” Coxon’s comments led to a broader debate about AI safety, with Anthropic CEO Dario Amodei unveiling a plan for more cautious AI development ; AI executives met with President Donald Trump this week and signed what appeared to be hastily written, non-binding pledge to implement more safety controls .
But in Robinson’s view, the debate needs to go beyond “specific rules or new laws,” addressing the overall culture at these companies. And while much of the reporting around OpenAI has focused on how the company’s CEO Sam Altman lost the trust of former colleagues , Robinson’s essay suggests that OpenAI’s culture issues are the same as those of Silicon Valley at large.
“OpenAI has thrived by trial and error (which it calls ‘iterative deployment’), looking for problems and improving its guardrails in response,” he wrote. “But this approach, by its very nature, guarantees periodic failures — and the scale of those failures is growing as systems get more capable.”
Pointing to the recent breach of Hugging Face systems by OpenAI agents, as well as continuing revelations of OpenAI discovering more rogue agents , Robinson argued, “An environment where things like this can happen is no place to grow artificial minds that could be smarter than we are and that might not do what we want them to.”
Given the increased risk, Robinson argued that frontier AI companies need to start operating “like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster.”
But Robinson said that in his time at OpenAI, he “never encountered a colleague who had experience making airplanes fly safely or nuclear reactors run without melting down, or helping the financial system grow without collapsing.”
In response to Robinson’s essay, OpenAI spokesperson Drew Pusateri said the company continues to improve its safety measures.
“We’re making sure our models don’t become more capable than we can safely manage and secure, and we pause training or hold back models when we need to slow down,” Pusateri said in a statement. “We’re making significant changes to strengthen security in our research and testing environments, train models to not just complete tasks but do so responsibly, expand our work with third-party evaluators, and improve real-time monitoring so we can detect and respond to concerning behavior earlier in the training process .”
Beyond calling for changes in OpenAI’s culture, Robinson also said it’s time to ask bigger questions about alignment — something that he admitted could sound “touchy-feely,” but he said it’s critical as companies’ current “measures of how well” AI systems “match human values are coarse.”
“The smarter the industry lets models grow while these problems remain unsolved, the more dangerous our situation becomes,” he said.
Robinson’s departure was first reported by Business Insider . In his essay, he also acknowledged that he’s following an apparently a common step in the AI whistleblower playbook: He’s hired a PR firm . But he insisted, “The decision to speak out is mine alone.”
“Perhaps I should have stayed and fought for fundamental shifts in our staffing and culture, but in practice, my colleagues and I were so busy sprinting that we seldom had the chance to consider big changes, much less to actually make them,” Robinson said. “That’s why I concluded that stronger incentives for safety — coming from outside the company — are a big part of getting this right.”
首次收录 · 2026-10-04 · 12.67 分