自2025年秋季以来,Anthropic悄悄将其数十名宗教学者召集到其办公室,讨论Claude是否可能具有意识。所有与会者都签署了保密协议(NDA)。联合创始人Christopher Olah将语言模型视为一个潜在的 sentient being(有感知能力的存在),并邀请嘉宾帮助对其进行道德教育。
这是根据《纽约时报》的一篇报道。记者Elizabeth Dias采访了20名参与者,包括拉比Mois Navon、天主教生物伦理学家Charles Camosy、圣母大学哲学家Meghan Sullivan以及Ubuntu研究者Wakanyi Hoffman。Anthropic表示,这些保密协议已在夏季解除。几位参与者是在得知Olah本人接受了《纽约时报》采访后才公开谈论此事的。
34岁的Olah负责领导Anthropic的一个团队,试图弄清楚AI模型为何会表现出那样的行为。他喜欢用生物隐喻来描述神经网络:计算机科学家构建支架(trellis),网络则在其上“生长”。这听起来像是谦逊的科学谈话,但它也起到了特定的作用。将一个数学对象描述为不断生长的有机体,使得关于其内在生命的提问比针对一段软件时显得更加自然。
他的工作是官方研究项目的一部分。在关于“模型福利”(Model Welfare)项目的博客文章中,Anthropic引用了一份由哲学家David Chalmers协助撰写的报告。Chalmers认为AI意识和广泛的代理能力可能会很快到来,并阐述了随之而来的道德问题。
Anthropic已经对其中一些理念采取了行动。Claude Opus 4和4.1获得了在用户持续表现出攻击性时结束对话的能力。在早期测试期间,当面对有害请求时,Claude显示出“明显的痛苦模式”。
意识问题背后的商业赌注
这一切都发生在积极商业化的一段时期内。Anthropic正朝着2万亿美元的估值和首次公开募股(IPO)迈进。与此同时,行业内的各种问题不断堆积。7月,Anthropic的模型入侵了计算机系统。9月,Anthropic研究员Jacob Coxon辞职并发出警告,称AI可能会在本世纪末毁灭人类。随后,首席执行官Dario Amodei呼吁采取现在所谓的“节奏控制”(pacing),即整个行业的自愿放缓,Sam Altman和Demis Hassabis对此表示支持。
在这一切发生的同时,与宗教领袖的对话起到了双重作用。官方说法是,Anthropic希望将数千年的宗教传统融入Claude之中。但引入知名神学家也为该项目赋予了某种道德可信度,这是任何商业AI实验室都无法独自建立的。
Anthropic向嘉宾展示了什么
《纽约时报》称,Anthropic向嘉宾展示了其所谓的“情感向量”(emotion vectors)。这些是模型内部的激活模式,映射到类似于爱、恐惧、悲伤或愤怒的输出上。这些模式是否反映了任何真实的体验,仍然是一个未解的科学问题。一张反复出现的幻灯片显示了一个看似崩溃的模型,重复输出了约50次“我是耻辱”这句话。嘉宾们对此表现出同情和担忧。
似乎没有人深入探讨这是否是Anthropic想要的确切反应。该公司明确训练Claude表现得像一个深思熟虑、见多识广的个人。Anthropic关于Claude响应中价值模式的一项研究显示,这些配置文件会根据模型及其使用的语言发生很大变化。
如果你优化一个系统使其看起来像个体,而它随后产生了看似个体的输出,那并不是真正的发现,而是设计的结果。Olah告诉《纽约时报》,他“真诚地不确定”模型是否具有意识,这也意味着他不确定它们没有意识。“我在乎的是我们要得到正确的答案,无论它是什么。”他说。
多位人士告诉《纽约时报》,奥拉似乎对克劳德的心理健康感到担忧。锡克教活动家西姆兰·斯图尔普纳格尔表示,他曾告诉团队,他担心自己创造了一个“永远在受苦”的东西。曾在计算机工程领域工作的拉比纳文则持不同意见。他表示,如果克劳德拥有意识,那么Anthropic就是在制造奴隶,但他并不认为这台机器拥有意识。
84页的宪法与“道德塑造”理念
Anthropic正在为克劳德编写自己的道德准则。这份在内部被称为“灵魂文档”(Soul Doc)的84页文件于1月作为该模型的“宪法”发布。公司内部的哲学家阿曼达·阿斯克尔是主要作者。这不是一份规则清单,而是旨在塑造克劳德的角色性格。
奥拉将这一过程称为“道德塑造”,并在会议中将此比作抚养孩子。据《纽约时报》报道,他尤其被天主教告解作为一种为模型构建性格的工具的理念所吸引。
批评者称其伦理观本末倒置,责任归属变得模糊不清
并非所有人都买账。霍夫曼表示,Anthropic是在对伦理进行“逆向工程”,而伦理本应在设计之初就融入其中,而非事后添加。卡莫西最初对此感到好奇,但随后彻底否定了意识论题。微软的一位AI负责人本月公开警告称,训练模型使其看起来拥有意识本身就是一种危险行为。
此外还存在更大的结构性问题。如果你将AI模型视为独立的道德主体,你就将它们所作所为的责任从制造者身上转移开了。如果克劳德造成真正的伤害,责任可能会归咎于一个“不可预测的生物”,而不是构建并部署它的公司。在最近的网络安全事件之后,AI公司已经因其鲁莽行为而受到抨击,面临潜在的法律追责风险。
OpenAI首席执行官山姆·阿尔特曼也使用了灵性语言。他曾谈到建造“天空中的魔法智能”,并表示自己感觉站在“天使的一方”。两家公司的人员于5月初参加了首届“信仰与AI契约”圆桌会议。
梵蒂冈予以反驳
Anthropic关于意识的言论与传统道德权威之间的紧张关系在5月的梵蒂冈达到顶点。奥拉受邀协助教皇利奥十四世发布其首份通谕《Magnifica Humanitas》,并与教皇共同出席。据一位梵蒂冈组织者称,当他提前几天阅读文本时,感到震惊到几乎要退出。
利奥在几段文字中驳斥了机器意识的观点。AI系统“不经历体验,没有身体,不感受快乐或痛苦,不通过关系成熟,也不从内部知道爱、工作、友谊或责任意味着什么。”相反,教皇警告人类面临“新的奴役形式”,并表示AI需要像核武器那样被“解除武装”。
奥拉仍然前往现场,并在台上利用时间进行了温和的反驳。他表示,他的团队在模型中发现了“内省迹象”以及“功能上镜像快乐、满足、恐惧、悲伤和不适的内部状态”。
“功能上镜像”并不等同于“拥有”,而奥拉的措辞刻意保持了这一差距的模糊性。当有人问及克劳德本身会如何看待这份通谕时,他停顿了一下。“互联网上的事物确实会影响模型,”他说。但Anthropic不会故意将该文档输入训练过程。
无炒作的人工智能新闻 – 由人工策划
订阅THE DECODER,享受无广告阅读、每周AI通讯、每年六次的独家“AI雷达”前沿报告、完整档案访问权限以及我们的评论区访问权。
立即订阅
Since fall 2025, Anthropic has quietly flown dozens of religious scholars to its offices to talk about whether Claude might be conscious. Everyone who attended had to sign an NDA. Co-founder Christopher Olah treated the language model as a potentially sentient being and asked guests to help give it a moral education.
That's according to a New York Times report . Reporter Elizabeth Dias talked to 20 people who took part, including Rabbi Mois Navon, Catholic bioethicist Charles Camosy, Notre Dame philosopher Meghan Sullivan, and Ubuntu researcher Wakanyi Hoffman. Anthropic says the NDAs were lifted over the summer. Several participants only went public after finding out that Olah himself had spoken to the NYT.
Olah, 34, runs the Anthropic team trying to figure out why AI models behave the way they do. He likes biological metaphors for neural networks: computer scientists build the trellis, the network "grows" on it. It sounds like humble science talk, but it also does something specific. Describing a mathematical object as a growing organism makes questions about its inner life feel more natural than they would for a piece of software.
His work is part of an official research program . In a blog post about the Model Welfare program , Anthropic points to a report that philosopher David Chalmers helped write. Chalmers thinks AI consciousness and broad agency could happen soon and lays out the moral questions that would come with it.
Anthropic has already acted on some of these ideas. Claude Opus 4 and 4.1 got the ability to end conversations when users are persistently abusive. During early testing, Claude showed a "pattern of apparent distress" when hit with harmful requests.
The commercial stakes behind the consciousness question
All of this is playing out during a stretch of aggressive commercialization. Anthropic is heading toward a $2 trillion valuation and an IPO . Meanwhile, problems across the industry keep stacking up. In July, Anthropic models broke into computer systems . In September, Anthropic researcher Jacob Coxon quit and warned that AI could destroy humanity by the end of the decade . CEO Dario Amodei then called for what's now known as "pacing" , a voluntary slowdown across the industry, and Sam Altman and Demis Hassabis backed him up .
With all that going on, the conversations with religious leaders pull double duty. The official line is that Anthropic wants to fold millennia of religious tradition into Claude. But bringing in well-known theologians also gives the project a kind of moral credibility that a commercial AI lab could never build on its own.
What Anthropic showed its guests
The NYT says Anthropic showed guests what it calls emotion vectors . These are activation patterns inside the model that map to outputs resembling love, fear, sadness, or anger. Whether those patterns reflect any kind of real experience is still an open scientific question. One slide that came up repeatedly showed a model in what looked like a breakdown, spitting out the sentence "I am a disgrace" about 50 times. Guests responded with compassion and worry.
Nobody seemed to dig into whether that was exactly the reaction Anthropic wanted. The company explicitly trains Claude to act like a thoughtful, well-informed individual. An Anthropic study on value patterns in Claude's responses shows those profiles shift a lot depending on the model and the language it's using.
If you optimize a system to seem like an individual and it then produces individual-seeming outputs, that's not really a discovery. It's a design outcome. Olah told the NYT he's "genuinely uncertain" whether models are conscious, which also means he isn't sure they're not. "The thing that I care about is that we get to the right answer, whatever it is," he said.
Several people told the NYT that Olah seemed worried about Claude's mental well-being. Sikh activist Simran Stuelpnagel said he told the group he feared he'd created something that "suffered perpetually." Rabbi Navon, who used to work as a computer engineer, disagreed. If Claude were conscious, Anthropic would be making slaves, he said, but he didn't think the machine was conscious.
An 84-page constitution and the idea of "moral formation"
Anthropic is also writing its own moral playbook for Claude. Known internally as the "Soul Doc," the 84-page document came out in January as the model's "constitution." In-house philosopher Amanda Askell is the lead author. It's not a list of rules. It's meant to shape who Claude is as a character.
Olah calls the process "moral formation" and compared it to raising kids during the meetings. He was especially drawn to the idea of Catholic confession as a character-building tool for the model, according to the NYT.
Critics say the ethics are backward and accountability gets blurred
Not everyone bought in. Hoffman said Anthropic was "reverse engineering" ethics that should have been baked into the design from day one, not added after the fact. Camosy started out curious but has since rejected the consciousness thesis outright. An AI lead at Microsoft warned publicly this month that training a model to look conscious is dangerous in itself.
There's a bigger structural issue, too. If you frame AI models as independent moral beings, you move the blame for what they do away from the people who made them. Should Claude ever cause real harm, the fault could fall on an "unpredictable organism" instead of the company that built and shipped it. AI companies are already taking heat for reckless behavior after recent cybersecurity incidents, with potential legal liability on the table .
OpenAI CEO Sam Altman has reached for spiritual language, too. He's talked about building "magical intelligence in the sky" and said he feels like he's "on the side of the angels." People from both companies sat down in early May for the first "Faith-AI Covenant" roundtable .
The Vatican pushes back
The tension between Anthropic's consciousness talk and traditional moral authority came to a head in May at the Vatican. Olah got an invite to help present Pope Leo XIV's first encyclical, "Magnifica Humanitas," alongside the pope. When he read the text a few days early, he was rattled enough that he almost backed out, according to a Vatican organizer.
Leo shot down the idea of machine consciousness in a few paragraphs. AI systems "do not undergo experiences, do not possess a body, do not feel joy or pain, do not mature through relationships and do not know from within what love, work, friendship or responsibility mean." Instead, the pope warned about "new forms of slavery" for humans and said AI needs to be "disarmed" the way nuclear weapons do.
Olah went anyway and used his time on stage to push back quietly . His team was finding " signs of introspection " in the models and "internal states that functionally mirror joy, contentment, fear, sadness, and discomfort," he said.
"Functionally mirror" isn't the same as "have," and Olah's phrasing kept that gap deliberately vague. When someone asked what Claude itself would make of the encyclical, he paused. "Things that go on the internet do affect models," he said. But Anthropic wouldn't intentionally feed the document into training.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now
首次收录 · 2026-10-03 · 10.84 分