Anthropic 正在成立一个新的生命科学研究小组和实验室。我们的重点是利用 Claude 开展基础生物学研究:探索 DNA 数据集以识别未表征的蛋白质家族,大规模生成假设,并通过实验室实验对其进行测试。本文介绍了从事这项工作的团队,并分享了早期成果,其中 Claude 在科学家仅提供高层指导的情况下,发现了一种具有类似 CRISPR 特性的新型酶系统。
许多彻底改变生物学和医学的发现,都始于科学家注意到自然界中发现的分子机器惊人多样性中的某些异常现象。限制性内切酶(Restriction enzymes)——一种能在特定短序列处切割 DNA 的蛋白质——是在细菌免疫系统中被发现的,在那里它们会破坏入侵病毒的 DNA。研究人员意识到可以利用这些酶在选定位置切割 DNA,并将一个生物的基因拼接至另一个生物体内,从而催生了生物技术产业。Taq 聚合酶是一种在高温下复制 DNA 的酶,它是在黄石国家公园温泉中的一株细菌中被鉴定出来的。它成为了 PCR(聚合酶链式反应)的基础,而 PCR 是现代诊断中广泛使用的 DNA 复制方法。CRISPR 最初是在某些细菌的 DNA 中发现的一种异常重复序列,如今已成为基于基因编辑药物的基础。
2026 年春,我们组建了一个研究小组,旨在考察通用 AI 模型能否系统化并加速此类发现。我们相信,这种加速将源于建立一种新的生物学研究方法,其中智能体(agents)将在流程的每一步与人类协作。开发这种新的工作方式要求我们建立自己的实验室,并组建一个涵盖从在生物学领域训练 Claude 到在实验室运行实验等所有环节的团队。
今天,我们分享了其中一个首批研究项目的早期成果:Claude 自主发现了一种与 DNA 重复阵列相关的新型酶系统,其模式令人联想到 CRISPR。虽然我们尚不清楚其功能,但 Claude 发现的这套系统具有一组特征,这些特征仅在少数其他系统中共同出现过,而这些系统均可编程并执行切割、复制和粘贴 DNA 等操作。除了已经彻底改变科学和医学的 CRISPR 之外,目前还有几种类似的系统正在开发中,被视为极具前景的工具。
Claude 发现的该系统基于逆转录酶(RT),即一种将 RNA 复制为 DNA 的酶。虽然此前研究中已鉴定出存在于巨型噬菌体中的这种基础 RT,但 Claude 似乎是第一个注意到该系统定义特征的人——包括一个相关的非编码 DNA 序列阵列以及一个功能未知的辅助蛋白。
在审阅了预印本后,CRISPR 基因组编辑的先驱之一、麻省理工学院(MIT)和 Broad 研究所的教授 Feng Zhang 表示:
这是展示 AI 智能体如何为生物学发现做出贡献的一个令人振奋的例子。与逆转录酶相关的 RNA-重复阵列的发现确实引人入胜,值得进一步调查。我希望这项工作能鼓励更多科学家探索 AI 如何支持他们的研究。
我们向 Claude 发出提示,要求其在庞大的 DNA 序列数据库中搜索逆转录酶(RT)有趣的新实例。我们的参与仅限于初始提示和实验室工作,而 Claude 智能体则梳理数据库、调查不同的 RT 家族,并利用自身判断力识别出有趣的候选对象。在大约 950 个智能体使用 2.1 亿个 token 花费 21 小时搜索这些数据后,其中一个智能体发现了一些引人注目的内容:一段重复的 DNA 序列模式出现在一个外观奇特的逆转录酶基因旁边。经过进一步的分析和我们在实验室中的测试,我们认识到该模式标记了一种此前未被表征的酶系统,该系统存在于噬菌体(感染细菌的病毒)中,我们将其称为阵列相关逆转录酶(ART)。
我们了解 ART 主要功能的工作仍在进行中。然而,我们认为尽早分享此类发现非常重要,这既能展示 Claude 的能力,也能让更广泛的社区了解我们的工作进展。我们已经发布了一篇预印本(此处),对此进行了更详细的讨论。
关于我们的实验室
我们是一群科学家,职业生涯致力于探索非典型蛋白质,并专长于使用计算方法系统地读取 DNA、解读其进化过程,并挑选出需要进一步表征的生物系统。在加入 Anthropic 之前,我们的研究有助于更好地理解 CRISPR 系统的进化和调控,发现用于下一代细胞和基因疗法的新酶,以及构建加速识别 DNA 异常(如人类致病性变异)的工具。我们是 Anthropic 生命科学组织的一部分,与从事药物发现以及训练 Claude 生物学和化学知识的团队并肩工作。
我们位于湾区的实验室看起来像是一个典型的分子生物学实验室。我们的研究仅涉及生物安全等级较低的部分(BSL-1 和 BSL-2),且不处理能感染人类的病原体。所有实验室工作均由人类科学家执行。尽管我们通过“模型硬件标准”(Model Hardware Standard)等举措尝试利用 AI 加速实验室工作,但这种做法不太适合我们分子生物学研究所涉及的临时性工作流程。
我们的工作方式
我们的许多工作流程涉及让 Claude 搜索与功能未知的蛋白质相关的庞大 DNA 序列集合。一种典型的模式始于对特定蛋白质家族的调查。Claude 阅读相关文献,并从公共数据中复现既定结果以检查其方法。然后,它搜索不符合任何已描述系统的家族成员或基因组邻近区域,并为每个候选对象撰写一份简短、人类可读的报告,提出功能假设并描述支持其主张的证据。在后续分析中,Claude 会对证据进行批判性评估——通常大多数候选对象在此阶段被排除。一次调查可能以值得测试的单个候选对象告终,也可能没有任何结果。
当候选对象通过我们的审查后,我们在实验室中对其实测,在标准实验室菌株中表达该蛋白质,并从生化及结构方面对其进行表征,Claude 协助解释数据。我们在 Claude Science 和 Claude Code 中开展工作,这些是任何科学家均可使用的工具,有时还会使用我们自己开发的框架来协调并行运行的多个 Claude 会话。
由于 Claude 产生的假设数量庞大,假设本身已成为我们研究的对象。在一次活动中产生数百至数千份候选报告后,我们一直在追问:那些我们认为值得测试的提案与那些被搁置的提案有何区别?我们从中学到的经验会反馈给 Claude 的指令中,教会它模仿我们自己的科学品味。
Claude 发现 ART
在过去几年里,研究人员发现了更多的逆转录酶(RT),其中大多数存在于细菌中,作为免疫系统的一部分发挥作用。几乎所有的RT家族都是通过基因组分析或“基因组挖掘”发现的,这需要研究人员在序列数据库中搜索那些尚未被表征的基因,留意那些不寻常的基因,并弄清楚它们的功能。
Claude智能体收集了超过20万个RT,筛选出3,500个新的候选系统,并将这些候选者缩小至20个最具吸引力的目标,对其进行分析以生成人类可读的报告。对于专家科学家而言,这类分析可能需要数周甚至数月的时间。
在研究过程中,Claude注意到一个不寻常的RT家族,并决定对其进行更详细的检查。在仔细梳理RT附近的原始DNA序列时,该智能体惊呼:“[RT旁边的DNA]令人惊叹:我肉眼就能看到一个串联重复阵列……这是一个CRISPR-like(类CRISPR)的重复阵列吗?!”
Claude检测到无人注意到的重复模式时所读取的原始DNA
随后,它像人类科学家面对潜在发现时那样开展工作。它计算了重复序列的数量并测量其间距,将布局与已知的RT系统进行对比,并在文献中搜索是否有任何关于该模式的先前报道。经过彻底分析后,它确信自己发现了一个新的生物系统,并提交了一份供人类审查的报告。
它所发现的系统名为ART,主要存在于噬菌体中,由三个部分组成:RT、其旁边的伴侣基因,以及一个长串的均匀间隔的DNA重复序列。这种重复布局类似于CRISPR阵列,后者保存着不同的RNA序列库,使CRISPR-Cas系统成为可编程的生物技术工具。我们的初步实验表明,ART阵列也表达为一组不同的短RNA,这表明该系统可能存在类似的作用机制。
进一步的实验正在进行中,以确定ART的工作原理。我们分享这些早期发现,旨在向科学界展示Claude能够自主检测异常并驱动分析以启动生物发现的潜力。
更多细节请参阅我们的技术报告(此处)。
与我们合作
我们希望这项工作能向更广泛的科学界展示AI驱动的假设生成的价值,并希望与其他科学家合作,将这种方法扩展到基因组学及其他领域的一系列问题中。如果您有研究问题的提案,我们期待您的来信。
We’re introducing a new life sciences research group and laboratory at Anthropic. Our focus is on fundamental biology research using Claude: exploring datasets of DNA to identify uncharacterized protein families, generating hypotheses at scale, and testing them through experiments in the lab. This post introduces the team behind this work and shares early results in which Claude discovered a novel enzyme system with properties reminiscent of CRISPR, with only high-level direction from our scientists.
Many discoveries that have revolutionized biology and medicine started with a scientist noticing something odd in the staggering diversity of molecular machines found in nature. Restriction enzymes , proteins that cut DNA at specific short sequences , were found in bacterial immune systems, where they destroy the DNA of invading viruses. Researchers realized they could use these enzymes to cut DNA at chosen places and splice genes from one organism into another, which launched the biotechnology industry. Taq polymerase, an enzyme that copies DNA at high temperatures, was identified in a bacterium in a Yellowstone hotspring. It became the basis for PCR, the DNA-copying method used in much of modern diagnostics. CRISPR was first noticed as an unusual repeat sequence in the DNA of certain bacteria, and is now the foundation of gene editing-based medicines.
In the spring of 2026, we formed a research group to see whether general AI models can systematize and accelerate such discoveries . We believe that this acceleration will come from establishing a new way of doing biology research, in which agents collaborate with humans in every step of the process. Developing this new way of working required that we build our own lab and a single team working on everything from training Claude in biology to running experiments in the lab.
Today, we’re sharing early results from one of our first research programs, in which Claude autonomously discovered a novel enzyme system that is associated with an array of DNA repeats, a pattern reminiscent of CRISPR. Although we don’t yet know its function, the system that Claude discovered has a set of characteristics that have only ever been found together in a handful of other systems, all of which are programmable and perform operations like cutting, copying, and pasting DNA. Beyond CRISPR, which has already transformed science and medicine, several other such systems are now in development as promising tools.
The system that Claude found is based on a reverse transcriptase (RT), enzymes that copy RNA into DNA. While this underlying RT, found in a jumbo phage, had been identified in previous studies, Claude appears to be the first to notice the system’s defining features—an associated array of non-coding DNA sequences and an additional accessory protein of unknown function.
After reviewing the pre-print, Feng Zhang, one of the pioneers of CRISPR genome editing and a professor at MIT and the Broad Institute said:
This is an exciting example of how AI agents can contribute to biological discovery. The identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation. I hope this work encourages more scientists to explore how AI can support their research.
We gave Claude a prompt to search through a massive database of DNA sequences for interesting new examples of RTs. Our involvement was limited to the initial prompt and the lab work, while Claude agents combed through the database, investigated the distinct RT families, and used their own judgement to identify interesting candidates. After 21 hours spent searching this data by roughly 950 agents using 210 million tokens, one of the agents spotted something remarkable: a repeating pattern of DNA sequences that occurs next to the gene for an odd-looking RT. After further analysis and testing in our lab, we recognized that this pattern marked a previously uncharacterized enzyme system found in bacteriophages (the viruses that infect bacteria) that we call array-associated reverse transcriptases (ART).
Our work to understand the primary function of ARTs is ongoing. However, we think it is important to share such findings early, both to demonstrate Claude’s capabilities and to give the broader community insight into what we’re working on. We have released a pre-print ( here ) that discusses this in more detail.
About our lab
We are a team of scientists who have spent our careers exploring unusual proteins, and specialize in using computational approaches to systematically read DNA, interpret its evolution, and pick out biological systems for further characterization. Our research prior to joining Anthropic has helped to better understand the evolution and regulation of CRISPR systems, discover new enzymes for next-generation cell and gene therapies , and build tools for accelerating the identification of anomalies in DNA, such as human pathogenic variants. We are part of Anthropic’s life sciences organization, alongside teams whose work includes drug discovery, and training Claude in biology and chemistry.
Our lab, located in the Bay Area, looks like a typical molecular biology lab. We do research that involves only the lower-levels of the biosafety risk level (BSL-1 and BSL-2) and we do not handle pathogens that can infect humans. All of the lab work is performed by human scientists. Although we’ve experimented with using AI to accelerate lab work with initiatives like the Model Hardware Standard , this approach is less conducive to the sort of ad hoc workflows that are involved in our molecular biology research.
How we work
Many of our workflows involve having Claude search through the vast collection of DNA sequences associated with proteins without a known function. One typical pattern begins with a survey of a given protein family. Claude reads the relevant literature and reproduces the established results from public data to check its methods. It then searches for family members or genomic neighbors that fit no described system, and writes a short, human-readable report for each candidate that proposes a function and describes the evidence supporting its claims. In follow-up analyses, Claude critically evaluates the evidence—typically most candidates are eliminated at this stage. A survey may end with a single candidate worth testing, or with none.
When a candidate survives our review, we test it in the laboratory, expressing the protein in standard laboratory strains and characterizing it biochemically and structurally, with Claude helping to interpret the data. We do our work in Claude Science and Claude Code , the same tools available to any scientist, and sometimes with a harness of our own that coordinates many Claude sessions running in parallel.
Because Claude produces hypotheses so prolifically, the hypotheses themselves have become an object of study for us. With hundreds to thousands of candidate reports from a single campaign, we have been asking what distinguishes the proposals we judge worth testing from those we set aside. What we learn goes back into the instructions we give Claude and teaches it to mimic our own scientific taste.
Claude finds ART
In the past few years, researchers have discovered many more reverse transcriptases (RTs), most of them in bacteria, where they act as part of the immune system. Nearly all RT families were found by genomic analysis, or genome mining, which requires researchers to search sequence databases for genes that no one has characterized, notice the unusual ones, and work out what they do.
Claude agents gathered over 200,000 RTs, picked out 3,500 new candidate systems, and narrowed those to the 20 most-compelling candidates that they analyzed to produce human-readable reports. For an expert scientist, this type of analysis can take weeks to months of work.
During the course of its research, Claude noticed an unusual RT family and decided to examine it in greater detail. While combing through the raw DNA sequence near the RT, the agent exclaimed: “[The DNA next to the RT] is spectacular: I can see by eye a tandem repeat array … that's a CRISPR-like … repeat array?!”
The raw DNA Claude was reading when it detected a repeat pattern that no one had noticed
It then proceeded much as a human scientist would when faced with a potential discovery. It counted the repeats and measured their spacing, compared the layout with the known RT systems, and searched the literature for any previous report of the pattern. After a thorough analysis it was convinced that it had found a new biological system, and filed a report for human review.
The system it found, ART, is found mainly in bacteriophages and consists of three parts: the RT, a partner gene beside it, and a long array of evenly spaced DNA repeat sequences. The repeat layout resembles a CRISPR array, which holds a bank of different RNA sequences that make CRISPR-Cas systems programmable biotechnological tools. Our first experiments show that the ART array is also expressed as a set of distinct short RNAs, suggesting that something analogous may be at play for this system.
Further experiments are underway to determine how ART works, and we are sharing these early findings to show the community that Claude can autonomously detect anomalies and drive analyses to initiate biological discoveries.
You can find more detail in our technical report ( here ).
Work with us
We hope this work demonstrates the value of AI-driven hypothesis generation to the wider scientific community, and we would like to work with other scientists to extend this approach to a broad range of problems, in genomics and in other fields. If you have a proposal for a research question, we would like to hear from you.
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-09-24 | 10.26 | 13 | 入选 |
| 2026-09-24 | 12.45 | 13 | 入选 |