这是一个问题,因为在我们都关心的高风险应用中,例如医学、工程和科学研究,不仅系统的结论至关重要,其得出结论的过程同样重要。当错误发生时——例如在医疗诊断和治疗中——我们需要能够 pinpoint 出哪里出了问题:是系统的推理有误,还是它依据了无效的证据,亦或是做出了错误的假设?
这就是为什么我最近离开了谷歌DeepMind的职位。我相信我们需要一种全新的机器推理方法——这种方法应借鉴AlphaGo的架构。AlphaGo会维护一份关于当前局势已知信息的记录:即棋局树(game tree)。这一数据结构包含了AlphaGo所考虑的所有变化路径和可能的未来,每一步走法和每个局面都标注了其神经网络做出的判断。随着推理的推进,AlphaGo会更新棋局树,并最终综合其中的信息以决定下一步的走法。
同样地,对于通用推理而言,系统应维护一个认知状态(epistemic state),用以表征系统视为确定的内容、其怀疑之处、已排除的可能性以及仍悬而未决的问题。推理可以被理解为一连串改变认知状态的动作序列,旨在推进知识并减少不确定性:推导后果、将问题分解为部分,以及——至关重要地——决定下一步要问什么问题、执行何种计算或进行哪项实验。
当然,开放世界中的推理比下围棋或国际象棋等棋类游戏要困难得多。在现实世界中,当前局势仅被部分知晓,可用动作的集合庞大且多变,而动作的后果则是随机的或未知的。
This is a problem because in the high-stakes applications we all care about, such as medicine, engineering, and scientific research, it matters not only what a system concludes but also how it arrives at its conclusion. When mistakes happen—for example, in medical diagnosis and treatment—we need to be able to pinpoint what went wrong: Was the system’s reasoning at fault, did it draw on invalid evidence, or did it make incorrect assumptions?
This is why I recently left my position at Google DeepMind. I believe we need a fresh approach to machine reasoning—one that draws on AlphaGo’s architecture. AlphaGo maintains a record of what it knows about a given position: the game tree. This data structure contains all the variations, the possible futures, that AlphaGo has considered, each move and position being annotated with judgments made by its neural networks. As its reasoning progresses, AlphaGo updates the game tree and eventually synthesizes the information in it to decide which move to make.
Similarly, for general reasoning a system should maintain an epistemic state that represents what the system holds as settled, what it doubts, what it has ruled out, which questions stay open. Reasoning can then be understood as a sequence of moves that change the epistemic state to advance knowledge and reduce uncertainty: deducing consequences, breaking problems into parts, and—crucially—deciding what question to ask, calculation to perform, or experiment to run next.
Of course, open-world reasoning is harder than playing a board game such as Go or chess. In the real world the current state of affairs is only partially known, the set of available actions is large and variable, and the consequences of actions are stochastic or unknown.
首次收录 · 2026-10-03 · 10.27 分