论文
arxiv:2609.34563
复制 Markdown
重新思考潜在视觉推理:将潜在推理建立在视觉证据之上
发布于 9月28日
·
提交者
Xi Xiao
于 10月1日
·
亚马逊
作者:
Xi Xiao
,
Tianchen Zhao
,
Youngeun Kim
,
Zhuowei Li
,
Linghan Xu
,
Jiaye Wu
,
Zheng Zhang
,
Xiang Xu
,
Xuanbai Chen
,
Farhan Tejani
,
Jakub Zablocki
,
Julia Xu
,
Yifan Xing
摘要
潜在视觉推理(LVR)使多模态大语言模型(MLLMs)能够在连续的潜在标记中进行中间计算,而不是将每个推理步骤都用文字表达出来。然而,与文本思维链(CoT)不同,潜在推理无法直接观察,这使得监督潜在标记学习什么变得困难。在本工作中,我们首先对潜在标记的行为进行了彻底分析,并发现了一个“潜在证据-信用差距”:潜在标记对改变正确答案的图像扰动反应微弱。我们假设这个问题源于 GRPO 训练期间缺乏显式监督。这些发现表明,最终答案奖励在指导保留哪些视觉证据或如何在潜在标记之间分配信用方面提供的信息太少。为了弥合这一差距,我们提出了 ReaLVR,它将视觉证据监督引入模型自身的自由运行潜在轨迹中。ReaLVR 通过对比正确答案和模型生成的错误答案来确定需要更强监督的位置,并通过相关和不匹配的视觉证据来指定应保留的内容。在三个模型系列中,ReaLVR 始终优于所评估的 LVR 基线,在 Qwen2.5-VL-7B 上取得了 63.7% 的最高五项任务平均成绩。至关重要的是,我们是第一个在潜在空间中扩展视觉推理的研究者,表明我们的框架在高达 235B 的前沿模型规模上持续提供稳健的改进。进一步的分析显示,潜在标记的位置对问题更敏感,与相关视觉区域的对齐更强,并且对最受关注的潜在标记具有更大的固定上下文依赖性。
查看 arXiv 页面
查看 PDF
项目页面
GitHub
7
添加到收藏集
社区
MarkShaw99
论文提交者
大约 8 小时前
请查看最新的潜在视觉推理论文。
查看翻译
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞成票
200
+188
在您的代理中获取此论文:
hf papers read 2609.34563
没有最新的 CLI?
引用此论文的模型
0
无链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.34563 以从本页面链接它。
引用此论文的数据集
0
无链接此论文的数据集
在数据集的 README.md 中引用 arxiv.org/abs/2609.34563 以从本页面链接它。
引用此论文的 Spaces
0
无链接此论文的 Space
在 Space 的 README.md 中引用 arxiv.org/abs/2609.34563 以从本页面链接它。
包含此论文的收藏集
0
无包含此论文的收藏集
将此论文添加到收藏集以从本页面链接它。
Papers
arxiv:2609.34563
Copy markdown
Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
Published on Sep 28
·
Submitted by
Xi Xiao
on Oct 1
·
Amazon
Authors:
Xi Xiao
,
Tianchen Zhao
,
Youngeun Kim
,
Zhuowei Li
,
Linghan Xu
,
Jiaye Wu
,
Zheng Zhang
,
Xiang Xu
,
Xuanbai Chen
,
Farhan Tejani
,
Jakub Zablocki
,
Julia Xu
,
Yifan Xing
Abstract
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.
View arXiv page
View PDF
Project page
GitHub
7
Add to collection
Community
MarkShaw99
Paper submitter
about 8 hours ago
Please Check out the latest Latent Visual Reasoning paper.
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
200
+188
Get this paper in your agent:
hf papers read 2609.34563
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.34563 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.34563 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.34563 in a Space README.md to link it from this page.
Collections including this paper
0
No Collection including this paper
Add this paper to a collection to link it from this page.
首次收录 · 2026-10-02 · 10.95 分