论文
arxiv:2609.23038
复制 Markdown
Spatial-Interactor:通过与可观测物理世界的交互学习空间推理
发布于 9月19日
·
提交者
Wenqi Zhang
于 9月24日
·
ZJU-OmniAI
作者:
Kaixiang Yao
,
Xu Wang
,
Miao Pan
,
Hu Xiyue
,
Weishi Wang
,
Daniel Dahlmeier
,
Jintao Chen
,
Yongliang Shen
,
Xuhong Zhang
,
Wenqi Zhang
摘要
空间推理对于视觉-语言模型(VLMs)理解和在物理世界中行动至关重要。在动态环境中的推理要求 VLM 感知由物体运动和视角变化引起的局部状态转换,并将这些转换整合到长轨迹中以维持更新的空间状态,然而现有的 VLM 在这两方面都受到限制。当前的空间训练主要关注关于物体属性和空间关系的静态问题,为状态转换提供的直接监督有限;相比之下,交互轨迹自然地连接了前一个观察、动作和后一个观察,为局部状态转换提供了直接监督,而完整的轨迹揭示了连续转换之间的依赖关系。因此,我们引入了 Spatial-Interactor,这是一个训练 VLM 通过交互对物理世界状态转换进行建模的框架,将这一学习过程组织为涵盖 L1 被动世界状态转换、L2 主动自我状态转换和 L3 长视界交互轨迹的三个级别课程。据此,我们从模拟和真实的交互轨迹中构建了来自空间交互的学习数据集(LSI-108K),任务与每个级别的目标保持一致。我们的两阶段训练策略对 L1 和 L2 应用监督微调(SFT)以进行局部转换建模,然后使用特权自蒸馏的在线策略蒸馏(OPD):一个教师分支提供片段级转换描述,监督学生的在线思维链(CoT),帮助学生学习在 L3 长轨迹上整合连续转换。在多个 VLM 和空间基准测试中的实验表明,局部转换建模和长视界整合方面均取得了持续的提升。
查看 arXiv 页面
查看 PDF
项目页面
GitHub
45
添加到收藏
社区
kagakouko
大约 18 小时前
•
此评论已被隐藏
zwq2018
论文作者
论文提交者
大约 17 小时前
理解场景不仅仅意味着识别可见的内容:它还要求推理通过交互空间如何变化。我们介绍了 Spatial-Interactor,这是一个用于学习局部状态转换和长视界空间推理的框架。我们的 LSI-108K 数据集支持三级课程,其中 SFT 用于局部转换建模,在线策略蒸馏用于长视界整合。代码、数据集资源以及四个模型检查点均可通过我们的项目页面获取。
查看翻译
回复
researchstudio-bot
大约 9 小时前
这是来自 ResearchStudio 团队的自动消息。
我们为这篇论文创建了一个交互式 ResearchStudio Reel。它包括一张视觉海报、一段视频和一个博客,所有文件均可下载且格式可编辑。
打开 ResearchStudio Reel →
从 Hugging Face 下载所有文件
如果您觉得这个 Reel 有帮助,请给此评论点赞!
想探索或为更多论文创建 Reels?请访问 ResearchStudio 演示。
查看翻译
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞成票
43
+31
在您的代理中获取此论文:
hf papers read 2609.23038
没有最新的 CLI?
引用此论文的模型
0
没有链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.23038 以从本页面链接它。
引用此论文的数据集
0
没有链接此论文的数据集
在数据集的 README.md 中引用 arxiv.org/abs/2609.23038 以从本页面链接它。
引用此论文的 Spaces
0
本文未链接至任何 Space。
在 README.md 文件中引用 arxiv.org/abs/2609.23038,即可从本页面链接到该论文。
包含此论文的集合
1
Spatial-Interactor
集合
Spatial-Interactor 的模型与数据:通过与可观测的物理世界进行交互来学习空间推理。
•
6 项内容
•
更新于 1 天前
•
2
Papers
arxiv:2609.23038
Copy markdown
Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
Published on Sep 19
·
Submitted by
Wenqi Zhang
on Sep 24
·
ZJU-OmniAI
Authors:
Kaixiang Yao
,
Xu Wang
,
Miao Pan
,
Hu Xiyue
,
Weishi Wang
,
Daniel Dahlmeier
,
Jintao Chen
,
Yongliang Shen
,
Xuhong Zhang
,
Wenqi Zhang
Abstract
Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial state, yet existing VLMs remain limited in both capabilities. Current spatial training primarily focuses on static questions about object attributes and spatial relations, providing limited direct supervision for state transitions; in contrast, interaction trajectories naturally connect a preceding observation, an action, and a subsequent observation, offering direct supervision for local state transitions, while complete trajectories reveal dependencies among consecutive transitions. We therefore introduce Spatial-Interactor, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories. Accordingly, we construct the Learning from Spatial Interaction dataset (LSI-108K) from simulated and real interaction trajectories, with tasks aligned with the objective of each level. Our two-stage training strategy applies Supervised Fine-Tuning (SFT) to L1 and L2 for local transition modeling, and On-Policy Distillation (OPD) then uses privileged self-distillation: a teacher branch given segment-level transition descriptions supervises the student's on-policy CoT, helping the student learn to integrate consecutive transitions over L3 long trajectories. Experiments across multiple VLMs and spatial benchmarks show consistent gains in local transition modeling and long-horizon integration.
View arXiv page
View PDF
Project page
GitHub
45
Add to collection
Community
kagakouko
about 18 hours ago
•
This comment has been hidden
zwq2018
Paper author
Paper submitter
about 17 hours ago
Understanding a scene means more than recognizing what is visible: it also requires reasoning about how space changes through interaction. We introduce Spatial-Interactor, a framework for learning local state transitions and long-horizon spatial reasoning. Our LSI-108K dataset supports a three-level curriculum, with SFT for local transition modeling and on-policy distillation for long-horizon integration. Code, dataset resources, and four model checkpoints are available through our project page.
See translation
Reply
researchstudio-bot
about 9 hours ago
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
43
+31
Get this paper in your agent:
hf papers read 2609.23038
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.23038 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.23038 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.23038 in a Space README.md to link it from this page.
Collections including this paper
1
Spatial-Interactor
Collection
Models and data for Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World.
•
6 items
•
Updated 1 day ago
•
2
首次收录 · 2026-09-25 · 10.95 分