论文
arxiv:2609.39378
复制 Markdown
EgoTools:迈向真实世界第一人称视频中的以工具为中心的推理
发布于 9月30日
·
提交者
Shulin Tian
于 10月2日
·
Ropedia
作者:
Shulin Tian
,
Junsu Kim
,
Shuai Liu
,
Hao Li
,
Yujiao Shen
,
Sihan Li
,
Zhe Yang
,
Yeongon Kim
,
Feiyu Li
,
Jialin Wu
,
Yichi Zhang
,
Wenhui Wang
,
Runmao Yao
,
Yuhao Dong
,
Zhaoxi Chen
,
Fangzhou Hong
,
Antonino Furnari
,
Jingkang Yang
,
Hongyuan Zhu
,
Ziwei Liu
摘要
现实世界中的具身任务,从日常活动到专业流程,要求智能体在物理约束下行动,同时跟踪不断变化的物体和任务状态。工具使用处于此类任务的核心地位,因为许多日常和专业活动都是通过工具介导的。理解这些任务需要对可供性、手-工具-物体的几何关系、程序进展以及对目标对象的因果效应进行推理。然而,尽管在诸如图像描述和通用视频问答等以感知为导向的视频任务中表现强劲,当前的多模态视频模型在以工具为中心的具身推理方面仍然受限。这一方向的进展受到缺乏真实世界第一人称数据和诊断基准的限制。为了弥补这一空白,我们引入了 EgoTools,这是首个针对第一人称工具使用理解的综合性套件。它由两个互补的部分组成:EgoTools-Data,一个包含100小时以工具为中心的第一人称录制的庞大语料库,配有同步音频、密集描述、侧重推理的旁白以及补充的3D信息;以及 EgoTools-Bench,一个涵盖四个赛道的诊断基准测试,包含1,000个问答对,覆盖了从感知和几何到程序和因果推理的工具使用理解。实验结果表明,当前模型在将工具使用与视觉证据对应方面仍面临困难:Gemini-3.1-Pro 的整体准确率为 66.9%,但在“感知与定位”赛道上仅为 51.7%。除了评估之外,我们还验证了 EgoTools-Data 作为训练资源的价值。在完整的1,000题基准测试中,严格分离源视频的情况下,全监督微调使 Qwen3-VL-8B-Instruct 的性能从 50.0% 提升至 60.9%。综上所述,这些结果确立了 EgoTools 作为用于训练和诊断评估真实世界第一人称工具使用理解的统一资源。
查看 arXiv 页面
查看 PDF
项目主页
GitHub
7
添加到收藏
社区
shulin16
论文提交者
3天前
项目页:https://ropedia.github.io/egotools/
代码:https://github.com/Ropedia/EgoTools
数据:https://huggingface.co/datasets/ropedia-ai/egotools-data
查看翻译
回复
librarian-bot
约20小时前
这是来自图书管理员机器人(Librarian Bot)的自动消息。我发现以下论文与本文相似。
以下论文由 Semantic Scholar API 推荐
RoboChrono:一个用于流式任务理解的真实机器人基准测试 (2026)
CapMem:一个基于第一人称视频中基于描述的片段记忆的基准测试 (2026)
EgoMonth:一个用于长期时空记忆的第一人称视频月度级基准测试 (2026)
本地视频理解能否跨场景迁移?EgoGears 基准测试 (2026)
GST-Bench:视觉语言模型(VLMs)能否从视频中发展出全局空间意识?(2026)
ICM-Bench:具有长期记忆的多模态智能体中的人级身份推理 (2026)
MachEmbodied-U0:具身智能的统一理解与生成模型 (2026)
如果您觉得此评论有帮助,请点赞!
如果您希望获取 Hugging Face 上任何论文的推荐,请访问此空间
您可以通过在评论中标记它直接向 Librarian Bot 请求论文推荐:@librarian-bot recommend
查看翻译
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞成票
57
+45
在您的智能体中获取此论文:
hf papers read 2609.39378
没有最新的 CLI?
引用此论文的模型
0
无链接此论文的模型
在模型 README.md 中引用 arxiv.org/abs/2609.39378,以便从本页面链接到该论文。
引用此论文的数据集
0
没有数据集链接此论文
在数据集 README.md 中引用 arxiv.org/abs/2609.39378,以便从本页面链接到该论文。
引用此论文的 Spaces
0
没有 Space 链接此论文
在 Space README.md 中引用 arxiv.org/abs/2609.39378,以便从本页面链接到该论文。
包含此论文的集合
2
推理
集合
161 项
•
更新于 2 天前
•
6
工具
集合
12 项
•
更新于 2 天前
Papers
arxiv:2609.39378
Copy markdown
EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
Published on Sep 30
·
Submitted by
Shulin Tian
on Oct 2
·
Ropedia
Authors:
Shulin Tian
,
Junsu Kim
,
Shuai Liu
,
Hao Li
,
Yujiao Shen
,
Sihan Li
,
Zhe Yang
,
Yeongon Kim
,
Feiyu Li
,
Jialin Wu
,
Yichi Zhang
,
Wenhui Wang
,
Runmao Yao
,
Yuhao Dong
,
Zhaoxi Chen
,
Fangzhou Hong
,
Antonino Furnari
,
Jingkang Yang
,
Hongyuan Zhu
,
Ziwei Liu
Abstract
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning about affordances, hand-tool-object geometry, procedural progress, and causal effects on target objects. Yet despite strong performance on perception-oriented video tasks such as captioning and general video QA, current multimodal video models remain limited in this form of tool-centric embodied reasoning. Progress in this direction has been limited by the lack of real-world egocentric data and diagnostic benchmarks. To address this gap, we introduce EgoTools, the first comprehensive suite for egocentric tool-use understanding. It consists of two complementary components: EgoTools-Data, a large-scale corpus of 100 hours of tool-centric egocentric recordings with synchronized audio, dense captions, reasoning-heavy narrations, and supplementary 3D information; and EgoTools-Bench, a diagnostic benchmark of 1,000 QA pairs across four tracks that cover tool-use understanding from perception and geometry to procedure and causal reasoning. Experimental results show that current models still struggle to ground tool use in visual evidence: Gemini-3.1-Pro achieves 66.9% overall accuracy but only 51.7% on Perception & Grounding. Beyond evaluation, we validate EgoTools-Data as a training resource. On the full 1,000-question benchmark, full supervised fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9%, under strict source-video separation. Together, these results establish EgoTools as a unified resource for both training and diagnostic evaluation of real-world egocentric tool-use understanding.
View arXiv page
View PDF
Project page
GitHub
7
Add to collection
Community
shulin16
Paper submitter
3 days ago
Proj page: https://ropedia.github.io/egotools/
Code: https://github.com/Ropedia/EgoTools
Data: https://huggingface.co/datasets/ropedia-ai/egotools-data
See translation
Reply
librarian-bot
about 20 hours ago
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
RoboChrono: A Real Robot Benchmark for Streaming Task Understanding (2026)
CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video (2026)
EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory (2026)
Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark (2026)
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? (2026)
ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory (2026)
MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
57
+45
Get this paper in your agent:
hf papers read 2609.39378
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.39378 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.39378 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.39378 in a Space README.md to link it from this page.
Collections including this paper
2
Reasoning
Collection
161 items
•
Updated 2 days ago
•
6
Tool
Collection
12 items
•
Updated 2 days ago
首次收录 · 2026-10-05 · 9.77 分