论文
arxiv:2610.01762
复制 Markdown
OneStreamer:在流式视频交互中统一感知、记忆与主动响应
发表于 10月1日
·
提交者
曾祥宇(Xiangyu Zeng)
于 10月2日提交
·
南京大学
作者:
曾祥宇
,
杨远东(Yuandong Yang)
,
张志秋(Zhiqiu Zhang)
,
朱雨涵(Yuhan Zhu)
,
李新浩(Xinhao Li)
,
司清一(Qingyi Si)
,
姚丁宇(Dingyu Yao)
,
马长莲(Changlian Ma)
,
陈浩然(Haoran Chen)
,
陈心宇(Xinyu Chen)
,
史彦松(Yansong Shi)
,
周俊豪(Junhao Zhou)
,
李一飞(Yifei Li)
,
张军(Jun Zhang)
,
秦传宇(Chuanyu Qin)
,
杨晨旭(Chenxu Yang)
,
余新磊(Xinlei Yu)
,
欧阳坤(Kun Ouyang)
,
邵雨辰(Yuchen Shao)
,
魏千山(Qianshan Wei)
,
周长海(Changhai Zhou)
,
高军(Jun Gao)
+2 位作者
摘要
流式视频大语言模型必须在证据与未来任务的相关性明确之前保留这些证据,并在获得足够证据时做出响应。挑战在于形成可复用的事实记忆,同时不损害实时感知能力。我们提出了 OneStreamer,它通过共享的主动生成过程联合学习独立于查询的证据记录与任务响应。其主动分层字幕记忆(PHCM)生成时间定位的局部细节描述以及已完成事件的摘要。流式字幕目标在训练期间监督对已观察视频前缀的解释。在推理阶段,模型生成的记录补充了最近的视觉窗口,提供了可复用的事实上下文,而无需重新访问历史视觉特征。主动状态转换学习(PSTL)通过在所有输出锚点保留监督信号并选择代表性的状态变化与状态持久化标记,减少了重复等待状态的主导地位。我们进一步开发了一种流式数据合成管道,将输出内容与时间对齐至可用证据。结合由此产生的流式字幕和问答以及清洗后的开源数据,我们构建了 OneStreamer-1M,这是一个覆盖广泛的流式视频交互数据集,包含超过一百万条记录,涵盖多样化的任务。我们的 4B 模型在所有八个评估的流式视频理解基准测试中取得了优于对比方法的最佳结果。消融研究表明,保留生成的字幕在提升历史问答能力的同时,并未降低实时感知性能。PSTL 也优于密集状态监督,而仅对 27.5% 的标注状态标记进行监督。综上所述,这些结果支持将主动生成作为连接流式视频交互中感知、记忆形成和及时响应的共享学习接口。
查看 arXiv 页面
查看 PDF
项目页面
GitHub
45
添加到合集
社区
Lanxingxuan
论文作者
论文提交者
约 14 小时前
大家好!我们分享了 OneStreamer,这是一个 4B 参数量的模型,旨在统一流式视频交互中的感知、记忆与主动响应。
核心思想是学习“记住什么”以及“何时响应”。OneStreamer 在视频展开过程中记录证据,此时未来的问题尚不明确,并在获得足够证据时做出响应。
主要亮点:
超越近期视觉上下文的记忆:主动分层字幕记忆(PHCM)将观察结果转化为带时间戳的字幕和事件摘要,在原始帧离开视觉窗口后,为后续问题保留证据。
学习何时响应:主动状态转换学习(PSTL)优于密集状态监督,而仅对 27.5% 的标注状态标记进行监督。
4B 规模下的强劲性能:OneStreamer 在所有八个评估基准上取得了综合最佳结果,涵盖感知、记忆和主动响应。
OneStreamer-1M:我们引入了一个包含超过一百万条流式视频交互记录的数据集,结合了合成的字幕和问答以及清洗后的现有数据。
📺 项目页面和视频演示
💻 代码
我们很乐意讨论该方法、数据和评估,并非常期待听到您的反馈!
查看翻译
🔥
2
❤️
1
🚀
1
👍
1
+
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞成票
145
+133
在您的智能体中获取此论文:
hf papers read 2610.01762
尚未安装最新的命令行界面?
引用此论文的模型
1
MCG-NJU/OneStreamer-4B
视频到文本
•
4B参数
•
约9小时前更新
•
5
引用此论文的数据集
1
MCG-NJU/OneStreamer-1M
预览
•
约13小时前更新
•
64
•
3
引用此论文的Spaces应用
1
🎬
hugging-apps/onestreamer
包含此论文的合集
2
agents
合集
4项内容
•
约12小时前更新
OneStreamer
合集
统一流式视频交互中的感知、记忆与主动响应
•
3项内容
•
约14小时前更新
Papers
arxiv:2610.01762
Copy markdown
OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
Published on Oct 1
·
Submitted by
Xiangyu Zeng
on Oct 2
·
Nanjing University
Authors:
Xiangyu Zeng
,
Yuandong Yang
,
Zhiqiu Zhang
,
Yuhan Zhu
,
Xinhao Li
,
Qingyi Si
,
Dingyu Yao
,
Changlian Ma
,
Haoran Chen
,
Xinyu Chen
,
Yansong Shi
,
Junhao Zhou
,
Yifei Li
,
Jun Zhang
,
Chuanyu Qin
,
Chenxu Yang
,
Xinlei Yu
,
Kun Ouyang
,
Yuchen Shao
,
Qianshan Wei
,
Changhai Zhou
,
Jun Gao
+2 authors
Abstract
Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.
View arXiv page
View PDF
Project page
GitHub
45
Add to collection
Community
Lanxingxuan
Paper author
Paper submitter
about 14 hours ago
Hi everyone! We’re sharing OneStreamer, a 4B model that unifies perception, memory, and proactive responses for streaming video interaction.
The core idea is to learn both what to remember and when to respond. OneStreamer records evidence as video unfolds, before a future question is known, and responds when sufficient evidence becomes available.
Key highlights:
Memory beyond the recent visual context: Proactive Hierarchical Caption Memory (PHCM) turns observations into timestamped captions and event summaries, preserving evidence for later questions after the original frames leave the visual window.
Learning when to respond: Proactive State Transition Learning (PSTL) outperforms dense state supervision while supervising only 27.5% of annotated state tokens.
Strong results at 4B: OneStreamer achieves the best aggregate results among the compared methods on all eight evaluated benchmarks, covering perception, memory, and proactive response.
OneStreamer-1M: We introduce a dataset with over one million streaming video interaction records, combining synthesized captions and QA with cleaned existing data.
📺 Project page and video demo
💻 Code
We’re happy to discuss the method, data, and evaluation, and would love to hear your feedback!
See translation
🔥
2
❤️
1
🚀
1
👍
1
+
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
145
+133
Get this paper in your agent:
hf papers read 2610.01762
Don't have the latest CLI?
Models citing this paper
1
MCG-NJU/OneStreamer-4B
Video-Text-to-Text
•
4B
•
Updated about 9 hours ago
•
5
Datasets citing this paper
1
MCG-NJU/OneStreamer-1M
Preview
•
Updated about 13 hours ago
•
64
•
3
Spaces citing this paper
1
🎬
hugging-apps/onestreamer
Collections including this paper
2
agents
Collection
4 items
•
Updated about 12 hours ago
OneStreamer
Collection
Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
•
3 items
•
Updated about 14 hours ago
首次收录 · 2026-10-03 · 10.95 分