论文
arxiv:2609.24984
复制 Markdown
WorldCrafter:具有隐式 3D 感知记忆的一致性视频世界模型
发布于 9月21日
·
提交者
Wangbo Yu
于 9月22日
·
ARC Lab,腾讯
作者:
Wangbo Yu
,
Kunhao Liu
,
Wenbo Hu
,
Shenghai Yuan
,
Chaoran Feng
,
Haiyang Zhou
,
Yukun Huang
,
Yiran Wang
,
Wang Zhao
,
Yingmin Luo
,
Ying Shan
摘要
视频世界模型使得对动态环境的交互式探索成为可能,但在长视域和跨视角下难以保持对先前观察结果的一致性。我们提出了 WorldCrafter,这是一种视频世界模型,它学习了一种可通过相机查询的隐式 3D 感知记忆以实现此目的。其关键见解在于,让请求的视角决定如何将多视图证据压缩到视频生成器的有限令牌预算中。与世界生成器联合训练,记忆编码器和姿态条件读取模块在去噪之前将历史观察结果整合为一组固定的目标视图特定令牌,无需显式的基于深度的对应关系。通过将这种记忆与近期的时间上下文以及少步蒸馏相结合,WorldCrafter 能够从单张输入图像或文本提示开始实现流式场景探索。在静态和动态场景上的实验表明,在分钟级探索期间保持视觉质量的同时,长视域一致性和相机控制精度均获得了显著提升。
查看 arXiv 页面
查看 PDF
项目主页
GitHub
194
添加到合集
社区
Drexubery
论文提交者
约 19 小时前
我们介绍了 WorldCrafter,这是一种具有可通过相机查询的隐式 3D 感知记忆的视频世界模型,用于一致且长视域的场景探索。WorldCrafter 从单张图像或文本提示开始,结合历史观察结果、近期时间上下文和少步蒸馏,以支持流式探索,从而提高一致性和相机控制能力。
项目主页:https://drexubery.github.io/WorldCrafter
GitHub:https://github.com/TencentARC/WorldCrafter
查看翻译
🔥
3
+
回复
researchstudio-bot
约 13 小时前
这是 ResearchStudio 团队发送的自动消息。
我们为此论文创建了一个交互式 ResearchStudio Reel。它包括视觉海报、视频和博客,所有文件均可下载且格式可编辑。
打开 ResearchStudio Reel →
从 Hugging Face 下载所有文件
如果您觉得该 Reel 有帮助,请给此评论点赞!
想探索或为更多论文创建 Reels?请访问 ResearchStudio 演示。
查看翻译
回复
wbhu-tc
论文作者
约 8 小时前
演示线程:https://x.com/wbhu_cuhk/status/2102260092146196527
查看翻译
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞成票
112
+100
在您的代理中获取此论文:
hf papers read 2609.24984
没有最新的 CLI?
引用此论文的模型
2
TencentARC/WorldCrafter-Fast
图像到视频
•
更新于约 19 小时前
•
7
TencentARC/WorldCrafter-Base
图像到视频
•
14B
•
更新于约 19 小时前
•
26
•
3
引用此论文的数据集
0
无链接此论文的数据集
在数据集 README.md 中引用 arxiv.org/abs/2609.24984 以从此页面链接它。
引用此论文的 Spaces
2
🌍
hugging-apps/worldcrafter-demo
🌍
akhaliq/worldcrafter-demo-workflow
包含此论文的合集
0
无包含此论文的合集
将此论文添加到合集中以从此页面链接它。
Papers
arxiv:2609.24984
Copy markdown
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
Published on Sep 21
·
Submitted by
Wangbo Yu
on Sep 22
·
ARC Lab, Tencent
Authors:
Wangbo Yu
,
Kunhao Liu
,
Wenbo Hu
,
Shenghai Yuan
,
Chaoran Feng
,
Haiyang Zhou
,
Yukun Huang
,
Yiran Wang
,
Wang Zhao
,
Yingmin Luo
,
Ying Shan
Abstract
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.
View arXiv page
View PDF
Project page
GitHub
194
Add to collection
Community
Drexubery
Paper submitter
about 19 hours ago
We introduce WorldCrafter, a video world model with a camera-queryable implicit 3D-aware memory for consistent, long-horizon scene exploration. Starting from a single image or text prompt, WorldCrafter combines historical observations, recent temporal context, and few-step distillation to support streaming exploration with improved consistency and camera control.
Project page: https://drexubery.github.io/WorldCrafter
GitHub: https://github.com/TencentARC/WorldCrafter
See translation
🔥
3
+
Reply
researchstudio-bot
about 13 hours ago
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
See translation
Reply
wbhu-tc
Paper author
about 8 hours ago
Demo thread: https://x.com/wbhu_cuhk/status/2102260092146196527
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
112
+100
Get this paper in your agent:
hf papers read 2609.24984
Don't have the latest CLI?
Models citing this paper
2
TencentARC/WorldCrafter-Fast
Image-to-Video
•
Updated about 19 hours ago
•
7
TencentARC/WorldCrafter-Base
Image-to-Video
•
14B
•
Updated about 19 hours ago
•
26
•
3
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.24984 in a dataset README.md to link it from this page.
Spaces citing this paper
2
🌍
hugging-apps/worldcrafter-demo
🌍
akhaliq/worldcrafter-demo-workflow
Collections including this paper
0
No Collection including this paper
Add this paper to a collection to link it from this page.
首次收录 · 2026-09-23 · 10.95 分