评估世界模型需要同时评估其生成世界的质量,以及在探索、交互和修改过程中的稳定性和响应能力。我们推出了 HappyWorld-Bench,这是一个全面的基准测试,用于评估生成的世界在智能体与其交互时是否保持可靠。我们的设计基于六个世界能力(W1-W6)的分层能力框架,从生成式构建到统一的世界建模,并在三个独立的评估赛道中实现:视频世界模型、空间世界模型和具身世界模型。HappyWorld-Bench 包含 1,138 个视频提示、300 个空间场景和 254 个具身测试用例。在所有三个赛道中,我们构建并运行了 HappyWorld-Arena,以组织人类 A/B 比较并得出模型级别的 Elo 评级,这些评级与旨在捕捉行为正确性的全新设计的自动化指标相辅相成。在此统一框架下,我们评估了 14 个视频世界模型、9 个空间系统和 8 个具身候选模型。结果表明,所有三个赛道仍存在可靠性差距:视频模型在长程生成和重新访问期间表现出一致性下降;空间模型最多仅能达到 70.14% 的放置准确率和 73.33% 的编辑执行率;而具身模型难以在多步动作中保持状态一致,并对改变的动作条件和物理规则做出精确响应。这些发现凸显了评估世界模型时不仅要看视觉质量,还要看状态一致性以及其对动作和干预响应的正确性。
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-09-27 | 9.77 | 15 | 入选 |
| 2026-09-26 | 10.22 | 22 | 未入选 |
| 2026-09-25 | 10.95 | 15 | 未入选 |