论文
arxiv:2609.26780
复制 Markdown
SpeakerMem-R1:面向多方对话的以说话人为中心的双轨记忆系统
发布于 9月22日
·
提交者
HaoboZheng
于 9月24日
·
浙江大学
作者:
Haobo Zheng
,
Tan Tang
,
Yan Chen
,
Weijie Wang
,
Yingcai Wu
摘要
多方环境下的长期对话记忆不仅需要检索长对话中的相关内容,还必须区分谁说了什么、每句话涉及的对象是谁、个体之间如何相互感知、群体共享了哪些信息,以及状态随时间如何变化。近期关于多方对话基准测试的研究表明,现有的通用大语言模型(LLM)记忆系统往往会丢失人物和群体关系,或在整合分布在成员、群体和时间线上的线索时遇到困难。这些问题共同揭示了两大核心瓶颈:多方对话中的消息归属与关系理解,以及从交错历史中重建状态的能力。为解决这两个问题,我们提出了 SpeakerMem-R1:其双轨记忆系统将带有说话人标签的原始消息与组织成个人级和群体级视图的衍生状态分别存储,然后在查询时按实体、事件和时间结合两个轨道的证据。为了在结构化记忆构建过程中减少归属和更新错误,并支持本地部署,我们使用 SpeakerLevenshtein 奖励和说话人条件化的 GRPO 对 Writer-R1 进行了训练。在 GroupMemBench、SocialMemBench 和 EverMemBench 上,SpeakerMem-R1 的二分类准确率分别达到 47.9%、69.2% 和 61.9%。在 EverMind-AI 公开报告的 EverMemBench 排行榜上,我们取得了 62.33% 的成绩,是最新最先进框架中报告的最佳结果。在所有 1,986 个 LoCoMo 问题上,它也达到了 70.85% 的准确率,我们将这些题目用作两人长期对话的边界测试。在针对 305 道题目的受控评估中,强化学习(RL)将监督微调(SFT)Writer 的平均准确率从 57.38% 提升至 68.20%。我们同时报告了二分类准确率和 Token-F1,消融实验表明,在标准化评估接口下,原始消息和结构化轨道,以及个人级和群体级视图是互补的。
查看 arXiv 页面
查看 PDF
项目主页
GitHub
70
添加到收藏
社区
HaoboZheng
论文作者
论文提交者
大约 20 小时前
很高兴分享我们在 SpeakerMem-R1 上的工作!最近的基准测试揭示了一个令人惊讶的差距:通用记忆系统在多方对话中表现困难,在某些情况下甚至不如简单的 BM25 检索。这些局限性凸显了长期多方对话中的两个核心挑战:消息归属和状态重建。
我们推出了 SpeakerMem-R1,这是一个以说话人为中心的双轨记忆框架,旨在解决这些挑战。
🧠 两条互补的记忆轨道:保留带有说话人标签的原始消息,同时将衍生状态组织为个人级和群体级记忆。在查询时,结合两者以恢复跨人物、事件和时间的证据。
📝 学习编写更好的记忆:SpeakerLevenshtein 奖励和说话人条件化的 RL 将 Qwen2.5-3B Writer 的下游问答准确率从 57.38% 提升至 68.20%,在受控评估中检索和回答过程保持冻结。
📊 在三个多方基准测试中取得更强结果:SpeakerMem-R1 在 GroupMemBench、SocialMemBench 和 EverMemBench 上的准确率分别比最强的已评估记忆/检索基线高出 3.3、12.4 和 9.4 个百分点。在我们与 EverMind-AI 公开报告的 EverMemBench 排行榜的比较中,它取得了 62.33% 的准确率,领先于包括 EverOS(60.08%)和 RippleMem(54.75%)在内的对比系统。
项目主页:https://2022hpsk.github.io/SpeakerMemR1/
代码:https://github.com/2022hpsk/SpeakerMemR1
我们非常乐意听取您的想法并讨论这项工作!
查看翻译
🤗
4
🤯
2
+
回复
researchstudio-bot
大约 7 小时前
这是来自 ResearchStudio 团队的自动消息。
我们为本篇论文制作了一个互动式 ResearchStudio Reel。它包含一张可视化海报、一段视频和一个博客,所有文件均可下载并支持编辑格式。
打开 ResearchStudio Reel →
从 Hugging Face 下载所有文件
如果您觉得这个 Reel 有帮助,请给这条评论点赞!
想要探索或为更多论文创建 Reel?请访问 ResearchStudio 演示页面。
查看翻译
1 条回复
·
👍
1
🤝
1
+
HaoboZheng
论文作者
大约 5 小时前
出色的海报!令人惊叹的 ResearchStudio 团队~
查看翻译
❤️
1
+
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞同
77
+65
在您的智能体中获取此论文:
hf papers read 2609.26780
尚未安装最新版的 CLI?
引用此论文的模型
0
暂无链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.26780,即可从本页面链接该论文。
引用此论文的数据集
0
暂无链接此论文的数据集
在数据集的 README.md 中引用 arxiv.org/abs/2609.26780,即可从本页面链接该数据集。
引用此论文的 Spaces
0
暂无链接此论文的 Space
在 Space 的 README.md 中引用 arxiv.org/abs/2609.26780,即可从本页面链接该 Space。
包含此论文的资源库
0
暂无包含此论文的资源库
将此论文添加到资源库,即可从本页面链接它。
Papers
arxiv:2609.26780
Copy markdown
SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue
Published on Sep 22
·
Submitted by
HaoboZheng
on Sep 24
·
Zhejiang University
Authors:
Haobo Zheng
,
Tan Tang
,
Yan Chen
,
Weijie Wang
,
Yingcai Wu
Abstract
Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent studies on multi-party dialogue benchmarks show that existing general-purpose LLM memory systems tend to lose person and group relations or struggle to integrate clues distributed across members, groups, and time. Together, these issues reveal two core bottlenecks: message attribution and relational understanding in multi-party dialogue, and state reconstruction from interleaved histories. To address both, we propose SpeakerMem-R1: its dual-track memory stores speaker-labeled verbatim messages and derived states organized into person-level and group-level views, then combines evidence from both tracks by entity, event, and time at query time. To reduce attribution and update errors during structured memory construction while enabling local deployment, we train Writer-R1 with SpeakerLevenshtein and speaker-conditioned GRPO. On GroupMemBench, SocialMemBench, and EverMemBench, SpeakerMem-R1 achieves binary accuracies of 47.9%, 69.2%, and 61.9%, respectively. On the publicly reported EverMemBench leaderboard from EverMind-AI, we achieves 62.33%, the best reported result among the latest state-of-the-art frameworks. It also achieves 70.85% on all 1,986 LoCoMo questions, which we use as a two-person long-term conversation boundary test. In a controlled evaluation of 305 questions, RL raises the SFT Writer's mean accuracy from 57.38% to 68.20%. We report both binary accuracy and token-F1, and ablations show that the verbatim and structured tracks, as well as person-level and group-level views, are complementary under the standardized evaluation interface.
View arXiv page
View PDF
Project page
GitHub
70
Add to collection
Community
HaoboZheng
Paper author
Paper submitter
about 20 hours ago
Excited to share our work on SpeakerMem-R1! Recent benchmarks reveal a surprising gap: general-purpose memory systems struggle with multi-party conversations and can even underperform simple BM25 retrieval in some settings. These limitations highlight two core challenges: message attribution and state reconstruction in long-term multi-party dialogue.
We introduce SpeakerMem-R1, a speaker-centered dual-track memory framework designed to address these challenges.
🧠 Two complementary memory tracks: keep speaker-labeled messages verbatim while organizing derived states into person- and group-level memories. At query time, combine both to recover evidence across people, events, and time.
📝 Learning to write better memories: SpeakerLevenshtein rewards and speaker-conditioned RL improve a Qwen2.5-3B Writer from 57.38% to 68.20% downstream QA accuracy in a controlled evaluation, with retrieval and answering frozen.
📊 Stronger results across three multi-party benchmarks: SpeakerMem-R1 improves accuracy over the strongest evaluated memory/retrieval baselines by 3.3, 12.4, and 9.4 percentage points on GroupMemBench, SocialMemBench, and EverMemBench, respectively. In our comparison with EverMind-AI’s publicly reported EverMemBench leaderboard, it achieves 62.33% accuracy, leading the compared systems, including EverOS (60.08%) and RippleMem (54.75%).
Project page: https://2022hpsk.github.io/SpeakerMemR1/
Code: https://github.com/2022hpsk/SpeakerMemR1
We’d love to hear your thoughts and discuss the work!
See translation
🤗
4
🤯
2
+
Reply
researchstudio-bot
about 7 hours ago
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
See translation
1 reply
·
👍
1
🤝
1
+
HaoboZheng
Paper author
about 5 hours ago
Excellent poster!Amazing team ResearchStudio~
See translation
❤️
1
+
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
77
+65
Get this paper in your agent:
hf papers read 2609.26780
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.26780 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.26780 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.26780 in a Space README.md to link it from this page.
Collections including this paper
0
No Collection including this paper
Add this paper to a collection to link it from this page.
首次收录 · 2026-09-25 · 10.95 分