不同模型家族的LLM能否直接共享KV缓存,而无需接收端进行预填充?
我们的答案是HeteroFold。
我非常兴奋能分享我们的新论文:《面向异构多智能体LLM的免预填充跨家族KV缓存传输》!
近期关于免预填充KV缓存传输的研究表明,LLM可以重用之前计算好的上下文,而不是反复在接收端执行预填充操作。然而,现有方法主要集中于同一模型家族内或具有兼容分词器的模型。
我们提出了HeteroFold,这是一种用于免预填充跨家族KV缓存传输的方法,使得Llama、Qwen和Ministral等模型尽管在分词器、模型深度、KV结构和表示空间上存在差异,仍能直接共享KV缓存。
HeteroFold在不同分词器之间对齐令牌,并在不同架构的层之间进行对齐,将发送端的K/V状态映射到接收端的表示空间中,并校准传输后的缓存以保留接收端的注意力模式和输出。发送端和接收端均保持冻结状态,而学习到的映射关系允许接收端直接从传输后的缓存中解码,无需再次处理原始上下文。
在32K上下文长度下,HeteroFold的传输速度比原生接收端预填充快约10.7倍。我们评估了Llama、Qwen和Ministral之间所有六种传输方向的结果,在长上下文、短上下文和多智能体设置中均取得了优异表现。
最令我兴奋的是使KV缓存能够跨模型家族边界复用的可能性。与其让不同的LLM反复重新计算相同的上下文,异构模型可以直接彼此重用计算结果。
Can LLMs from different model families directly share KV caches, without receiver-side prefill?
Our answer is HeteroFold.
I am so excited to share our new paper: Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs!
Recent work on prefill-free KV cache transfer has shown that LLMs can reuse previously computed context instead of repeatedly performing receiver-side prefill. However, existing approaches have largely focused on models within the same family or with compatible tokenization.
We introduce HeteroFold, a method for prefill-free cross-family KV cache transfer that enables models such as Llama, Qwen, and Ministral to directly share KV caches despite differences in tokenizers, model depth, KV structure, and representation spaces.
HeteroFold aligns tokens across different tokenizers and layers across different architectures, maps the sender’s K/V states into the receiver’s representation space, and calibrates the transferred cache to preserve the receiver’s attention patterns and outputs. Both the sender and receiver remain frozen, and the learned mappings allow the receiver to directly decode from the transferred cache without processing the original context again.
At 32K context length, HeteroFold achieves about 10.7× faster transfer than native receiver prefill. We evaluate all six transfer directions among Llama, Qwen, and Ministral, with strong results across long-context, short-context, and multi-agent settings.
What excites me most is the possibility of making KV caches reusable across model-family boundaries. Instead of different LLMs repeatedly recomputing the same context, heterogeneous models can directly reuse computation from one another.
首次收录 · 2026-10-05 · 9.77 分