半监督联邦学习(SSFL)利用教师模型在客户端的无标签数据上训练模型,并在服务器端使用少量带标签的种子数据集。自动语音识别(ASR)在此场景下尤为脆弱:伪标签错误会在输出序列中累积,并在多个训练轮次中叠加导致发散,从而使得其与全监督联邦学习之间存在巨大差距。我们表明,缩小这一差距取决于两个耦合的设计维度——教师模型(生成伪标签的模型)和锚点(服务器端在带标签数据上的更新,用于稳定训练)。在教师模型维度上,每个客户端的在线教师模型(即各客户端自身不断演化的模型)会自行发散,但一旦稳定下来,其表现就能匹配甚至超越广播式全局教师模型(在一轮内固定的单个服务器模型),在域内场景中具有决定性优势,在域偏移场景下也具有竞争力。随着种子数据集质量的提升以及在线教师模型优势的缩小,一种过渡型教师模型(在第 r 轮从全局切换为在线)的表现能匹配或优于两者。在锚点维度上,服务器必须在各轮次之间继续在带标签数据上进行训练——否则在线教师模型会发生漂移——这种交错机制对收敛性的影响比种子模型更大。这两个维度密不可分:激进的教师模型选择只有在锚点稳定了训练后才能奏效,而锚点的稳定性高度依赖于数据增强和批量大小——这些设置决定了服务器注入的输入噪声和梯度噪声的量。所需的稳定程度取决于具体领域,由种子数据的分布及其与客户端数据的重叠程度决定。这些发现为 ASR 训练中的 SSFL 提供了指导原则,在 11 组对比中有 9 组优于最强的先前方法,域内平均提升 20.8%,跨域平均提升 10.0%,从而缩小了与全监督联邦学习的差距。
Semi-supervised federated learning (SSFL) trains models on clientsâ unlabeled data using a teacher to generate pseudo-labels, with a small labeled seed dataset on the server. Automatic Speech Recognition (ASR) is particularly fragile here: pseudo-label errors compound across the output sequence and across training rounds into divergence, leaving a large gap to fully-supervised FL. We show that closing this gap turns on two coupled design axesâthe teacher (which model generates the pseudo-labels) and the anchor (the server-side updates on labeled data that stabilize training). On the teacher axis, a per-client online teacher (each clientâs own evolving model) diverges on its own, but once stabilized it matches or beats the broadcast global teacher (one server model, fixed within a round)âdecisively in-domain and competitively under domain shift. As the seed grows stronger and the online teacherâs advantage narrows, a transitioning teacher (global â online at round r) matches or beats both. On the anchor axis, the server must keep training on labeled data between roundsâotherwise the online teacher driftsâand this interleaving, more than the seed model, governs convergence. The two axes are inseparable: aggressive teacher choices pay off only once the anchor stabilizes training, which is highly sensitive to data augmentation and batch sizeâthe settings that govern how much input and gradient noise the server injects. How much stabilization is needed is domain-dependent, governed by the dispersion of the seed data and its overlap with client data. These findings yield guidelines for SSFL in ASR training, improving over the strongest prior method on 9 of 11 pairs, by 20.8% on average in-domain and 10.0% cross-domain, narrowing the gap to fully-supervised FL.
首次收录 · 2026-09-25 · 8.95 分