首页 博客
迈向可证明的联邦数据隐私学习
2026年10月2日
凯瑟琳·达利(Katharine Daly),软件工程师,丹尼尔·拉梅奇(Daniel Ramage),研究总监,Google Research
我们宣布推出一种新的联邦学习系统,该系统在将计算任务转移至服务器以提高训练速度、准确性和设备覆盖率的同时,提供可外部验证的隐私保障。
2017年,谷歌推出了联邦学习(Federated Learning, FL)这一机器学习技术,它能够在去中心化且私有的数据上训练模型。该技术已被用于赋能诸多日常实用功能,包括Gboard上的下词预测和智能撰写、Google Messages中的回复建议以及Android系统中的智能文本选择。
我们的联邦学习系统开发遵循四项核心隐私原则:(1)数据最小化,(2)数据匿名化,(3)透明度与控制权,以及(4)可验证性与可审计性。多年来在匿名化方面的研究与开发,通过矩阵分解DP-FTRL(MF-DP-FTRL)算法以及结合安全聚合的分布式差分隐私(Differential Privacy, DP),为生产环境中的模型带来了强有力的差分隐私保障。2025年,我们引入了一种基于这四项原则的联邦学习演进定义:
联邦学习(FL)是一种机器学习场景,多个实体(客户端)在服务提供者的协调下协作解决机器学习问题。一个完整的联邦学习系统应使客户端能够对其数据、允许访问其数据的负载集合以及这些负载的匿名化属性保持完全的控制权。联邦学习系统应为其数据由联邦学习客户端管理的相关用户提供适当的透明度与控制权。
在《迈向可证明的联邦数据隐私学习》一文中,我们宣布了我们下一代联邦学习系统的发布,该系统利用可信执行环境(Trusted Execution Environments, TEEs)提供完全可验证且可审计的数据匿名化保障。在TEE中运行的逻辑具备远程可证明性(第三方可以验证正在执行的逻辑),同时获得机密性(其内部状态无法被观察)和完整性(逻辑不会被干扰),当然这需以当前一代TEE的限制为前提。我们全新的基于TEE的联邦学习系统建立在TEE在单台机器层面提供的这些特性之上,构建了一个完全可验证的端到端联邦学习系统,实现了更强的隐私保障和更高的准确性。Gboard已经采用了这一新系统,并受益于比我们要之前的联邦学习系统快得多的计算时间。
利用单个TEE构建可验证私有的联邦学习系统
我们的基于TEE的联邦学习系统建立在我们早期在机密联邦分析和可证明隐私洞察方面工作的技术基础之上。
该系统协调四个核心操作概念:
数据上传:客户端设备在本地对训练样本进行加密并上传。这些设备会预先授权一个访问策略,该策略定义了允许处理此数据的可信执行环境(TEE)计算集合,且这些计算仅发布匿名化结果。客户端要求将访问策略发布到公共透明度日志中。
密钥管理与策略验证:由实施 RAFT 共识协议的 TEE 集群组成的密钥管理系统(KMS),仅向与访问策略中的计算相匹配的服务端 TEE 工作负载提供解密密钥。
工作负载执行:一个数据处理 TEE 执行实现训练循环的 Python 程序。这个“根”Tee 将可并行化的子任务委派给一组工作器 TEE 集群。分布式逻辑使用联邦语言(Federated Language)表达,这是一种源自 TensorFlow Federated 的开源、框架无关的编排语言,它曾驱动我们早期的联邦学习系统。训练循环会定期向数据分析师发布匿名化的模型权重。
容错恢复:Python 程序在每轮训练结束时保存由 KMS 加密的恢复状态。这可用于从间歇性的根节点或工作器故障中恢复,而不会泄露任何额外的隐私敏感信息。
该图表展示了数据在系统中的流动过程,以及验证系统上运行的工作负载隐私保证所需的步骤。
有关基于 TEE 的联邦学习系统设计的更多详细信息,请参阅我们的白皮书《迈向可证明隐私保护的联邦数据学习》(Toward provably private learning from federated data)。
基于 TEE 的联邦学习如何加强隐私保护
在早期的联邦学习系统中,设备数据被上传用于即时聚合,但外部观察者无法验证数据是否从未被记录或检查。后来,安全聚合允许通过密码学手段保护上传过程,但它与最先进的中心差分隐私(DP)保证不兼容。我们新的基于 TEE 的系统代表了我们持续努力完全消除对服务器运营商信任的下一个里程碑。
在我们新的基于 TEE 的联邦学习系统中,工作负载运营商只能看到指标和差分隐私模型权重。从设备收集的加密训练数据只能在运行访问策略中表示的 Python 训练程序的 TEE 内部解密和处理,且仅在上传后的有限时间内有效。
公共透明度日志与可重现构建
参与我们新的基于 TEE 的联邦学习系统的设备知道可能访问其上传数据的服务器工作负载的完整集合。代表这些潜在未来服务器工作负载的访问策略被发布到 Rekor(一个公共透明度日志),外部审计员能够观察这些日志,以跟踪设备可能参与的服务器端工作负载的完整集合。
我们联邦学习系统中使用的 KMS 和数据处理二进制文件可以从 Confidential Federated Compute Github 仓库中发布的开源代码进行可重现构建。
带有动态侧载的可验证执行
在早期的联邦学习系统中,运行在服务端的逻辑既无法由设备也无法由审计员验证,因此我们需要被信任以正确地向梯度总和添加随机噪声,从而提供差分隐私保护。
在我们新的基于 TEE 的联邦学习系统中,发布到 Rekor 的访问策略直接描述了表达联邦学习训练逻辑的 Python 程序。为了在保持可审计性的同时保护专有模型架构和数据预处理逻辑,我们的数据处理 TEE 支持在运行时将序列化信息侧载(sideloading)到 Python 程序中。只要所有与隐私相关的逻辑都硬编码在 Python 程序中,这种侧载功能就允许必须保持专有的逻辑在 TEE 中运行,同时仍提供强大的外部可验证隐私保证(关于侧信道观察的讨论见下文)。
Gboard 如何利用基于 TEE 的联邦学习
Gboard 已部署这套基于 TEE 的联邦学习系统,推出了具有更强隐私保障和更高准确性的英语及日语下一个词预测模型。这些改进可归因于新系统设计的几个方面。
缓解昼夜可用性限制
通过在执行服务器端训练工作负载之前收集所有设备上传的数据,我们不再需要担心设备的昼夜可用性波动影响训练进度。在服务器端执行时,我们可以在程序中动态计算最优的设备参与计划,并利用它来调整其他差分隐私(DP)参数。
该图表展示了通过优化设备参与计划,新 TEE 系统能够实现更强的隐私保障和/或更小的噪声乘数。这些曲线是在两个系统上对英语下一个词预测模型进行 5000 轮训练、每轮涉及 6500 台设备群体后得出的。
训练加速
过去,训练这些联邦学习模型每个可能需要 1-2 个月的时间,进度受限于设备可用性、设备端计算能力以及同一组设备资源在多个训练工作负载之间的竞争。在新的基于 TEE 的系统中,瓶颈已转移至服务器,跨多台机器的计算并行化使我们能够实现训练时间的显著加速,目前仅受限于 TEE 资源的可用性。
未来展望?
在我们新的基于 TEE 的联邦学习系统中,客户端梯度的计算被移至服务器端,从而克服了早期系统中与设备端计算资源相关的限制。这为使用联邦学习技术训练越来越大的模型铺平了道路。将 TEE 与加速器集成将在此类用例中发挥重要作用。
上述描述的系统不仅能够以可验证的方式执行联邦学习训练工作负载,还能执行任何可以用 Python 表达的工作负载。我们正在尝试在此基础设施上运行其他类型的工作负载,例如合成数据生成工作负载。另一个探索领域是:使用能够执行任意 Python 的数据处理 TEE,与我们专门用于大语言模型(LLM)推理等功能的其他数据处理 TEE 相结合。
这项工作朝着严格证明服务器端处理保护个人隐私迈出了重要一步。由于外部验证者能够检查我们在 Google 运行的确切代码,因此我们能够提供强有力的保证,即数据在服务器端的处理方式与描述完全一致。我们预计,未来的 TEE 硬件以及针对缓解侧信道观测的持续研究,将为动态加载的工作负载提供更深的防护,抵御恶意服务器端攻击。我们 anticipate(预期)像我们这样的系统未来可能会附带差分隐私算法和系统组件软件实现的完整正确性证明。
致谢
作者感谢为基础设施设计和实施做出贡献的合作者:Arun Ganesh、Brendan McMahan、Brett McLarnon、Chunxiang (Jake) Zheng、Emily Glanz、Maya Spivak、Michael Reneer、Nova Fallen、Stefan Dierauf、Suxin Guo、Timon Van Overveldt、Yu Xiao、Zachary Charles 和 Zachary Garrett。我们还要感谢支持 Gboard 集成的紧密合作伙伴:Haicheng Sun、Heng Su、Jianpeng Hou、Liyang Jiang、Noriyuki Takahashi、Wenzhi Mao、Xiaojuan Fang、Yanxiang Zhang、Yingjie Liu、Yuanbo Zhang 和 Yun Wang。这项工作得到了 Corinna Cortes、Shumin Zhai 和 Yossi Matias 的支持。我们还要感谢 Maysam Moussalem 对本文撰写的反馈。
标签:
移动系统
安全、隐私与滥用预防
软件系统与工程
快速链接
论文
分享
其他相关文章
2026 年 6 月 26 日
通过冻结多令牌预测在 Pixel 上加速 Gemini Nano 模型
机器智能 ·
移动系统 ·
自然语言处理
2026年6月10日
用于审计机器遗忘的新框架
算法与理论 ·
负责任的AI ·
安全、隐私与滥用预防
2026年5月27日
通过零信任聚合实现私有分析
安全、隐私与滥用预防
Home Blog
Toward provably private learning from federated data
October 2, 2026
Katharine Daly, Software Engineer, and Daniel Ramage, Research Director, Google Research
We announce a new Federated Learning system that provides externally verifiable privacy guarantees while shifting computation to the server to improve training speed, accuracy, and device coverage.
In 2017, Google introduced Federated Learning (FL) a machine learning technique that trains models across decentralized, private data. It has been used to power everyday helpful features, including next-word prediction and Smart Compose on Gboard, reply suggestions in Google Messages, and Smart Text Selection in Android.
Our FL systems development is guided by four essential privacy principles: (1) data minimization, (2) data anonymization, (3) transparency and control, and (4) verifiability and auditability. Years of research development on anonymization have led to strong differential privacy (DP) guarantees for production models through algorithms like matrix factorization DP-FTRL (MF-DP-FTRL) and distributed differential privacy coupled with Secure Aggregation. In 2025, we introduced an evolved definition of FL centered on these four principles:
Federated learning (FL) is a machine learning setting where multiple entities (clients) collaborate in solving a machine learning problem, under the coordination of a service provider. A complete FL system should enable clients to maintain full control over their data, the set of workloads allowed to access their data, and the anonymization properties of those workloads. FL systems should provide appropriate transparency and control to the users whose data is managed by FL clients.
In “Toward provably private learning from federated data”, we announce the next generation of our FL system, which leverages Trusted Execution Environments (TEEs) to provide fully verifiable and auditable data anonymization guarantees. Logic that runs in TEEs is remotely attestable (third parties can verify the logic that is being executed), and it also gains confidentiality (its internal state cannot be observed) and integrity (the logic cannot be disrupted), subject to current-generation TEE limitations. Our new TEE-based FL system builds on these properties which TEEs offer at the level of a single machine to form a fully verifiable end-to-end FL system that achieves stronger privacy guarantees and improved accuracy. Gboard has already adopted the new system and is benefiting from substantially faster compute times than our previous FL system.
Building a verifiably private FL system out of individual TEEs
Our TEE-based FL system builds on techniques developed in our earlier work on confidential federated analytics and provably private insights.
The system coordinates four core operational concepts:
Data upload: Client devices locally encrypt training examples and upload them. The devices pre-authorize an access policy, which is the set of TEE computations that will be allowed to process this data, and these computations only release anonymized results. Clients require access policies to be published to a public transparency log.
KMS and policy verification: The Key Management System (KMS), which consists of a cluster of TEEs implementing the RAFT consensus protocol, only gives decryption keys to server-side TEE workloads that match the computations in the access policy.
Workload execution: A data processing TEE executes a Python program that implements a training loop. This "root" TEE delegates parallelizable subtasks to a cluster of worker TEEs. Distributed logic is expressed using Federated Language, an open-source, framework-agnostic orchestration language derived from TensorFlow Federated, which powered our earlier FL system. The training loop periodically releases anonymized model weights to the data analyst.
Fault-tolerant recovery: The Python program saves a KMS-encrypted recovery state at the end of the training round. This can be used to recover from intermittent root or worker failures, without leaking any additional privacy-sensitive information.
This diagram shows both the flow of data through the system as well as the steps needed to validate the privacy guarantees of workloads running on the system.
For more details on the TEE-based FL system design please see our whitepaper, Toward provably private learning from federated data.
How TEE-based FL strengthens privacy
In earlier FL systems, device data was uploaded for the purpose of immediate aggregation, but there was no way for external observers to verify that the data was never logged or inspected. Later, Secure Aggregation allowed uploads to be protected cryptographically, but was not compatible with state-of-the-art central DP guarantees. Our new TEE-based system represents the next milestone in our ongoing effort to completely remove the need to trust the server operator.
In our new TEE-based FL system, only metrics and differentially private model weights are visible to workload operators. Encrypted training data collected from devices can only be decrypted and processed within TEEs running Python training programs represented in the access policies, and only for a limited amount of time after upload.
Public transparency log and reproducible builds
Devices participating in our new TEE-based FL system know the full set of server workloads that may access data they’re uploading. The access policies representing these potential future server workloads are published to Rekor, a public transparency log, and external auditors are able to observe these logs to track the full set of server-side workloads that devices could potentially be participating in.
The KMS and data processing binaries used in our FL system can be reproducibly built from open source code published in the Confidential Federated Compute Github repository.
Verifiable execution with dynamic sideloading
In our earlier FL systems, the logic running on the server could be verified neither by devices nor auditors, and thus we needed to be trusted to correctly add random noise to gradient sums to provide differential privacy.
In our new TEE-based FL system, the access policies that are published to Rekor directly describe the Python program that expresses the FL training logic. To protect proprietary model architectures and data preprocessing logic while preserving auditability, our data processing TEEs support sideloading serialized information into the Python program at runtime. As long as all privacy-relevant logic remains hardcoded in the Python program, this sideloading functionality allows logic that must remain proprietary to run in the TEE while still providing strong externally verifiable privacy guarantees (see below for a discussion on side-channel observations).
How Gboard is using TEE-based FL
Gboard has deployed this TEE-based FL system to launch English and Japanese next word prediction models with stronger privacy guarantees and improved accuracy. These improvements can be attributed to several aspects of the new system design.
Mitigating diurnal availability constraints
By collecting all device uploads before running the server-side training workload, we no longer have to worry about diurnal variations in device availability impacting training progress. At server-side execution time, we can dynamically calculate the optimal device participation schedule within the program, and can use it to tune other DP parameters.
This graph shows the stronger privacy guarantees and/or smaller noise multipliers that become possible with the new TEE-based system by optimizing device participation schedules. These curves were derived from training an English next word prediction model for 5000 rounds with cohorts of 6500 devices on both systems.
Training speedups
In the past, training these FL models could take 1-2 months each, with progress limited by device availability, on-device compute, and competition across multiple training workloads for the same set of device resources. With the new TEE-based system, bottlenecks have been moved to the server, and computation parallelization across many machines allows us to achieve significant speedups in training time, currently only limited by TEE resource availability.
What’s next?
In our new TEE-based FL system, computation of client gradients is shifted to the server, lifting limitations related to on-device compute resources that were present in earlier systems. This paves the way for training increasingly larger models using FL techniques. Integrating TEEs with accelerators will play an important role in such use cases.
The system we have described above is capable of executing not just FL training workloads in a verifiable manner, but also arbitrary workloads that can be expressed using Python. We are experimenting with running other types of workloads on this infrastructure, such as synthetic data generation workloads. Another area of exploration: using the data processing TEEs that execute arbitrary Python in combination with our other data processing TEEs that specialize in functionality such as LLM inference.
This work is a step toward rigorous proof that server side processing preserves individual privacy. With external verifiers able to inspect exactly what code we run at Google, we are able to offer strong assurances that data is processed server-side exactly as described. We expect future TEE hardware, along with ongoing research into mitigating side-channel observations, to offer deeper protections for dynamically loaded workloads against malicious server-side attacks. We anticipate that systems like ours may one day come with full proofs of correctness of the software implementations of the DP algorithms and system components.
Acknowledgements
The authors would like to thank the collaborators who contributed to the infrastructure design and implementation: Arun Ganesh, Brendan McMahan, Brett McLarnon, Chunxiang (Jake) Zheng, Emily Glanz, Maya Spivak, Michael Reneer, Nova Fallen, Stefan Dierauf, Suxin Guo, Timon Van Overveldt, Yu Xiao, Zachary Charles, and Zachary Garrett. We also thank close partners who supported the Gboard integration: Haicheng Sun, Heng Su, Jianpeng Hou, Liyang Jiang, Noriyuki Takahashi, Wenzhi Mao, Xiaojuan Fang, Yanxiang Zhang, Yingjie Liu, Yuanbo Zhang, and Yun Wang. This work was supported by Corinna Cortes, Shumin Zhai, and Yossi Matias. We additionally thank Maysam Moussalem for feedback on the writing of this post.
Labels:
Mobile Systems
Security, Privacy and Abuse Prevention
Software Systems & Engineering
QUICK LINKS
Paper
Share
Other posts of interest
JUNE 26, 2026
Accelerating Gemini Nano models on Pixel with frozen Multi-Token Prediction
Machine Intelligence ·
Mobile Systems ·
Natural Language Processing
JUNE 10, 2026
New framework for auditing machine unlearning
Algorithms & Theory ·
Responsible AI ·
Security, Privacy and Abuse Prevention
MAY 27, 2026
Private analytics via zero-trust aggregation
Security, Privacy and Abuse Prevention
首次收录 · 2026-10-03 · 10.59 分