让大语言模型运行得更快的一种最廉价方法,也是最粗暴的方法:删除整个 Transformer 模块。因为模型在物理意义上变短了,
块移除(也称为深度剪枝)在节省内存的同时带来了可预测的推理加速,并且它可以与量化、低秩压缩以及其他技术无缝结合。难点在于决定
要切除哪些模块。如果切错了,模型就会崩溃;而且删除任何一个模块的影响取决于你同时删除了哪些其他模块,因此这些选择是相互作用的。这使得它成为一个组合问题,而非排序问题,而具有相互作用二元变量的组合问题正是自旋系统的物理学所擅长描述的对象。
我们最新的论文《通过受约束的二元优化进行 LLM 压缩》(LLM Compression by Block Removal with Constrained Binary Optimization)将这种对应关系付诸实践。我们将块选择重新表述为受约束的二元优化(CBO)问题,该问题直接映射到伊辛玻璃态(Ising glass),这是一种具有全连接相互作用和固定数量“向上”自旋的无序自旋系统。该自旋系统的能量被证明是剪枝模型在基准测试中实际得分的一个强大且廉价的代理指标,这意味着我们可以对大量候选配置进行排名,而无需对其中任何一个进行基准测试,并将困难实例交给我们在 Multiverse 其他地方使用的相同经典和类量子求解器。在深度压缩领域,回报巨大:在对 Llama-3.3-70B-Instruct 进行 50% 压缩时,我们在 MMLU 上比最佳竞争的块移除方法提高了近 23 个百分点。
为什么选择模块是一个多体问题
大多数现有的块移除方法单独评估每个模块,然后移除那些看起来最不重要的模块,使用幅度、敏感性或“块影响”启发式方法。在物理学术语中,这些是平均场方法:它们将每个模块视为其贡献独立于其他模块,就像平均场理论用单个平均场替换自旋的邻居一样。一个相关的捷径是仅删除连续的一段模块,这使问题保持较小,但丢弃了大部分搜索空间。
麻烦在于,模块并非相互独立,正如真实磁铁中的自旋并非相互独立一样。删除第 20 个模块是否损害模型,取决于你是否也删除了第 19 个或第 24 个模块,这是两个决策之间的相互作用,或耦合。随着模型变得更深且更异构,忽略这些耦合会导致质量损失,尤其是当你希望一次性删除大量模块时。你真正想要的是在考虑模块如何相互作用的情况下搜索模块组合,但组合的数量呈指数级增长,因此暴力求解看起来毫无希望。这正是统计物理工具大显身手的领域:具有成对耦合的指数级大的配置空间。
想法:将块选择转化为能量最小化问题
我们为每个 Transformer 模块附加一个二元变量:0 表示保留它,1 表示删除它,就像可以指向下或上的自旋一样。然后我们对模型损失关于这些变量进行二阶泰勒展开,从而产生(近似)海森矩阵。该海森矩阵的对角线表示每个模块单独的重要性;非对角线元素正是模块之间的成对耦合,即平均场方法所丢弃的多体物理。
这种重构将“我应该移除哪些块?”转化为一个清晰的优化问题:在恰好移除 N 个块中的 M 个块的约束下,找到一组 M 个块,使其移除后的能量 xᵀH⁰x 最小化。从数学上看,这是一个受约束的二元优化问题;从物理角度看,它是一个自旋玻璃(Ising glass),即一个具有守恒磁化强度(固定数量的被移除块扮演了固定总自旋的角色)的全耦合自旋系统。我们确立的关键性质是,这种能量是下游质量的强代理指标:自旋系统的低能态对应于高性能的剪枝模型。最小化能量和最大化基准分数变成了同一个搜索过程。
块的选择成为一个受约束的二元优化问题,等同于寻找自旋玻璃的低能态;每个解指出要删除 N 个块中的哪 M 个。右图:我们插入到每个块残差路径中以构建海森矩阵(Hessian)的耦合变量 α。来源:论文图 1。
这之所以具有实用性,原因在于成本。海森矩阵,即完整的耦合集合,仅通过在小规模校准数据集上进行前向和反向传播计算一次。此后,评估任何候选配置只需进行一次廉价的能量计算,无需运行实际模型,更不用说对其进行基准测试了。而且由于耦合不依赖于压缩目标,同一个海森矩阵可以重复使用来解决许多不同的 M 值。
求解方法:条件允许时求精确解,否则使用量子或类量子求解器
对于大多数模型而言,配置空间虽然庞大但仍可检查。由于计算一次能量非常便宜,我们在单个 GPU 上进行暴力搜索,检查多达数百亿个自旋配置。几百万个配置只需几秒钟;此处最难的可行案例是移除 Llama-3.3-70B 的 80 个块中的 8 个(约 290 亿个配置),耗时大约两天。
超出这个范围,精确方法就会失效,而将问题表述为自旋玻璃的优势在此时再次显现。在其等效的二次无约束二元优化(QUBO)形式中(约束被吸收进惩罚项),完全相同的任务可以交给为此类哈密顿量构建的高度优化的经典、量子和类量子求解器,即量子退火、QAOA、禁忌搜索以及专门的分支定界法。我们发现,开源的禁忌求解器能够在几秒钟内可靠地达到最低能量状态,即使在我们能够通过暴力搜索进行验证的最难案例中也是如此。因此,该方法可以扩展到无法枚举配置的模型,使用那些 squarely 处于 Multiverse 领域内的求解器。
这里有一个微妙但重要的观点,它与优化的常规思路背道而驰。通常,CBO 或退火求解器的评判标准是它是否找到了真正的基态(ground state)。我们实际上并不需要基态。我们需要的是快速生成少量良好低能态的方法,这是一个容易得多的目标,这就是为什么轻量级求解器对我们如此有效,以及为什么我们可以负担得起运行多个求解器的原因。
为何整个低能谱都至关重要
能量是质量的强代理指标,但并非完美指标,因此单一最低能量状态并不总是最好的模型。事实证明,这是一个特性而非缺陷:一旦哈密顿量设置完毕,读取基态和低激发态几乎是免费的,从而提供了一系列高质量的候选剪枝方案供尝试,而不是一个脆弱的单一答案。探索激发态而不仅仅是基态本身就是物理学研究的一个活跃领域,并且完美映射到从业者在此处真正需要的内容。
一个具体的例子:对于 Llama-3.1-8B-Instruct,在移除 16/32 个块时,大多数低能态倾向于移除模型末尾的块,这符合先前工作的预期。但是,第 17 个激发态是第一个提出移除模型开头附近块的配置,经过轻度重新训练后,该配置在多个基准测试中均优于基态(ground state)。这直接驳斥了“最佳剪枝是中间或靠后的连续块”这一常见假设,并展示了尊重问题的完整多体结构为何能带来收益。
左图:20 个最低能态各自移除的块(红色 = 被移除)。右图:第 17 个激发态,通过移除早期块,在重新训练后于多个基准测试中击败基态。最佳模型是一个激发态,而非基态。来源:论文图 2。
结果
在 Llama-3.1-8B-Instruct、Qwen3-14B 和 Llama-3.3-70B-Instruct 上,我们的方法(CBO)与最先进的块移除基线持平或更优,且随着压缩变得更加激进,这一优势差距进一步扩大。
最显著的胜利在于对 Llama-3.3-70B-Instruct 的深度压缩评估,且未进行重新训练。在最多移除 80 个块中的 24 个时,CBO 与 block influence(块影响力)大致持平。但在移除 32/80 和 40/80 个块时,它确立了决定性优势,在最深的设置下 MMLU 领先基线近 23 分,并在我们测试的所有基准测试中均优于基线。对于 Qwen3-14B,在移除 12/40 个块时,CBO 在 MMLU 上领先约 10 分。在较轻的压缩程度下,各方法表现相当,这是预期的:耦合(couplings)在你深入剪枝时最为重要。
Llama-3.3-70B-Instruct,无重新训练
移除块数 | MMLU
---|---
原始模型 | 0 | 82.2
CBO (ours) | 32 / 80 | 76.6
Block influence | 32 / 80 | 59.3
CBO (ours) | 40 / 80 | 76.9
Block influence | 40 / 80 | 54.0
在移除 40/80(50% 深度)时,CBO 将 MMLU 维持在 77 左右,而最强的基线则跌至 50 多分。来源:论文表 2。
其泛化能力超越稠密 Transformer
在现代异构架构上,块移除变得更加困难,因为不同类型的块是交错排列的,而 Ising 公式并不在意这一点:无论每个位置是什么类型的块,耦合就是耦合。为了对此进行压力测试,我们将该方法应用于 NVIDIA-Nemotron-3-Nano-30B-A3B-FP8,这是一个混合模型,它以非均匀模式交错排列 Mamba2、注意力(attention)和混合专家(MoE)层,且未进行任何重新训练。
我们的公式没有任何关于同质堆栈的假设,因此可以直接迁移。在移除 2–3 个 MoE 层或 2 个注意力层的情况下,CBO 找到的配置在 AIME25 和 GPQA 上击败了 block influence。结果还证实了这些混合模型中的冗余确实是真实存在的,但分布不均:某些专家层比其他层更易于舍弃,而该方法能够搜索耦合配置空间的能力正是其定位优质剪枝位置的关键。即使在这里,来自稠密模型的规律依然成立:最佳配置通常是一个激发态,而非基态。
为何这契合 Multiverse Computing 的理念
将混乱的机器学习问题重构为 Ising 哈密顿量,然后使用为物理学构建的经典和类量子优化机器来求解,这正是 Multiverse 的核心专长所在,这也正是贯穿我们压缩栈的相同直觉。块移除可以与该栈的其他部分(量化、低秩/SVD 压缩、宽度剪枝以及基于知识蒸馏的修复)组合,因此它融入的是一个更大的流水线,而非与之竞争。
想要获取完整的技术细节,包括泰勒展开推导、QUBO 映射、求解器基准测试、校准数据集消融实验以及完整的結果表格?请在 Hugging Face 上阅读完整论文,或联系我们的团队探讨将此应用于您自己的模型。代码已在 github.com/CompactifAI/Block_removal_through_constrained_binary_optimization 开源。
One of the cheapest ways to make a large language model faster is also one of the bluntest: delete whole transformer blocks. Because the model literally gets shorter,
block removal (also called depth pruning) buys predictable inference speedups on top of the memory savings, and it stacks cleanly with quantization, low-rank compression, and other techniques. The hard part is deciding
which blocks to cut. Remove the wrong ones and the model collapses; and the effect of removing any one block depends on which others you remove alongside it, so the choices interact. That makes it a combinatorial problem, not a ranking problem, and combinatorial problems with interacting binary variables are exactly what the physics of spin systems was built to describe.
Our latest paper, LLM Compression by Block Removal with Constrained Binary Optimization , takes that correspondence literally. We reformulate block selection as a constrained binary optimization (CBO) problem that maps directly onto an Ising glass , a disordered spin system with all-to-all interactions and a fixed number of "up" spins. The energy of that spin system turns out to be a strong, cheap proxy for how well the pruned model will actually score on benchmarks, which means we can rank a huge number of candidate configurations without benchmarking any of them, and hand the hard instances to the same classical and quantum-inspired solvers we use elsewhere at Multiverse. The payoff in the deep-compression regime is large: at 50% compression of Llama-3.3-70B-Instruct, we gain almost 23 percentage points on MMLU over the best competing block-removal method.
Why picking blocks is a many-body problem
Most existing block-removal methods score each block on its own, then remove the ones that look least important, using magnitude, sensitivity, or "block influence" heuristics. In physics terms these are mean-field methods: they treat each block as if its contribution were independent of the others, the way mean-field theory replaces a spin's neighbors with a single averaged field. A related shortcut is to only ever remove a single consecutive run of blocks, which keeps the problem small but throws away most of the search space.
The trouble is that blocks are not independent, any more than spins in a real magnet are. Whether removing block 20 hurts the model depends on whether you also removed block 19 or block 24, an interaction, or coupling , between the two decisions. As models get deeper and more heterogeneous, ignoring those couplings leaves quality on the table, especially when you want to remove a lot of blocks at once. What you really want is to search over combinations of blocks while accounting for how they interact, but the number of combinations grows exponentially, so brute force looks hopeless. This is precisely the regime, exponentially large configuration spaces with pairwise couplings, where the tools of statistical physics earn their keep.
The idea: turn block selection into an energy-minimization problem
We attach a binary variable to each transformer block: 0 means keep it, 1 means remove it, just like a spin that can point down or up. Then we do a second-order Taylor expansion of the model's loss with respect to those variables, which produces an (approximate) Hessian matrix. The diagonal of that Hessian is how much each block matters on its own; the off-diagonal entries are exactly the pairwise couplings between blocks, the many-body physics that mean-field methods throw away.
That reformulation turns "which blocks should I remove?" into a clean optimization: find the set of M blocks whose removal minimizes the energy xᵀH⁰x , subject to removing exactly M of the N blocks. Mathematically this is a constrained binary optimization problem; physically it is an Ising glass, an all-to-all coupled spin system with conserved magnetization (the fixed number of removed blocks plays the role of a fixed total spin). The key property we establish is that this energy is a strong proxy for downstream quality: low-energy states of the spin system correspond to high-performing pruned models. Minimizing energy and maximizing benchmark score become the same search.
Block selection becomes a constrained binary optimization problem, equivalent to finding low-energy states of an Ising glass; each solution says which M of N blocks to delete. Right: the coupling variable α we insert into each block's residual path to build the Hessian. Source: paper Figure 1.
The reason this is practical is cost. The Hessian, i.e. the full set of couplings, is computed just once, from forward and backward passes on a small calibration dataset. After that, evaluating any candidate configuration is a single cheap energy calculation, no need to run the actual model, let alone benchmark it. And because the couplings don't depend on the compression target, the same Hessian can be reused to solve for many different values of M .
Solving it: exact when you can, quantum or quantum-inspired when you can't
For most models the configuration space is large but still checkable. Because computing one energy is so cheap, we brute-force it on a single GPU, checking up to tens of billions of spin configurations. A few million take seconds; the hardest tractable case here, removing 8 of Llama-3.3-70B's 80 blocks (about 29 billion configurations), took roughly two days.
Beyond that the exact approach breaks down, and this is where casting the problem as an Ising glass pays off a second time. In its equivalent QUBO form (the constraint absorbed into a penalty term), the exact same task can be handed to the highly optimized classical, quantum, and quantum-inspired solvers built for this class of Hamiltonian, the machinery of quantum annealing, QAOA, tabu search, and specialized branch-and-bound. We find that an open-source tabu solver reliably reaches the lowest-energy states in seconds , even on the hardest cases we can verify against brute force. So the method scales to models where enumerating configurations is out of the question, using solvers that are squarely in Multiverse's domain.
There's a subtle but important point here, and it runs against the usual grain of optimization. Normally a CBO or annealing solver is judged by whether it finds the true ground state. We don't actually need the ground state. What we need is a fast way to generate a handful of good low-energy states, and that is a far easier bar, which is why lightweight solvers work so well for us and why we can afford to run several of them.
Why the whole low-energy spectrum matters
The energy is a strong proxy for quality, but not a perfect one, so the single lowest-energy state isn't always the best model. This turns out to be a feature, not a bug: once the Hamiltonian is set up, reading off the ground state and the low-lying excited states is essentially free, giving a spectrum of high-quality candidate prunings to try rather than one fragile answer. Exploring excited states, not just the ground state, is itself an area of active physics research, and it maps neatly onto what practitioners actually need here.
A concrete example: for Llama-3.1-8B-Instruct at 16/32 blocks removed, most of the top states cut blocks toward the end of the model, as prior work would expect. But the 17th excited state is the first to propose removing a block near the beginning of the model, and after light retraining that configuration outperforms the ground state across several benchmarks. That directly disproves the common assumption that the best pruning is one consecutive chunk of middle-or-late blocks, and it shows why respecting the full many-body structure of the problem pays off.
Left: which blocks each of the 20 lowest-energy states removes (red = removed). Right: the 17th excited state, which removes an early block, beats the ground state on several benchmarks after retraining. The best model is an excited state, not the ground state. Source: paper Figure 2.
Results
Across Llama-3.1-8B-Instruct, Qwen3-14B, and Llama-3.3-70B-Instruct, our method (CBO) is on par with or better than state-of-the-art block-removal baselines, and the gap widens as compression gets more aggressive.
The clearest win is deep compression of Llama-3.3-70B-Instruct, evaluated without retraining. Up to 24 of 80 blocks removed, CBO is roughly on par with block influence. But at 32/80 and 40/80, it pulls decisively ahead, with an almost 23-point MMLU advantage at the deepest setting, where it beats the baseline on every benchmark we tested. For Qwen3-14B at 12/40 removed, CBO leads MMLU by about 10 points. At lighter compression the methods are comparable, which is expected: the couplings matter most when you're cutting deep.
Llama-3.3-70B-Instruct, no retraining
Blocks removed
MMLU
Original
0
82.2
CBO (ours)
32 / 80
76.6
Block influence
32 / 80
59.3
CBO (ours)
40 / 80
76.9
Block influence
40 / 80
54.0
At 40/80 (50% depth), CBO holds MMLU near 77 while the strongest baseline falls to the mid-50s. Source: paper Table 2.
It generalizes beyond dense transformers
Block removal gets much harder on modern heterogeneous architectures, where different block types are interleaved, and the Ising formulation doesn't care: a coupling is a coupling regardless of what kind of block sits at each site. To stress-test that, we applied the method to NVIDIA-Nemotron-3-Nano-30B-A3B-FP8, a hybrid model that interleaves Mamba2, attention, and mixture-of-experts (MoE) layers in a non-uniform pattern, without any retraining.
Nothing about our formulation assumes a homogeneous stack, so it transfers directly. Removing 2–3 MoE layers or 2 attention layers, CBO finds configurations that beat block influence on AIME25 and GPQA. The results also confirm that redundancy in these hybrid models is real but unevenly distributed: some expert layers are far more disposable than others, and the method's ability to search the coupled configuration space is what locates the good cuts. Even here, the pattern from the dense models holds, the best configuration is often an excited state rather than the ground state.
Why this fits Multiverse Computing
Reframing a messy machine-learning problem as an Ising Hamiltonian, then solving it with the classical and quantum-inspired optimization machinery built for physics, is squarely in Multiverse's wheelhouse, it's the same instinct that runs through our compression stack. And block removal composes with the rest of that stack, quantization, low-rank/SVD compression, width pruning, and knowledge-distillation-based healing, so it slots into a larger pipeline rather than competing with it.
Want the full technical details, including the Taylor-expansion derivation, the QUBO mapping, the solver benchmarks, the calibration-dataset ablations, and the complete results tables? Read the full paper on Hugging Face , or get in touch with our team to talk about applying this to your own models. The code is open-sourced at github.com/CompactifAI/Block_removal_through_constrained_binary_optimization .
首次收录 · 2026-09-22 · 10.53 分