Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem
Deleting whole transformer blocks is one of the cheapest ways to speed up a large language model: because the model literally gets shorter, depth pruning delivers predictable inference speedups on top of memory savings, and it stacks cleanly with quantization, low-rank compression and other techniques. The hard part is deciding which blocks to cut. Remove the wrong ones and the model collapses, and the effect of removing any single block depends on which others are removed alongside it—the choices interact, making this a combinatorial problem rather than a ranking problem.
Most existing block-removal methods score each block on its own using magnitude, sensitivity, or "block influence" heuristics, then drop the ones that look least important. In physics terms these are mean-field methods: they treat each block as if its contribution were independent of the others, the way mean-field theory replaces a spin's neighbors with a single averaged field. A related shortcut is only ever removing a single consecutive run of blocks, which keeps the problem small but discards most of the search space. In reality blocks are not independent, any more than spins in a real magnet are—whether removing block 20 hurts depends on whether block 19 or block 24 was also removed, a coupling between decisions. Ignoring those couplings leaves quality on the table, especially in aggressive compression regimes, but searching over combinations directly is exponential and brute force is hopeless.
The paper, LLM Compression by Block Removal with Constrained Binary Optimization, takes the correspondence between interacting binary variables and spin systems literally. It reformulates block selection as a constrained binary optimization (CBO) problem that maps onto an Ising glass—a disordered spin system with all-to-all interactions and a fixed number of "up" spins. The energy of that spin system proves to be a strong, cheap proxy for how well the pruned model will actually score on benchmarks, allowing a huge number of candidate configurations to be ranked without benchmarking any of them, with hard instances handed off to the same classical and quantum-inspired solvers Multiverse uses elsewhere.
The payoff in the deep-compression regime is large: at 50% compression of Llama-3.3-70B-Instruct, the approach gains almost 23 percentage points on MMLU over the best competing block-removal method, suggesting physics-inspired combinatorial optimization can meaningfully outperform independent per-block heuristics when a lot of depth must be removed.