A neural network can be accurate and still be too expensive to run where it is needed. Perhaps a device has limited memory, or a service needs to answer many requests quickly. One way to reduce the cost is pruning: removing parts of a model that contribute relatively little to its predictions.

That sounds like a straightforward deletion problem. In practice, removing one part can force changes elsewhere. Structurally Prune Anything, or SPA, tackles this bookkeeping problem so that pruning methods can work across different network designs, software frameworks, and stages of training.

Zeros are not the same as a smaller model

A model’s weights are commonly stored in large arrays. Unstructured pruning sets selected individual weights to zero. But an ordinary dense computation may still process the same-sized arrays. Exploiting all those zeros often needs specialized software or hardware support.

Structured pruning removes larger pieces, such as entire channels. A channel is one feature dimension: in a convolutional network, it can correspond to a learned feature map. Removing whole channels makes the arrays themselves smaller, reducing the arithmetic and storage required by subsequent operations.

The catch is that channels are connected. If one layer produces fewer features, the next layer must expect fewer inputs. In networks with branching paths, the consequences can spread farther than a single neighboring layer.

Why a local deletion becomes a graph problem

Neural networks contain many such relationships: additions, concatenations, grouped convolutions, and attention operations all impose rules on their inputs and outputs. Writing a custom pruning routine for every architecture is tedious and fragile.

SPA instead asks the network’s computational graph what must change together. This graph records operations, the data flowing between them, their shapes, and their parameters. It contains more information than a simple list of connected layers.

A shared representation, then a coordinated cut

SPA first converts a model into ONNX, a common format for representing machine-learning models. Different frameworks may describe their layers differently, but this conversion expresses them through standardized operators. SPA can then analyze those operators instead of depending on how the original framework implemented each layer.

Next comes mask propagation. Think of a mask as a marker placed on a channel we might remove. SPA passes that marker through the graph using rules for each operator. Those rules identify which other channels are coupled to the original one and must be handled consistently.

Once the dependencies are known, SPA organizes coupled channels into groups and estimates their importance. The score for a candidate removal combines information across its linked parameters, rather than judging one isolated weight. Lower-scoring candidates can then be removed together while keeping the network’s structure consistent.

SPA converts models from source frameworks to a shared graph, propagates channel dependencies, groups and scores them, prunes them, and exports the smaller model.
Follow the model through the diagram. A common representation separates the source framework from the pruning logic. Inside that representation, dependency analysis determines what must be removed together; importance estimation determines which candidates to remove. Select the figure to inspect the channel groups.

Finally, SPA changes the parameter arrays and their shapes in the ONNX model. The smaller model can be converted to a suitable framework for further training or use. Compatibility still depends on model conversion and support for the operators involved; “anything” is the design ambition, not a guarantee that every possible custom operation works automatically.

Why the time of pruning matters

The best way to judge importance depends on when pruning happens. Before training, the weights have not yet learned the task. After training, their values and behavior provide different evidence. If more training is available after pruning, the remaining network can recover from some of the damage.

SPA separates the structural question—what must be removed together?—from the scoring question—what is least important? That lets it adapt different importance criteria to three workflows:

  • Prune, then train. Start with a smaller structure and train it on the task.
  • Train, prune, then fine-tune. Remove parts of a trained model, then give it additional training to recover accuracy.
  • Train, then prune. Produce a smaller model without a subsequent fine-tuning stage.

For the last setting, the paper introduces Optimal Brain SPA (OBSPA). It selects whole coupled channels for removal and adjusts the remaining weights to help preserve each layer’s outputs. This is more careful than deleting channels and leaving all surviving weights untouched.

There is an important detail behind “data-free.” OBSPA can operate without the original training data; the strict data-free setting uses synthetic inputs sampled from a uniform distribution for calibration. It does not mean the procedure performs no calibration computation or uses no inputs at all.

What does a smaller model retain?

The paper evaluates 11 architectures and demonstrates pruning models originating in PyTorch, TensorFlow, MXNet, and JAX. The architecture experiments include convolutional networks, a vision transformer, and DistilBERT, a language model used here for sentiment classification.

Selected architecture results with pruning followed by fine-tuning. Image classification uses CIFAR-10; sentiment classification uses SST-2.
ModelOriginal accuracyPruned accuracyFLOPs reduction factor
ResNet-5093.26%93.42%2.13×
ViT-base95.35%96.10%2.05×
DistilBERT91.06%88.88%2.04×

FLOPs count arithmetic operations. A 2× reduction factor means roughly half as many operations, not necessarily half the latency on a particular device. Accuracy can improve after fine-tuning in some settings and decline in others, as these examples illustrate. The practical question is the trade-off for the model, task, and deployment you care about.

The takeaway: making a network smaller is not just about choosing dispensable weights. It is about understanding which parts are structurally tied together. SPA makes those dependencies explicit, so different pruning criteria can remove coordinated pieces of the network instead of breaking it one channel at a time.

The paper

Structurally Prune Anything: Any Architecture, Any Framework, Any Time

1CISPA Helmholtz Center for Information Security2Pruna AI3Department of Computer Science & Munich Data Science Institute, Technical University of Munich

* Equal contribution. # Work was done at Technical University of Munich.

arXiv 2024. Figures and experimental results above are from this paper; the explanatory examples are illustrative.

Abstract

Neural network pruning serves as a critical technique for enhancing the efficiency of deep learning models. Unlike unstructured pruning, which only sets specific parameters to zero, structured pruning eliminates entire channels, thus yielding direct computational and storage benefits. However, the diverse patterns for coupling parameters, such as residual connections and group convolutions, the diverse deep learning frameworks, and the various time stages at which pruning can be performed make existing pruning methods less adaptable to different architectures, frameworks, and pruning criteria. To address this, we introduce Structurally Prune Anything (SPA), a versatile structured pruning framework that can prune neural networks with any architecture, from any framework, and at any stage of training. SPA leverages a standardized computational graph and ONNX representation to prune diverse neural network architectures without the need for manual intervention. SPA employs a group-level importance estimation method, which groups dependent computational operators, estimates their importance, and prunes unimportant coupled channels. This enables the transfer of various existing pruning criteria into a structured group style. As a result, SPA supports pruning at any time, either before training, after training with fine-tuning, or after training without fine-tuning. In the context of the latter, we introduce Optimal Brain SPA (OBSPA), an algorithm that achieves state-of-the-art pruning results needing neither fine-tuning nor calibration data. In extensive experiments, SPA shows competitive to state-of-the-art pruning performance across various architectures, from popular frameworks, at different pruning times.

BibTeX

@misc{wang2024structurallypruneanythingarchitecture,
        title={Structurally Prune Anything: Any Architecture, Any Framework, Any Time},
        author={Xun Wang and John Rachwan and Stephan Günnemann and Bertrand Charpentier},
        year={2024},
        eprint={2403.18955},
        archivePrefix={arXiv},
        primaryClass={cs.LG},
        url={https://arxiv.org/abs/2403.18955},
  }