Submitted:
23 October 2023
Posted:
24 October 2023
Read the latest preprint version here
Abstract
Keywords:
1. Introduction
2. Related work
3. LLM Pruning with Relative Importance and Activations
3.1. Post-training Pruning: Preliminaries
3.2. Relative Importance: A New Pruning Metric

3.3. Incorporating Activations into Relative Importance
4. Turning into N:M Sparsity
4.1. N:M Structured Pruning: Formulation
4.2. Channel Permutation for Improved N:M Sparsity
- Step 1:
- Heuristic Channel Allocation.
- Step 2:
- Linear Sum Assignment.
- Remarks.
5. Experiments
5.1. Setup
- Tasks and Metrics.
- Baselines.
- Calibration Data.
5.2. Unstructured Pruning
- Main Results.
- Ablation Studies.
5.3. N:M Structured Pruning
- Main Results.
5.4. Running time analysis
- Pruning and permutation running time.
- Pruning time. We test each algorithm with 128 calibration data. For SparseGPT, the pruning time amounts to 5756 seconds (approximately 1.5 hours). In contrast, Wanda and RIA demonstrate significantly reduced runtimes, with times of approximately 611 seconds (approximately 10 minutes) and 627 seconds (approximately 10 minutes)
- Channel permutation time. For a comparative analysis of execution duration, we present the processing time of a single matrix constructed using our algorithm for the N:M sparsity. This is compared against the greedy method and its variants, which employ escaping strategies to circumvent getting trapped in local minima, as discussed in (Pool and Yu 2021). The results for different dimensions are provided in Table A3.
6. Discussion
7. Conclusion
Appendix A. Sensitivity test on calibration datasets
| Eval. dataset | PTB | wikitext2 | c4 | ||||||
| Calib. dataset | wikitext2 | c4 | PTB | wikitext2 | c4 | PTB | wikitext2 | c4 | PTB |
| Magnitude | 146.35 | 6.83 | 9.38 | ||||||
| Wanda | 68.49 | 69.70 | 63.58 | 5.85 | 5.97 | 5.89 | 8.43 | 8.30 | 8.25 |
| SparseGPT | 72.94 | 72.31 | 59.10 | 5.69 | 6.03 | 5.90 | 8.50 | 8.22 | 8.30 |
| RIA | 67.58 | 68.69 | 67.88 | 5.75 | 5.83 | 5.83 | 8.07 | 8.03 | 8.08 |
Appendix B. Weight Reconstruction
| wikitext2 | c4 | |||||
| wikitext2 | c4 | PTB | wikitext2 | c4 | PTB | |
| Magnitude | 6.38 | 9.38 | ||||
| Magnitude + rec | 5.84 | 6.07 | 6.17 | 8.48 | 8.30 | 8.62 |
| Wanda | 5.85 | 5.97 | 5.89 | 8.43 | 8.30 | 8.25 |
| Wanda+rec | 5.70 | 6.00 | 5.91 | 8.57 | 8.29 | 8.35 |
| SparseGPT | 5.69 | 6.03 | 5.90 | 8.50 | 8.22 | 8.30 |
| RIA | 5.75 | 5.83 | 5.83 | 8.07 | 8.03 | 8.08 |
| RIA+rec | 5.57 | 5.86 | 5.83 | 8.20 | 8.03 | 8.21 |
Appendix C. Implementation Details of Channel Permutation


Appendix D. Complexity analysis of RIA and Channel Permutation
| Greedy | 252.1 | 495.4 | 818.7 | 1563.6 |
| Greedy + 100 escapes | 845.3 | 1349.1 | 1896.4 | 3592.3 |
| Channel Permutation | 6.2 | 8.1 | 11.5 | 15.3 |
Appendix E. Hungarian Algorithm
References
- Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
- Evci, Utku, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen. 2020. Rigging the lottery: Making all tickets winners. In International Conference on Machine Learning, pp. 2943–2952. PMLR.
- Frantar, Elias and Dan Alistarh. 2022. Spdy: Accurate pruning with speedup guarantees. In International Conference on Machine Learning, pp. 6726–6743. PMLR.
- Frantar, Elias and Dan Alistarh. 2023. Sparsegpt: Massive language models can be accurately pruned in one-shot.
- Frantar, Elias, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323.
- Gao, Leo, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021, September. A framework for few-shot language model evaluation. [CrossRef]
- Han, Song, Jeff Pool, John Tran, and William Dally. 2015. Learning both weights and connections for efficient neural network. Advances in neural information processing systems 28. [CrossRef]
- Hassibi, Babak, David G Stork, and Gregory J Wolff. 1993. Optimal brain surgeon and general network pruning. In IEEE international conference on neural networks, pp. 293–299. IEEE. [CrossRef]
- Hoang, Duc NM, Shiwei Liu, Radu Marculescu, and Zhangyang Wang. 2022. Revisiting pruning at initialization through the lens of ramanujan graph. In The Eleventh International Conference on Learning Representations.
- Hubara, Itay, Brian Chmiel, Moshe Island, Ron Banner, Joseph Naor, and Daniel Soudry. 2021. Accelerated sparse neural training: A provable and efficient method to find n: m transposable masks. Advances in neural information processing systems 34, 21099–21111. [CrossRef]
- Kuhn, Harold W. 1955. The hungarian method for the assignment problem. Naval Research Logistics Quarterly 2(1-2), 83–97. [CrossRef]
- LeCun, Yann, John Denker, and Sara Solla. 1989. Optimal brain damage. Advances in neural information processing systems 2.
- Lee, Namhoon, Thalaiyasingam Ajanthan, and Philip HS Torr. 2018. Snip: Single-shot network pruning based on connection sensitivity. arXiv preprint arXiv:1810.02340.
- Li, Jiajun and Ahmed Louri. 2021. Adaprune: An accelerator-aware pruning technique for sustainable cnn accelerators. IEEE Transactions on Sustainable Computing 7(1), 47–60. [CrossRef]
- Lin, Ji, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2023. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978.
- Liu, Shiwei, Tianlong Chen, Xiaohan Chen, Zahra Atashgahi, Lu Yin, Huanyu Kou, Li Shen, Mykola Pechenizkiy, Zhangyang Wang, and Decebal Constantin Mocanu. 2021. Sparse training via boosting pruning plasticity with neuroregeneration. Advances in Neural Information Processing Systems 34, 9908–9922.
- Marcus, Mitchell P., Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993. Building a large annotated corpus of English: The Penn Treebank. Computational Linguistics 19(2), 313–330. [CrossRef]
- Merity, Stephen, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models.
- Mishra, Asit, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. 2021. Accelerating sparse deep neural networks. arXiv preprint arXiv:2104.08378.
- Mocanu, Decebal Constantin, Elena Mocanu, Peter Stone, Phuong H Nguyen, Madeleine Gibescu, and Antonio Liotta. 2018. Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science. Nature communications 9(1), 2383. [CrossRef]
- Pool, Jeff and Chong Yu. 2021. Channel permutations for n: M sparsity. Advances in neural information processing systems 34, 13316–13327.
- Raffel, Colin, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv e-prints.
- Sanh, Victor, Thomas Wolf, and Alexander Rush. 2020. Movement pruning: Adaptive sparsity by fine-tuning. Advances in Neural Information Processing Systems 33, 20378–20389.
- Sun, Mingjie, Zhuang Liu, Anna Bair, and J Zico Kolter. 2023. A simple and effective pruning approach for large language models. arXiv preprint arXiv:2306.11695.
- Tillet, Philippe, Hsiang-Tsung Kung, and David Cox. 2019. Triton: an intermediate language and compiler for tiled neural network computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, pp. 10–19.
- Touvron, Hugo, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
- Touvron, Hugo, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
- Xiao, Guangxuan, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning, pp. 38087–38099. PMLR.
- Yuan, Geng, Xiaolong Ma, Wei Niu, Zhengang Li, Zhenglun Kong, Ning Liu, Yifan Gong, Zheng Zhan, Chaoyang He, Qing Jin, et al. 2021. Mest: Accurate and fast memory-economic sparse training framework on the edge. Advances in Neural Information Processing Systems 34, 20838–20850.
- Zhang, Susan, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
- Zhang, Yuxin, Mingbao Lin, Zhihang Lin, Yiting Luo, Ke Li, Fei Chao, Yongjian Wu, and Rongrong Ji. 2022. Learning best combination for efficient n: M sparsity. Advances in Neural Information Processing Systems 35, 941–953.
- Zhang, Y, J Zhao, W Wu, A Muscoloni, and CV Cannistraci. 2023. Epitopological sparse ultra-deep learning: A brain-network topological theory carves communities in sparse and percolated hyperbolic anns.
- Zhou, Aojun, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li. 2021. Learning n: m fine-grained structured sparse neural networks from scratch. arXiv preprint arXiv:2102.04010.
- Zhu, Michael and Suyog Gupta. 2017. To prune, or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878.
| 1 |
https://huggingface.co/meta-llama, https://huggingface.co/facebook |




| Method | LLaMA 7b | LLaMA 13b | LLaMA 30b | LLaMA 65b | LLaMA2 7b | LLaMA2 13b | LLaMA2 70b | OPT 1.3b | OPT 13b |
| Dense | 5.68 | 5.09 | 4.77 | 3.56 | 5.47 | 4.88 | 3.32 | 14.62 | 10.13 |
| Magnitude | 17.28 | 20.22 | 7.54 | 5.90 | 16.02 | 6.83 | 5.36 | 1712 | 11561 |
| Wanda | 7.26 | 6.15 | 5.24 | 4.57 | 6.92 | 5.99 | 4.22 | 18.41 | 11.92 |
| SparseGPT | 7.24 | 6.20 | 5.32 | 4.57 | 6.99 | 6.10 | 4.25 | 27.00 | 11.18 |
| RIA (Ours) | 7.12 | 6.08 | 5.08 | 4.38 | 6.81 | 5.83 | 4.11 | 18.08 | 11.05 |
| LLaMA-13B | LLaMA-30B | |
| 20.22 | 7.55 | |
| 11.97 | 6.73 | |
| 7.80 | 5.55 | |
| 6.57 | 5.27 | |
| 6.14 | 5.13 | |
| 6.08 | 5.08 |
| Method | Hellaswag | BoolQ | ARC-C | MNLI | RTE | AVG |
| Dense | 64.77 | 83.70 | 54.44 | 45.81 | 67.87 | 63.32 |
| magnitude | 60.58 | 71.10 | 49.32 | 32.80 | 60.65 | 54.89 |
| wanda | 62.70 | 83.27 | 52.50 | 43.19 | 70.84 * | 62.50 |
| sparsegpt | 62.36 | 84.26 * | 53.07 | 40.29 | 70.76 * | 62.15 |
| RIA | 63.22 | 84.77 * | 52.56 | 42.80 | 71.48 * | 62.97 |
| Method | Unstructured 50% | 2:4 | 2:4+CP w/o LSA | 2:4+CP | 4:8 | 4:8+CP | |
| LLaMA2-13b (Dense 4.88) |
Magnitude | 6.83 | 8.74 | 8.89 | 8.72 | 7.32 | 7.52 |
| Wanda | 5.99 | 9.00 | 8.79 | 8.60 | 7.00 | 6.73 | |
| SparseGPT | 6.10 | 8.77 | 8.61 | 8.53 | 7.01 | 6.70 | |
| RIA (Ours) | 5.83 | 8.41 | 8.03 | 7.97 | 6.74 | 6.33 | |
| LLaMA2-70b (Dense 3.32) |
Magnitude | 5.36 | 6.76 | 6.77 | 6.73 | 5.89 | 6.03 |
| Wanda | 4.23 | 5.48 | 5.29 | 5.29 | 4.77 | 4.65 | |
| SparseGPT | 4.25 | 5.68 | 5.37 | 5.34 | 4.91 | 4.76 | |
| RIA (Ours) | 4.11 | 5.36 | 5.18 | 5.11 | 4.68 | 4.46 |
| Method | Hellaswag | BoolQ | ARC-C | MNLI | RTE | AVG |
| Dense | 64.77 | 83.70 | 54.44 | 45.81 | 67.87 | 63.32 |
| Wanda (2:4) | 57.35 | 81.44 | 46.01 | 37.69 | 68.59 * | 58.22 |
| Wanda (2:4+CP) | 59.37 | 84.50 * | 48.55 | 43.09 | 66.43 | 60.39 |
| Wanda (4:8+CP) | 60.86 | 82.73 | 49.94 | 40.15 | 67.87 | 60.51 |
| RIA (2:4) | 57.13 | 82.78 | 46.76 | 37.39 | 69.31 * | 58.68 |
| RIA (2:4+CP) | 58.48 | 85.14 * | 49.15 | 49.08 * | 68.95 * | 62.16 |
| RIA (4:8+CP) | 60.44 | 83.58 | 50.43 | 48.69 * | 70.04 * | 62.64 |
| Method | Q/K/V/Out | Up/Gate | Down |
| unstructured 50% (CPU) | 1.91× | 1.84× | 1.80× |
| 2:4 (CPU) | 1.91× | 1.84× | 1.80× |
| unstructured 50% (GPU) | 0.98× | 0.98× | 0.97× |
| 2:4 (GPU) | 1.64× | 1.65× | 1.62× |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).