1.
Structured Pruning of Small Language Models: An Empirical Study on GQA-Aware Attention and MLP Compression with LoRA Recovery. Int. J. Adv. Soft Comput. Appl. [Internet]. 2026 Oct. 9 [cited 2026 Oct. 10];18(3):113–139. Available from: https://ijasca.zuj.edu.jo/index.php/IJASCA/article/view/171