Structured Pruning of Small Language Models: An Empirical Study on GQA-Aware Attention and MLP Compression with LoRA Recovery. International Journal of Advances in Soft Computing and its Applications , [S. l.], v. 18, n. 3, p. 113–139, 2026. DOI: 10.15849/ijasca.171. Disponível em: https://ijasca.zuj.edu.jo/index.php/IJASCA/article/view/171. Acesso em: 10 oct. 2026.