Return to Article Details Structured Pruning of Small Language Models: An Empirical Study on GQA-Aware Attention and MLP Compression with LoRA Recovery Download Download PDF