Structured Pruning of Small Language Models: An Empirical Study on GQA-Aware Attention and MLP Compression with LoRA Recovery
DOI:
https://doi.org/10.15849/ijasca.171Keywords:
Large Language Models, Structured Pruning, Grouped-Query Attention, Low-Rank Adaptation, Model Compression, TinyLlamaAbstract
Large language models have shown strong performance across many natural language tasks, but their size makes deployment difficult on hardware with limited memory and compute. Compression techniques aim to reduce this cost while preserving as much quality as possible. This study investigates structured pruning applied to a small transformer-based language model that uses Grouped-Query Attention (GQA), a design now common in modern LLMs but rarely addressed explicitly by existing pruning methods. The research problem is therefore how to compress such an architecture along two structural axes, attention heads and feed-forward neurons, without breaking the GQA grouping that ties query heads to their shared key-value projections. We design a two-stage pruning pipeline, guided by data-dependent importance scores computed from attention entropy and mean activation magnitude, followed by a lightweight LoRA fine-tuning step that recovers part of the lost quality without adding inference cost. The pipeline is evaluated on TinyLlama-1.1B across seven configurations covering attention-only pruning, MLP-only pruning, their combination, and LoRA recovery. For TinyLlama-1.1B on WikiText-2 at the ratios we evaluate, the results show that feed-forward pruning offers a more favorable quality-compression trade-off than head pruning, that combining both forms of pruning produces a non-linear quality drop, that LoRA recovery partially restores the lost quality at no extra inference cost, and that inference latency is unchanged by pruning at these ratios, with every configuration including the unpruned baseline falling within a narrow band on our setup. These findings suggest practical implications for practitioners deploying small GQA-based language models, although every measurement was taken on a single GPU, a Tesla P100 for quality and a Tesla T4 for latency, and no CPU, mobile or embedded target was evaluated, and they point to layer-wise adaptive pruning and pruning-quantization combinations as promising directions for further reducing deployment cost.
Downloads
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 ADNANE BAKKOU, Badr Hirchoua, Mouad Banane, Naoufal Er-rajiCopyright © The Author(s).
Articles published in the International Journal of Advances in Soft Computing and its Applications (IJASCA) are licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
This license permits anyone to copy, redistribute, remix, transform, and build upon the material for any purpose, including commercial use, provided appropriate credit is given to the original author(s), a link to the license is provided, and any modifications are indicated.
Authors retain the copyright of their published work and grant the journal right of first publication, with the work simultaneously licensed under the terms above.
Link
