Structured Pruning of Small Language Models: An Empirical Study on GQA-Aware Attention and MLP Compression with LoRA Recovery

Authors

  • ADNANE BAKKOU AIMCE Laboratory, École Nationale Supérieure d'Arts et Métiers (ENSAM), Université Hassan II de Casablanca, Casablanca, Morocco.
  • Badr Hirchoua AIMCE Laboratory, École Nationale Supérieure d'Arts et Métiers (ENSAM), Université Hassan II de Casablanca, Casablanca, Morocco.
  • Mouad Banane AIMCE Laboratory, École Nationale Supérieure d'Arts et Métiers (ENSAM), Université Hassan II de Casablanca, Casablanca, Morocco.
  • Naoufal Er-raji Ecole Marocaine des Sciences de l’Ingénieur, Casablanca, Morocco.

DOI:

https://doi.org/10.15849/ijasca.171

Keywords:

Large Language Models, Structured Pruning, Grouped-Query Attention, Low-Rank Adaptation, Model Compression, TinyLlama

Abstract

Large language models have shown strong performance across many natural language tasks, but their size makes deployment difficult on hardware with limited memory and compute. Compression techniques aim to reduce this cost while preserving as much quality as possible. This study investigates structured pruning applied to a small transformer-based language model that uses Grouped-Query Attention (GQA), a design now common in modern LLMs but rarely addressed explicitly by existing pruning methods. The research problem is therefore how to compress such an architecture along two structural axes, attention heads and feed-forward neurons, without breaking the GQA grouping that ties query heads to their shared key-value projections. We design a two-stage pruning pipeline, guided by data-dependent importance scores computed from attention entropy and mean activation magnitude, followed by a lightweight LoRA fine-tuning step that recovers part of the lost quality without adding inference cost. The pipeline is evaluated on TinyLlama-1.1B across seven configurations covering attention-only pruning, MLP-only pruning, their combination, and LoRA recovery. For TinyLlama-1.1B on WikiText-2 at the ratios we evaluate, the results show that feed-forward pruning offers a more favorable quality-compression trade-off than head pruning, that combining both forms of pruning produces a non-linear quality drop, that LoRA recovery partially restores the lost quality at no extra inference cost, and that inference latency is unchanged by pruning at these ratios, with every configuration including the unpruned baseline falling within a narrow band on our setup. These findings suggest practical implications for practitioners deploying small GQA-based language models, although every measurement was taken on a single GPU, a Tesla P100 for quality and a Tesla T4 for latency, and no CPU, mobile or embedded target was evaluated, and they point to layer-wise adaptive pruning and pruning-quantization combinations as promising directions for further reducing deployment cost.

Downloads

Download data is not yet available.

Downloads

Published

2026-10-09

How to Cite

Structured Pruning of Small Language Models: An Empirical Study on GQA-Aware Attention and MLP Compression with LoRA Recovery (A. BAKKOU, B. Hirchoua, M. Banane, & N. . Er-raji, Trans.). (2026). International Journal of Advances in Soft Computing and Its Applications , 18(3), 113–139. https://doi.org/10.15849/ijasca.171
Total Downloads: 0

Google Scholar Link

Similar Articles

1-10 of 47

You may also start an advanced similarity search for this article.