Meta-Learning for Dataset Selection: An Adaptive Utility Scoring Framework with Interpretable Meta-Feature Analysis

Authors

  • Faisal AL-Saqqar Computer Science Department, Al al-Bayt University, Mafraq, Jordan.
  • Amjad H. Alkilan University of Arizona Global Campus, Chandler, Arizona, USA.
  • Mohammad I. Nusir CBM Integrated Software Inc, San Diego, CA, USA
  • Wael AlQassas Computer Science Department, Al al-Bayt University, Mafraq, Jordan.
  • Mohammad El-Bashir Computer Science Department, Al al-Bayt University, Mafraq, Jordan.

DOI:

https://doi.org/10.15849/ijasca.160

Keywords:

meta-learning, dataset selection, meta-features, classifier utility prediction, interpretable machine learning

Abstract

Dataset selection is commonly based on dataset size, prior experience, or computationally expensive trial-and-error evaluation. We treat that choice as a meta-learning problem instead. The Adaptive Dataset Utility Score (ADUS) ranks candidate datasets without training the downstream classifier portfolio on each candidate, using a scoring function developed from a separate collection of datasets whose classifier performance has already been measured. It reads 31 meta-features covering a dataset's statistics, structure, information content, and complexity, then combines them with weights that are either set by hand or fit to past results using non-negative least squares under a nested leave-one-dataset-out protocol. Across 20 datasets tested against five classifiers, both ADUS variants ranked the true best-performing dataset first in the aggregated leave-one-dataset-out ranking, a result no baseline or learned meta-learner matched, and the default-coefficient variant tracked realised utility at a Spearman correlation of 0.665, against 0.614 for a random forest meta-learner and 0.155 for ranking on sample count. Information-theoretic and complexity-based meta-features contributed most strongly to predictive performance, with the complexity group carrying the single largest share. Unexpectedly, the bell-shaped complexity term in the default formula improved raw ranking correlation when removed, even though it also reduced the fraction of top-performing datasets recovered in the top-3 -- a genuine trade-off documented in the component ablation rather than a reason to drop the term outright.

Downloads

Download data is not yet available.

Downloads

Published

2026-10-07

How to Cite

Meta-Learning for Dataset Selection: An Adaptive Utility Scoring Framework with Interpretable Meta-Feature Analysis (F. AL-Saqqar, A. H. . Alkilan, M. I. . Nusir, W. . AlQassas, & M. . El-Bashir, Trans.). (2026). International Journal of Advances in Soft Computing and Its Applications , 18(3), 87–112. https://doi.org/10.15849/ijasca.160
Total Downloads: 8

Google Scholar Link

Similar Articles

1-10 of 42

You may also start an advanced similarity search for this article.