Use this link to cite:
https://hdl.handle.net/2183/48859 Enhancing Classification Performance on Imbalanced Datasets Through Complexity-Guided Oversampling With SMOTE
Loading...
Identifiers
Publication date
Authors
Advisors
Other responsabilities
Journal Title
Bibliographic citation
Morillo-Salas, J. L., V. Bolón-Canedo, L. Morán-Fernández, and A. Alonso-Betanzos. 2026. “ Enhancing Classification Performance on Imbalanced Datasets Through Complexity-Guided Oversampling With SMOTE.” Expert Systems 43, no. 8: e70352. https://doi.org/10.1111/exsy.70352
Type of academic work
Academic degree
Abstract
[Abstract]: Improving classification performance on imbalanced datasets remains a challenging problem in machine learning. Synthetic oversampling techniques such as SMOTE are widely used to address class imbalance; however, their random interpolation strategy often ignores structural data properties, which may affect classifier generalisation. This work proposes a set of SMOTE-based strategies that guide the generation of synthetic samples in order to produce structurally simpler training datasets that are easier for classifiers to learn. The first strategy generates more dispersed (outer) synthetic samples to increase class separability with minimal computational overhead. The second and main contribution, SMOTE-Complex, formulates synthetic sample selection as an explicit optimisation process that minimises measurable training dataset complexity. A clustering-based variant, SMOTE-Complex-Clustering, reduces computational cost by restricting optimisation to feature subspaces while preserving most structural and predictive benefits. The underlying hypothesis is that reducing the structural complexity of the training data can lead to improved predictive behaviour. Extensive experiments on binary and multiclass datasets, using multiple classifiers and complementary evaluation metrics, provide empirical support for this hypothesis across diverse structural conditions. The results indicate moderate but stable improvements in structurally favourable scenarios—particularly binary and moderately complex problems—without systematic degradation in imbalance-sensitive metrics, while the clustering-based refinement offers a scalable trade-off between optimisation strength and computational efficiency.
Description
Editor version
Rights
Attribution 4.0 International








