Enhancing Classification Performance on Imbalanced Datasets Through Complexity-Guided Oversampling With SMOTE

UDC.coleccionInvestigación
UDC.departamentoCiencias da Computación e Tecnoloxías da Información
UDC.grupoInvLaboratorio de Investigación e Desenvolvemento en Intelixencia Artificial (LIDIA)
UDC.institutoCentroCITIC - Centro de Investigación de Tecnoloxías da Información e da Comunicación
UDC.issue8
UDC.journalTitleExpert Systems
UDC.startPagee70352
UDC.volume43
dc.contributor.authorMorillo-Salas, José Luis
dc.contributor.authorBolón-Canedo, Verónica
dc.contributor.authorMorán-Fernández, Laura
dc.contributor.authorAlonso-Betanzos, Amparo
dc.date.accessioned2026-07-13T07:20:41Z
dc.date.available2026-07-13T07:20:41Z
dc.date.issued2026
dc.description.abstract[Abstract]: Improving classification performance on imbalanced datasets remains a challenging problem in machine learning. Synthetic oversampling techniques such as SMOTE are widely used to address class imbalance; however, their random interpolation strategy often ignores structural data properties, which may affect classifier generalisation. This work proposes a set of SMOTE-based strategies that guide the generation of synthetic samples in order to produce structurally simpler training datasets that are easier for classifiers to learn. The first strategy generates more dispersed (outer) synthetic samples to increase class separability with minimal computational overhead. The second and main contribution, SMOTE-Complex, formulates synthetic sample selection as an explicit optimisation process that minimises measurable training dataset complexity. A clustering-based variant, SMOTE-Complex-Clustering, reduces computational cost by restricting optimisation to feature subspaces while preserving most structural and predictive benefits. The underlying hypothesis is that reducing the structural complexity of the training data can lead to improved predictive behaviour. Extensive experiments on binary and multiclass datasets, using multiple classifiers and complementary evaluation metrics, provide empirical support for this hypothesis across diverse structural conditions. The results indicate moderate but stable improvements in structurally favourable scenarios—particularly binary and moderately complex problems—without systematic degradation in imbalance-sensitive metrics, while the clustering-based refinement offers a scalable trade-off between optimisation strength and computational efficiency.
dc.description.sponsorshipThis work has been supported by the National Plan for Scientific and Technical Research and Innovation of the Spanish Government (Grant PID2023-147404OB-I009), and by the Ministry for Digital Transformation and Civil Service and ‘Next-GenerationEU’/PRTR under Grant TSI-100925-2023-1. CITIC, as a center accredited for excellence within the Galician University System and a member of the CIGUS Network, receives subsidies from the Department of Education, Science, Universities, and Vocational Training of the Xunta de Galicia. Additionally, it is co-financed by the EU through the FEDER Galicia 2021-27 operational program (Ref. ED431G 2023/01). Grant ED431C 2022/44 funded by Xunta de Galicia.
dc.description.sponsorshipXunta de Galicia; ED431G 2023/01
dc.description.sponsorshipXunta de Galicia; ED431C 2022/44
dc.identifier.citationMorillo-Salas, J. L., V. Bolón-Canedo, L. Morán-Fernández, and A. Alonso-Betanzos. 2026. “ Enhancing Classification Performance on Imbalanced Datasets Through Complexity-Guided Oversampling With SMOTE.” Expert Systems 43, no. 8: e70352. https://doi.org/10.1111/exsy.70352
dc.identifier.doi10.1111/exsy.70352
dc.identifier.issn1468-0394
dc.identifier.urihttps://hdl.handle.net/2183/48859
dc.language.isoeng
dc.publisherWiley
dc.relation.projectIDinfo:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2021-2023/PID2023-147404OB-I00/ES/APRENDIZAJE AUTOMATICO FRUGAL: POTENCIANDO LA IA EN ENTORNOS CON RECURSOS LIMITADOS PARA LOS DESAFIOS DEL MUNDO REAL
dc.relation.projectIDinfo:eu-repo/grantAgreement/MTDPF//TSI-100925-2023-1/ES/CÁTEDRA UDC-INDITEX DE IA EN ALGORITMOS VERDES
dc.relation.urihttps://doi.org/10.1111/exsy.70352
dc.rightsAttribution 4.0 Internationalen
dc.rights.accessRightsopen access
dc.rights.urihttp://creativecommons.org/licenses/by/4.0/
dc.subjectClassification
dc.subjectComplexity measures
dc.subjectImbalanced datasets
dc.subjectSMOTE
dc.titleEnhancing Classification Performance on Imbalanced Datasets Through Complexity-Guided Oversampling With SMOTE
dc.typejournal article
dc.type.hasVersionVoR
dspace.entity.typePublication
relation.isAuthorOfPublicationc114dccd-76e4-4959-ba6b-7c7c055289b1
relation.isAuthorOfPublicationdfd64126-0d31-4365-b205-4d44ed5fa9c0
relation.isAuthorOfPublicationa89f1cad-dbc5-471f-986a-26c021ed4a95
relation.isAuthorOfPublication.latestForDiscoveryc114dccd-76e4-4959-ba6b-7c7c055289b1

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
BolonCanedo_Veronica_2026_Enhancing_Classification_Performance_on_Imbalanced_Datasets.pdf
Size:
2.91 MB
Format:
Adobe Portable Document Format