Use this link to cite:
https://hdl.handle.net/2183/49397 Selección de características basada en información mutua para datos positive-unlabeled
Loading...
Identifiers
Publication date
Authors
Vila Vivero, Verónica
Other responsabilities
Universidade da Coruña. Facultade de Informática
Journal Title
Bibliographic citation
Type of academic work
Academic degree
Abstract
[Resumen] En muchos problemas reales de aprendizaje automático, obtener etiquetas para todos los ejemplos resulta inviable. El escenario Positive-Unlabeled (PU) modela esta situación: solo se dispone de positivos etiquetados, mientras que el resto permanece sin etiquetar, pudiendo contener tanto negativos reales como positivos ocultos. Tratar directamente los ejemplos no etiquetados como negativos introduce un sesgo sistemático que afecta a la selección de características. Este trabajo propone un método de selección de características basado en información mutua adaptado al escenario PU, que corrige dicho sesgo estimando la probabilidad real de pertenencia a la clase positiva mediante el marco probabilístico de Elkan y Noto. La propuesta se evalúa experimentalmente sobre nueve conjuntos de datos de distinta naturaleza y dimensionalidad, comparando los rankings obtenidos frente a enfoques alternativos y analizando su impacto en el rendimiento de clasificación.
[Abstract] In many real-world machine learning problems, obtaining labels for all instances is infeasible. The Positive-Unlabeled (PU) learning scenario models this situation: only a subset of positive examples is labeled, while the remaining instances are unlabeled and may contain both true negatives and hidden positives. Treating unlabeled examples directly as negatives introduces a systematic bias that affects feature selection. This work proposes a mutual information-based feature selection method adapted to the PU setting, which corrects this bias by estimating the true probability of belonging to the positive class using the probabilistic framework of Elkan and Noto. The proposal is experimentally evaluated on nine datasets of varying nature and dimensionality, comparing the obtained rankings against alternative approaches and analyzing their impact on classification performance.
[Abstract] In many real-world machine learning problems, obtaining labels for all instances is infeasible. The Positive-Unlabeled (PU) learning scenario models this situation: only a subset of positive examples is labeled, while the remaining instances are unlabeled and may contain both true negatives and hidden positives. Treating unlabeled examples directly as negatives introduces a systematic bias that affects feature selection. This work proposes a mutual information-based feature selection method adapted to the PU setting, which corrects this bias by estimating the true probability of belonging to the positive class using the probabilistic framework of Elkan and Noto. The proposal is experimentally evaluated on nine datasets of varying nature and dimensionality, comparing the obtained rankings against alternative approaches and analyzing their impact on classification performance.
Description
Keywords
Aprendizaje Positive-Unlabeled Selección de características Información mutua Estimación de α Hipótesis SCAR Correlación de Spearman Alta dimensionalidad Sesgo de etiquetado Positive-Unlabeled learning Feature selection Mutual information α estimation SCAR assumption Spearman correlation High dimensionality Labeling bias
Editor version
Rights
Attribution-NonCommercial-NoDerivatives 4.0 International








