Use this link to cite:
https://hdl.handle.net/2183/49468 Validación de biclusters en sistemas paralelos de memoria compartida
Loading...
Identifiers
Publication date
Authors
Dios Touriñán, Lucas
Advisors
Other responsabilities
Universidade da Coruña. Facultade de Informática
Journal Title
Bibliographic citation
Type of academic work
Academic degree
Abstract
[Resumen] El biclustering es una técnica muy empleada dentro de la minería de datos para encontrar enlaces o patrones bidimensionales. En la actualidad existen muchos algoritmos de biclustering, cada uno con sus ventajas e inconvenientes. Sin embargo, tradicionalmente no se le ha prestado especial atención al proceso de validar dichos métodos de forma exhaustiva e independiente. Aunque existen varias métricas de validación, su elevado coste computacional dificulta la comparación y el procesamiento eficiente de grandes conjuntos de datos. En este Trabajo Fin de Grado se desarrolla una herramienta para la validación de biclusters basada en múltiples métricas estadísticas, incluyendo Pearson, Spearman, Kendall, Información Mutua, LogOdds y Matthews Correlation Coefficient. El sistema permite realizar tanto evaluaciones individuales como estrategias ensemble basadas en combinaciones lógicas AND y OR. La implementación se ha realizado en lenguaje C siguiendo una arquitectura modular y extensible. Asimismo, se han incorporado optimizaciones específicas, entre las que destaca una estrategia de corte temprano para la evaluación AND, así como mecanismos de paralelización mediante OpenMP sobre arquitecturas de memoria compartida. La evaluación experimental se realizó utilizando conjuntos de datos de expresión génica de gran tamaño y ejecutando los experimentos tanto en un entorno local como en el supercomputador FinisTerrae III del CESGA. Los resultados obtenidos muestran elevados niveles de escalabilidad y reducciones significativas de los tiempos de ejecución, alcanzando speedups próximos al comportamiento ideal en múltiples configuraciones. Finalmente, se analiza el comportamiento de las distintas métricas implementadas en el caso concreto de análisis génico. Se evalúa su precisión biológica, se justifica la selección del umbral de aceptación utilizado y se estudia el impacto computacional de las estrategias ensemble consideradas.
[Abstract] Biclustering is a widely used data mining technique for discovering bidimensional patterns and relationships within datasets. Nowadays, numerous biclustering algorithms exist, each presenting its own advantages and limitations. However, comparatively little attention has been devoted to the independent and exhaustive validation of the biclusters produced by these methods. Although several validation metrics have been proposed, their high computational cost makes the comparison and efficient processing of large datasets particularly challenging. This Bachelor’s Thesis presents the development of a bicluster validation tool based on multiple statistical metrics, including Pearson correlation, Spearman correlation, Kendall’s Tau, Mutual Information, Log Odds Ratio and Matthews Correlation Coefficient. The system supports both individual metric evaluation and ensemble strategies based on logical AND and OR combinations. The implementation has been developed in the C programming language following a modular and extensible architecture. In addition, several optimizations have been incorporated, including an early stopping strategy for AND-based ensemble evaluation and parallelization mechanisms based on OpenMP for shared-memory architectures. The experimental evaluation was conducted using large-scale gene expression datasets and executing the experiments both in a local environment and on the FinisTerrae III supercomputer at CESGA. The results obtained show high levels of scalability and significant reductions in execution times, achieving speedups close to ideal behavior in multiple configurations. Finally, the behavior of the different validation metrics is analyzed in the specific context of gene expression analysis. Their biological relevance is evaluated, the selection of the acceptance threshold is justified, and the computational impact of the considered ensemble strategies is studied.
[Abstract] Biclustering is a widely used data mining technique for discovering bidimensional patterns and relationships within datasets. Nowadays, numerous biclustering algorithms exist, each presenting its own advantages and limitations. However, comparatively little attention has been devoted to the independent and exhaustive validation of the biclusters produced by these methods. Although several validation metrics have been proposed, their high computational cost makes the comparison and efficient processing of large datasets particularly challenging. This Bachelor’s Thesis presents the development of a bicluster validation tool based on multiple statistical metrics, including Pearson correlation, Spearman correlation, Kendall’s Tau, Mutual Information, Log Odds Ratio and Matthews Correlation Coefficient. The system supports both individual metric evaluation and ensemble strategies based on logical AND and OR combinations. The implementation has been developed in the C programming language following a modular and extensible architecture. In addition, several optimizations have been incorporated, including an early stopping strategy for AND-based ensemble evaluation and parallelization mechanisms based on OpenMP for shared-memory architectures. The experimental evaluation was conducted using large-scale gene expression datasets and executing the experiments both in a local environment and on the FinisTerrae III supercomputer at CESGA. The results obtained show high levels of scalability and significant reductions in execution times, achieving speedups close to ideal behavior in multiple configurations. Finally, the behavior of the different validation metrics is analyzed in the specific context of gene expression analysis. Their biological relevance is evaluated, the selection of the acceptance threshold is justified, and the computational impact of the considered ensemble strategies is studied.
Description
Editor version
Rights
Todos os dereitos reservados. Neste caso prohíbese a reprodución, transformación, distribución e comunicación pública da obra por terceiros. En cambio, permítese a súa visualización e a descarga dunha copia privada para o uso persoal.







