Use this link to cite:
https://hdl.handle.net/2183/46176 Clasificación de cadeas de ADN a través de diferentes modelos de aprendizaxe máquina explicables
Loading...
Identifiers
Publication date
Authors
Dobarro Landeira, Daniel
Other responsabilities
Universidade da Coruña. Facultade de Informática
Journal Title
Bibliographic citation
Type of academic work
Academic degree
Abstract
[Resumo]: O presente traballo aborda a clasificación de secuencias de ADN co obxectivo de descifrar a lóxica funcional do regulador mestre SRRM4, un elemento clave na regulación de microexóns. Dada a alta dimensionalidade dos datos xenómicos, o proxecto céntrase no desenvolvemento de modelos de aprendizaxe automática explicables que non só acaden unha alta precisión na clasificación, senón que tamén ofrezan interpretabilidade sobre as súas decisións. Para iso, deséñase e impleméntase unha arquitectura de rede neuronal baseada nunha UNet unidimensional que actúa como extractor de características, combinada cunha capa de DDS. Esta última permite ao modelo identificar e seleccionar un subconxunto reducido dos nucleótidos máis relevantes para cada secuencia de ADN, proporcionando así unha xanela á súa lóxica interna. Avaliouse o rendemento do sistema en tarefas de supervisada e non supervisadas, demostrando que a selección dinámica de características é un enfoque viable que acada un rendemento competitivo. Os resultados amosan unha correlación positiva entre o número de características seleccionadas e a precisión do modelo, subliñando o compromiso entre a interpretabilidade e a capacidade predictiva. Máis aló da clasificación, o traballo destaca polo potencial do modelo como ferramenta para o descubrimento biolóxico, xa que a análise das características seleccionadas pode axudar a formular novas hipóteses sobre os mecanismos de regulación xenética.
[Abstract]: This project addresses the classification of DNA sequences with the aim of deciphering the functional logic of the master regulator SRRM4, a key element in the regulation of microexons. Given the high dimensionality of genomic data, the project focuses on the development of explainable artificial intelligence models that not only achieve high classification accuracy but also offer interpretability of their decisions. To this end, a neural network architecture based on a one-dimensional U-Net is designed and implemented to act as a feature extractor, combined with a dynamic feature selection (DDS) layer. The latter allows the model to identify and select a reduced subset of the most relevant nucleotides for each DNA sequence, thus providing a window into its internal logic. The system’s performance was evaluated on multiclass classification tasks, demonstrating that dynamic feature selection is a viable approach that achieves competitive performance. The results show a positive correlation between the number of selected features and the model’s accuracy, highlighting the trade-off between interpretability and predictive power. Beyond classification, the work stands out for the model’s potential as a tool for biological discovery, as the analysis of the selected features can help formulate new hypotheses about gene regulation mechanisms.
[Abstract]: This project addresses the classification of DNA sequences with the aim of deciphering the functional logic of the master regulator SRRM4, a key element in the regulation of microexons. Given the high dimensionality of genomic data, the project focuses on the development of explainable artificial intelligence models that not only achieve high classification accuracy but also offer interpretability of their decisions. To this end, a neural network architecture based on a one-dimensional U-Net is designed and implemented to act as a feature extractor, combined with a dynamic feature selection (DDS) layer. The latter allows the model to identify and select a reduced subset of the most relevant nucleotides for each DNA sequence, thus providing a window into its internal logic. The system’s performance was evaluated on multiclass classification tasks, demonstrating that dynamic feature selection is a viable approach that achieves competitive performance. The results show a positive correlation between the number of selected features and the model’s accuracy, highlighting the trade-off between interpretability and predictive power. Beyond classification, the work stands out for the model’s potential as a tool for biological discovery, as the analysis of the selected features can help formulate new hypotheses about gene regulation mechanisms.
Description
Keywords
Modelos Interpretables Xenómica Computacional U-Net Selección Dinámica de Características Clasificación Predicción Regresión Adestramento Interpretación Evaluación Interpretable Models Computational Genomics Dynamic Feature Selection Classification Prediction Regression Training Interpretation Evaluation
Editor version
Rights
Attribution 4.0 International








