Fernández-Blanco, EnriquePuente-Castro, AlejandroBarros Iglesias, MartínUniversidade da Coruña. Facultade de Informática2026-09-302026-09-302026-06https://hdl.handle.net/2183/49544[Resumen] Este Trabajo de Fin de Grado aborda la detección automática de la enfermedad de Alzheimer a partir del análisis multimodal del habla espontánea, en un contexto en el que el envejecimiento de la población ha convertido a esta patología en uno de los grandes retos sociosanitarios, dada la ausencia de tratamientos curativos y las limitaciones de las pruebas neuropsicológicas tradicionales para detectarla en fases tempranas. Para ello se ha empleado un corpus de referencia en el ámbito internacional, formado por grabaciones de pacientes y de participantes sanos describiendo oralmente una lámina estandarizada, lo que permite estudiar de forma natural las huellas lingüísticas y acústicas que deja la enfermedad en el discurso. El sistema desarrollado combina dos enfoques complementarios. Por un lado, se han extraído representaciones de alta dimensionalidad mediante modelos fundacionales de aprendizaje profundo preentrenados, tanto para el canal textual como para el canal acústico, que se combinan posteriormente mediante distintas estrategias de fusión multimodal. Por otro lado, se han diseñado módulos de caja blanca basados en biomarcadores clínicamente interpretables, centrados en la estructura gramatical del discurso y en la dinámica de las pausas y los silencios, con el objetivo de ofrecer al personal clínico una explicación comprensible de las decisiones del sistema. Todos los modelos se han evaluado bajo un protocolo estadístico común y riguroso, que permite comparar de forma objetiva su rendimiento y la significancia de las diferencias observadas entre ellos. Los resultados confirman que las representaciones textuales profundas capturan por sí solas la mayor parte de la información discriminativa disponible en el corpus, de modo que su combinación con el canal acústico no aporta una mejora adicional relevante, un fenómeno que en este trabajo se denomina dominancia lingüística. Asimismo, se ha logrado cuantificar con precisión el coste, en términos de rendimiento, de exigir explicabilidad clínica frente a los modelos de caja negra de mayor capacidad predictiva. Este enfoque pone de manifiesto que es posible construir sistemas de cribado automático del Alzheimer que resulten a la vez fiables, objetivos y comprensibles para el clínico, abriendo nuevas vías de trabajo orientadas a reforzar la aportación del canal acústico y a ampliar la validación clínica y externa del sistema.[Abstract] This undergraduate thesis addresses the automatic detection of Alzheimer’s disease through multimodal analysis of spontaneous speech, in a context where population ageing has turned this condition into one of the major sociosanitary challenges of our time, given the absence of curative treatments and the limitations of traditional neuropsychological tests for detecting it at early stages. To this end, an internationally recognised reference corpus has been used, composed of recordings of patients and healthy participants orally describing a standardised picture, allowing the linguistic and acoustic traces left by the disease in spontaneous discourse to be studied in a natural way. The developed system combines two complementary approaches. On one hand, high dimensional representations were extracted using pre-trained deep learning foundation models, for both the textual and the acoustic channel, which are subsequently combined through different multimodal fusion strategies. On the other hand, white-box modules based on clin ically interpretable biomarkers were designed, focusing on the grammatical structure of dis course and on the dynamics of pauses and silences, with the aim of offering clinicians an understandable explanation of the system’s decisions. All models were evaluated under a common, rigorous statistical protocol, allowing an objective comparison of their performance and the significance of the differences observed between them. The results confirm that deep textual representations alone capture most of the discrimi native information available in the corpus, so that combining them with the acoustic channel does not provide a relevant additional improvement, a phenomenon referred to in this work as linguistic dominance. Likewise, the cost, in terms of performance, of requiring clinical explainability relative to the highest-capacity black-box models has been precisely quantified.This approach shows that it is possible to build automatic Alzheimer’s screening systems that are simultaneously reliable, objective and understandable to clinicians, opening new lines of work aimed at reinforcing the contribution of the acoustic channel and extending the clinical and external validation of the system.spaAttribution-NonCommercial-ShareAlike 4.0 Internationalhttp://creativecommons.org/licenses/by-nc-sa/4.0/Enfermedad de AlzheimerHabla espontáneaProcesamiento del Lenguaje NaturalAprendizaje profundoFusión multimodalInterpretabilidad (XAI)Biomarcadores digitalesAlzheimer’s DiseaseSpontaneous SpeechNatural Language ProcessingDeep LearningMultimodal FusionExplainable AI (XAI)Digital BiomarkerDetección Multimodal del Alzheimer mediante Habla Espontánea: Fusión Profunda e Interpretabilidad Clínicabachelor thesisopen access