Aprendizaje profundo de representaciones multimodales para la alineación del espacio latente retina-cerebro en la predicción de alteración cognitiva

Loading...
Thumbnail Image

Identifiers

Publication date

Authors

Díaz Achitei, Raúl

Advisors

Torán-Monserrat, Pere

Other responsabilities

Universidade da Coruña. Facultade de Informática

Journal Title

Bibliographic citation

Type of academic work

Abstract

[Resumen] El fondo de ojo es una de las pocas estructuras del cuerpo humano donde la microvasculatura puede observarse de forma directa y no invasiva, lo que ha motivado un interés creciente por su uso como ventana de acceso a procesos neurológicos y sistémicos que de otro modo requerirían pruebas costosas o invasivas. Este trabajo explora hasta qué punto la imagen retiniana, por sí sola o combinada con otras fuentes de información, permite anticipar alteraciones cognitivas leves y detectar secuelas asociadas a la infección por SARS-CoV-2. Sobre una cohorte de poco más de cien pacientes con imagen de fondo de ojo y resonancia magnética cerebral, se entrena primero una red EfficientNetV2-S mediante fine-tuning en dos fases mediante el ajuste fino de las imágenes retinianas, alcanzando un AUC de 0,757 para la tarea de alteración cognitiva. Tomando como base este modelo de referencia se diseña e implementa un marco Teacher-Student basado en Learning Using Privileged Information (LUPI), en el que un segundo modelo (teacher) procesa cincuenta biomarcadores extraídos de la resonancia magnética seleccionados mediante pruebas de normalidad, contraste de hipótesis y corrección de comparaciones múltiples, y transmite dicho conocimiento al modelo retiniano mediante una pérdida de distilación por similitud coseno. El resultado es un estudiante que, sin necesitar resonancia en el momento de la predicción, eleva el AUC hasta 0,799, una mejora de más de cuatro puntos porcentuales atribuible exclusivamente a la información privilegiada incorporada durante el entrenamiento. Como contraste, se evalúan también variantes tabulares del mismo marco, un perceptrón residual y un FT-Transformer entrenadas sobre biomarcadores retinianos cuantitativos (calibre vascular, densidad, complejidad fractal), que no logran superar el azar de forma consistente, lo que sugiere que la señal cognitiva reside en patrones de la imagen difícilmente reducibles a un conjunto cerrado de medidas tabulares. En paralelo, se reutiliza la misma metodología para estudiar dos tareas derivadas de la pandemia: la discriminación entre pacientes con y sin COVID-19, y la distinción entre quienes desarrollaron síntomas persistentes (COVID persistente) y quienes se recuperaron por completo. Combinando embeddings de una red ResNet18 ajustada específicamente para cada tarea con los mismos biomarcadores vasculares, los clasificadores multimodales alcanzan un AUC de 0,933 para la primera tarea y de 0,950 para la segunda. Un experimento de control, en el que se compara la red sin ajustar frente a la red ajustada, confirma que la mejora proviene de una representación visual genuina del daño retiniano y no de un artefacto del procedimiento. Finalmente, se desarrolla una interfaz clínica que permite introducir una imagen retiniana o un conjunto de biomarcadores y obtener un informe descargable en PDF con el riesgo estimado, pensada como primer paso hacia una validación piloto en un entorno asistencial real. El trabajo concluye que la retina contiene información aprovechable tanto para el deterioro cognitivo como para el estado sistémico post-infeccioso, si bien la primera señal es considerablemente más débil y solo se vuelve explotable cuando se complementa, en fase de entrenamiento, con una fuente de información más directa como la resonancia magnética.
[Abstract] The ocular fundus is one of the few structures in the human body where microvasculature can be observed directly and non-invasively, which has motivated growing interest in its use as a window into neurological and systemic processes that would otherwise require costly or invasive examinations. This work explores to what extent retinal imaging, on its own or combined with other sources of information, can anticipate mild cognitive impairment and detect sequelae associated with SARS-CoV-2 infection. Using a cohort of just over one hundred patients with both fundus photographs and brain magnetic resonance imaging, an EfficientNetV2-S network is first fine-tuned in two stages directly on the retinal images, reaching an AUC of 0.757 for the cognitive impairment task. Starting from this baseline, a Teacher-Student framework grounded in Learning Using Privileged Information (LUPI) is designed and implemented: a second model (the teacher) processes fifty biomarkers extracted from the MRI scans selected through normality testing, hypothesis testing and multiple comparison correction, and transfers that knowledge to the retinal model via a cosine-similarity distillation loss. The resulting student, which requires no MRI data at prediction time, raises the AUC to 0.799, an improvement of over four percentage points attributable solely to the privileged information injected during training. As a point of contrast, tabular variants of the same framework are also evaluated a residual perceptron and an FT-Transformer, trained on quantitative retinal biomarkers (vascular caliber, density, fractal complexity); neither consistently outperforms chance, suggesting that the cognitive signal lies in image patterns that cannot be easily reduced to a closed set of tabular measurements. In parallel, the same methodology is reused to study two tasks derived from the pandemic: distinguishing COVID-19 positive from negative patients, and distinguishing those who developed persistent symptoms (long COVID) from those who fully recovered. By combining embeddings from a ResNet18 network fine-tuned separately for each task with the same vascular biomarkers, the resulting multimodal classifiers reach an AUC of 0.933 for the first task and 0.950 for the second, the latter being the best result obtained throughout the whole project. A control experiment comparing the unfine-tuned network against the fine-tuned one confirms that this improvement stems from a genuine visual representation of retinal damage rather than from an artifact of the procedure. Finally, a clinical interface is developed that accepts either a retinal image or a set of biomarkers and returns a downloadable PDF report with the estimated risk, intended as a first step towards a pilot validation in a real clinical setting. The work concludes that the retina carries exploitable information about both cognitive decline and post-infectious systemic state, although the former signal is considerably weaker and only becomes usable when complemented, at training time, with a more direct source of information such as MRI.

Description

Editor version

Rights

Attribution 4.0 International
Attribution 4.0 International

Except where otherwise noted, this item's license is described as Attribution 4.0 International