Improving Clinical Reliability of LLM Reasoning for Depression Assessment via Structured Generation and GRPO

UDC.coleccionInvestigación
UDC.departamentoCiencias da Computación e Tecnoloxías da Información
UDC.grupoInvInformation Retrieval Lab (IRlab)
UDC.institutoCentroCITIC - Centro de Investigación de Tecnoloxías da Información e da Comunicación
UDC.issue3
UDC.journalTitleScience Progress
UDC.volume109
dc.contributor.authorBao, Eliseo
dc.contributor.authorPérez, Anxo
dc.contributor.authorParapar, Javier
dc.date.accessioned2026-08-17T09:06:30Z
dc.date.available2026-08-17T09:06:30Z
dc.date.issued2026-07
dc.descriptionThe ReDSM5 dataset used in this study is publicly available at https://huggingface.co/datasets/irlab-udc/redsm5. Code for model training and evaluation is available at https://github.com/IRLab-UDC/grpo-dsm5/.
dc.description.abstract[Abstract]: Objective. Digital mental health screening is increasingly explored through the use of AI systems. Yet, most models provide limited insight into how predictions are derived, restricting clinical trust and patient-centered adoption. We investigate whether Large Language Models (LLMs) can generate structured DSM-5-aligned rationales for depression assessment when trained with Group Relative Policy Optimization (GRPO), a reinforcement learning method that encourages outputs aligned with DSM-5 diagnostic criteria. Methods. We fine-tuned LLMs (1B–27B parameters) on the ReDSM5 dataset, covering 1,484 Reddit posts annotated by a licensed psychologist for DSM-5 depressive symptoms and accompanied by expert rationales. We compared standard Supervised Fine-Tuning (SFT) with GRPO-based optimization using a composite reward integrating symptom classification accuracy and the quality of generated reasoning judged against clinical rationales. Results. GRPO consistently improved symptom detection over SFT, with relative gains exceeding 10% for mid-sized models and weighted F1 scores above 0.60. Models trained to generate structured DSM-5-aligned rationales exhibited additional performance boosts (0.09–0.39 F1), particularly for complex symptoms requiring complex contextual interpretation. Qualitative analysis shows that GRPO encourages models to reference symptom-relevant evidence rather than relying on superficial cues. Conclusions. GRPO enables LLMs to produce clinically grounded explanations while improving classification accuracy, representing a promising direction for interpretable AI in social media-based mental health screening, with potential to support patient-centered applications pending clinical validation.
dc.description.sponsorshipThe authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This workwas supported by the Department of Education, Science, Universities, and Vocational Training of the Xunta de Galicia (grant ED481A-2024-079); by the Ministry of Science, Innovation and Universities of the Government of Spain (project PID2022-137061OB-C21, MCIN/AEI/10.13039/501100011033); by the Department of Education, Science, Universities, and Vocational Training of the Xunta de Galicia (grant GRC ED431C 2025/49); and by the Xunta de Galicia and the European Union through the FEDER Galicia 2021–2027 OperationalProgramme (Ref. ED431G 2023/01).
dc.description.sponsorshipXunta de Galicia; ED481A-2024-079
dc.description.sponsorshipXunta de Galicia; ED431C 2025/49
dc.description.sponsorshipXunta de Galicia; ED431G 2023/01
dc.identifier.citationE. Bao, A. Perez, and J. Parapar, "Improving Clinical Reliability of LLM Reasoning for Depression Assessment via Structured Generation and GRPO", Science Progress, Vol. 109, Issue 3, Jul 2026, https://doi.org/10.1177/003685042614674
dc.identifier.doi10.1177/003685042614674
dc.identifier.issn2047-7163
dc.identifier.urihttps://hdl.handle.net/2183/49037
dc.language.isoeng
dc.publisherSage
dc.relation.isbasedonhttps://huggingface.co/datasets/irlab-udc/redsm5
dc.relation.isbasedonhttps://github.com/IRLab-UDC/grpo-dsm5/
dc.relation.projectIDinfo:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica, Técnica y de Innovación 2021-2023/PID2022-137061OB-C21/ES/BUSQUEDA, SELECCION Y ORGANIZACION DE CONTENIDOS PARA NECESIDADES DE INFORMACION RELACIONADAS CON LA SALUD - CONSTRUCCION DE RECURSOS Y PERSONALIZACION
dc.relation.urihttps://doi.org/10.1177/003685042614674
dc.rightsAttribution-NonCommercial 4.0 Internationalen
dc.rights.accessRightsopen access
dc.rights.urihttp://creativecommons.org/licenses/by-nc/4.0/
dc.subjectMental health
dc.subjectDepression
dc.subjectSocial media
dc.subjectDSM-5
dc.subjectReasoning
dc.subjectExplainability
dc.titleImproving Clinical Reliability of LLM Reasoning for Depression Assessment via Structured Generation and GRPO
dc.typejournal article
dc.type.hasVersionVoR
dspace.entity.typePublication
relation.isAuthorOfPublication99ed6581-6dee-442a-9b37-c35da63bef8a
relation.isAuthorOfPublicationc673c8b1-1afc-48f6-85e9-8f29f9cffb91
relation.isAuthorOfPublicationfef1a9cb-e346-4e53-9811-192e144f09d0
relation.isAuthorOfPublication.latestForDiscovery99ed6581-6dee-442a-9b37-c35da63bef8a

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Bao_Eliseo_2026_Improving_Clinical_Reliability_of_LLM_Reasoning.pdf
Size:
13.07 MB
Format:
Adobe Portable Document Format