Bao, EliseoPérez, AnxoParapar, Javier2026-08-172026-08-172026-07E. Bao, A. Perez, and J. Parapar, "Improving Clinical Reliability of LLM Reasoning for Depression Assessment via Structured Generation and GRPO", Science Progress, Vol. 109, Issue 3, Jul 2026, https://doi.org/10.1177/0036850426146742047-7163https://hdl.handle.net/2183/49037The ReDSM5 dataset used in this study is publicly available at https://huggingface.co/datasets/irlab-udc/redsm5. Code for model training and evaluation is available at https://github.com/IRLab-UDC/grpo-dsm5/.[Abstract]: Objective. Digital mental health screening is increasingly explored through the use of AI systems. Yet, most models provide limited insight into how predictions are derived, restricting clinical trust and patient-centered adoption. We investigate whether Large Language Models (LLMs) can generate structured DSM-5-aligned rationales for depression assessment when trained with Group Relative Policy Optimization (GRPO), a reinforcement learning method that encourages outputs aligned with DSM-5 diagnostic criteria. Methods. We fine-tuned LLMs (1B–27B parameters) on the ReDSM5 dataset, covering 1,484 Reddit posts annotated by a licensed psychologist for DSM-5 depressive symptoms and accompanied by expert rationales. We compared standard Supervised Fine-Tuning (SFT) with GRPO-based optimization using a composite reward integrating symptom classification accuracy and the quality of generated reasoning judged against clinical rationales. Results. GRPO consistently improved symptom detection over SFT, with relative gains exceeding 10% for mid-sized models and weighted F1 scores above 0.60. Models trained to generate structured DSM-5-aligned rationales exhibited additional performance boosts (0.09–0.39 F1), particularly for complex symptoms requiring complex contextual interpretation. Qualitative analysis shows that GRPO encourages models to reference symptom-relevant evidence rather than relying on superficial cues. Conclusions. GRPO enables LLMs to produce clinically grounded explanations while improving classification accuracy, representing a promising direction for interpretable AI in social media-based mental health screening, with potential to support patient-centered applications pending clinical validation.engAttribution-NonCommercial 4.0 Internationalhttp://creativecommons.org/licenses/by-nc/4.0/Mental healthDepressionSocial mediaDSM-5ReasoningExplainabilityImproving Clinical Reliability of LLM Reasoning for Depression Assessment via Structured Generation and GRPOjournal articleopen access10.1177/003685042614674