Evaluating Factual Grounding Strategies in Large Language Models

UDC.coleccionPublicacións UDC
UDC.conferenceTitleXoveTIC: impulsando el talento científico (8º. 2025. A Coruña)
UDC.departamentoCiencias da Computación e Tecnoloxías da Información
UDC.endPage356
UDC.grupoInvInformation Retrieval Lab (IRlab)
UDC.institutoCentroCITIC - Centro de Investigación de Tecnoloxías da Información e da Comunicación
UDC.startPage349
dc.contributor.authorFernández, Pablo
dc.contributor.authorPérez, Anxo
dc.contributor.authorParapar, Javier
dc.date.accessioned2026-09-10T15:09:27Z
dc.date.available2026-09-10T15:09:27Z
dc.date.issued2025
dc.descriptionPresentado en: VIII Congreso Xove TIC: impulsando el talento científico. Octubre, 2025, A Coruña.
dc.description.abstract[Abstract] Large Language Models (LLMs) often generate non-factual, or “hallucinated,” content, limiting their reliability in knowledge-intensive tasks. This challenge is particularly critical in multi-hop question answering (MHQA), where models must integrate and reason over multiple pieces of evidence. In this paper, we present an empirical study of prompting strategies aimed at improving the factual grounding of LLMs. Using the Llama-3.1 (8B) model on the HotpotQA benchmark, we evaluate five prompting techniques along three main design points: shot count (zero-shot vs. few-shot), context integration (with vs. without supporting documents), and output constraints (free-form vs. structured responses requiring evidence). We assess both answer accuracy and the precision of supporting fact identification, allowing us to analyze correctness from evidential grounding. Our results reveal different trade-offs: while few-shot prompting improves reasoning consistency, the gains diminish without high-quality supporting context. Similarly, structured outputs reduce variance and improve factual alignment, but their benefits depend critically on how evidence is presented and constrained. Compared to prior studies that focus primarily on accuracy, our analysis highlights the importance of balancing answer quality with verifiable evidence. These findings provide actionable guidance for the design of prompts in multi-hop QA and inform broader efforts to mitigate hallucinations in LLMs across retrieval-augmented and reasoning-intensive applications.
dc.identifier.citationFernández, P., Pérez, A., & Parapar, J. (2026). Evaluating Factual Grounding Strategies in Large Language Models. In Proceedings XoveTIC 2025: Impulsando el talento científico (pp. 349-356). Servizo de Publicacións UDC. https://doi.org/10.17979/spu.23.c55
dc.identifier.doi10.17979/spu.23.c55
dc.identifier.isbn978-84-9749-925-5
dc.identifier.urihttps://hdl.handle.net/2183/49193
dc.language.isoeng
dc.publisherUniversidade da Coruña, Servizo de Publicacións
dc.relation.urihttps://doi.org/10.17979/spu.23.c55
dc.rightsAttribution-NonCommercial-NoDerivatives 4.0 Internationalen
dc.rights.accessRightsopen access
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/4.0/
dc.subjectLarge Language Models (LLMs)
dc.subjectPrompting strategies
dc.subjectMulti-Hop Question Answering (MHQA)
dc.subjectHallucination mitigation
dc.titleEvaluating Factual Grounding Strategies in Large Language Models
dc.typeconference output
dspace.entity.typePublication
relation.isAuthorOfPublicationc673c8b1-1afc-48f6-85e9-8f29f9cffb91
relation.isAuthorOfPublicationfef1a9cb-e346-4e53-9811-192e144f09d0
relation.isAuthorOfPublication.latestForDiscoveryc673c8b1-1afc-48f6-85e9-8f29f9cffb91

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
XoveTIC_2025_proceedings_c55.pdf
Size:
1.83 MB
Format:
Adobe Portable Document Format