Enhancing Automatic Keyphrase Labelling with Text-to-Text Transfer Transformer (T5) Architecture: A Framework for Keyphrase Generation and Filtering
| UDC.coleccion | Investigación | |
| UDC.departamento | Ciencias da Computación e Tecnoloxías da Información | |
| UDC.grupoInv | Information Retrieval Lab (IRlab) | |
| UDC.institutoCentro | CITIC - Centro de Investigación de Tecnoloxías da Información e da Comunicación | |
| UDC.journalTitle | International Journal of Computational Intelligence Systems | |
| UDC.startPage | 270 | |
| UDC.volume | 19 | |
| dc.contributor.author | Gabín Brenlla, Jorge Juan | |
| dc.contributor.author | Ares Brea, Manuel Eduardo | |
| dc.contributor.author | Parapar, Javier | |
| dc.date.accessioned | 2026-07-20T10:21:58Z | |
| dc.date.available | 2026-07-20T10:21:58Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | [Abstract]: Automatic keyphrase labelling stands for the ability of models to retrieve words or short phrases that adequately describe documents’ content. Previous work has put much effort into exploring extractive techniques to address this task; however, most of these methods cannot produce keyphrases not found in the text. Given this limitation, keyphrase generation approaches have arisen lately. This paper presents a keyphrase generation model based on the Text-to-Text Transfer Transformer (T5) architecture. Having a document’s title and abstract as input, we learn a T5 model to generate keyphrases which adequately define its content. We name this model docT5keywords. We not only perform the classic inference approach, where the output sequence is directly selected as the predicted values, but we also report results from a majority voting approach. In this approach, multiple sequences are generated, and the keyphrases are ranked based on their frequency of occurrence across these sequences. Along with this model, we present a novel keyphrase filtering technique based on the T5 architecture. We train a T5 model to learn whether a given keyphrase is relevant to a document. We devise two evaluation methodologies to prove our model’s capability to filter inadequate keyphrases. First, we perform a binary evaluation where our model has to predict if a keyphrase is relevant for a given document. Second, we filter the predicted keyphrases by keyphrase generation models and check if the evaluation scores are improved. Experimental results show that our best docT5keywords variant yields relative improvements over the strongest baselines ranging from −7.5% to +60.5% across datasets and evaluation metrics; by evaluation type the gains are: present exact-match +23.9% to +60.5%, absent exact-match −7.5% to +60.4%, and partial-match (absent) −0.9% to +10.4%. The proposed filtering technique demonstrates strong performance in eliminating false positives across all datasets, even though identifying true keyphrases remains more challenging. | |
| dc.description.sponsorship | First and third authors acknowledge funding from the Ministry of Science, Innovation and Universities of the Government of Spain (projects PID2022-137061OB-C21, MCIN/AEI/10.13039/501100011033 and PID2025-167749OB-C22), as well as from the Department of Education, Science, Universities, and Vocational Training of the Xunta de Galicia (grant GRC ED431C 2025/49). CITIC, as a center accredited for excellence within the Galician University System and a member of the CIGUS Network, receives subsidies from the Department of Education, Science, Universities, and Vocational Training of the Xunta de Galicia. Additionally, it is co-financed by the EU through the FEDER Galicia 2021-27 operational program (Ref. ED431G 2023/01). Financial support was received from: Project PLEC2021-007662 (MCIN/AEI/10.13039/501100011033, Ministerio de Ciencia e Innovación, European Union NextGenerationEU/PRTR). Project PID2022-137061OB-C21 (MCIN/AEI/10.13039/501100011033, Ministerio de Ciencia e Innovación, ERDF “A way of making Europe” by the European Union). Consellería de Educación, Universidade e Formación Profesional, Spain (accreditation GRC ED431C 2025/49). Scholarship DIN2020-011582 financed by the MCIN/AEI/10.13039/501100011033 (acknowledged by the first author). | |
| dc.description.sponsorship | Xunta de Galicia; GRC ED431C 2025/49 | |
| dc.description.sponsorship | Xunta de Galicia; ED431G 2023/01 | |
| dc.identifier.citation | Gabín, J., Ares, M.E. & Parapar, J. Enhancing Automatic Keyphrase Labelling with Text-to-Text Transfer Transformer (T5) Architecture: A Framework for Keyphrase Generation and Filtering. Int J Comput Intell Syst 19, 270 (2026). https://doi.org/10.1007/s44196-026-01379-9 | |
| dc.identifier.doi | 10.1007/s44196-026-01379-9 | |
| dc.identifier.issn | 1875-6883 | |
| dc.identifier.uri | https://hdl.handle.net/2183/48897 | |
| dc.language.iso | eng | |
| dc.publisher | Springer Nature | |
| dc.relation.projectID | info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica, Técnica y de Innovación 2021-2023/PID2022-137061OB-C21/ES/BUSQUEDA, SELECCION Y ORGANIZACION DE CONTENIDOS PARA NECESIDADES DE INFORMACION RELACIONADAS CON LA SALUD - CONSTRUCCION DE RECURSOS Y PERSONALIZACION | |
| dc.relation.projectID | info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2024-2027/PID2025-167749OB-C22/ES/ | |
| dc.relation.projectID | info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2021-2024/PLEC2021-007662/ES/BIG-eRISK: PREDICCIÓN TEMPRANA DE RIESGOS PERSONALES EN CONJUNTOS DE DATOS MASIVOS | |
| dc.relation.projectID | info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020/DIN2020-011582/ES/PERSONALIZACIÓN DE LA EXPERIENCIA DE USUARIO EN ESCENARIOS DE BÚSQUEDA COMPLEJA PARA BUSINESS INTELLIGENCE | |
| dc.relation.uri | https://doi.org/10.1007/s44196-026-01379-9 | |
| dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 International | en |
| dc.rights.accessRights | open access | |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/4.0/ | |
| dc.subject | Keyphrase generation | |
| dc.subject | Keyphrase filtering | |
| dc.subject | Text-to-Text transfer transformer | |
| dc.subject | Large language models | |
| dc.subject | Natural language generation | |
| dc.subject | Document labelling | |
| dc.title | Enhancing Automatic Keyphrase Labelling with Text-to-Text Transfer Transformer (T5) Architecture: A Framework for Keyphrase Generation and Filtering | |
| dc.type | journal article | |
| dc.type.hasVersion | VoR | |
| dspace.entity.type | Publication | |
| relation.isAuthorOfPublication | fef1a9cb-e346-4e53-9811-192e144f09d0 | |
| relation.isAuthorOfPublication.latestForDiscovery | fef1a9cb-e346-4e53-9811-192e144f09d0 |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Parapar_Javier_2026_Enhancing_Automatic_Keyphrase_Labelling.pdf
- Size:
- 2.91 MB
- Format:
- Adobe Portable Document Format

