Bridging gaps in hate speech detection: meta-collections and benchmarks for low-resource Iberian languages
| UDC.coleccion | Investigación | |
| UDC.departamento | Ciencias da Computación e Tecnoloxías da Información | |
| UDC.grupoInv | Information Retrieval Lab (IRlab) | |
| UDC.institutoCentro | CITIC - Centro de Investigación de Tecnoloxías da Información e da Comunicación | |
| UDC.issue | 41 | |
| UDC.journalTitle | Language Resources and Evaluation | |
| UDC.volume | 60 | |
| dc.contributor.author | Piot, Paloma | |
| dc.contributor.author | Pichel Campos, José Ramom | |
| dc.contributor.author | Parapar, Javier | |
| dc.date.accessioned | 2026-05-18T07:04:14Z | |
| dc.date.available | 2026-05-18T07:04:14Z | |
| dc.date.issued | 2026-04 | |
| dc.description | Data is available upon request, and we have adhered to strict ethical standards, particularly those about the ethical development of AI (https://huggingface.co/datasets/irlab-udc/MetaHateES). The code to replicate the meta-collection construction, translation, analysis and running the baselines is available here https://github.com/palomapiot/mlhate. | |
| dc.description.abstract | [Abstract]: Hate speech poses a serious threat to social cohesion and individual well-being, particularly on social media, where it spreads rapidly. While research on hate speech detection has progressed, it remains largely focused on English, resulting in limited resources and benchmarks for low-resource languages. Moreover, many of these languages have multiple linguistic varieties, a factor often overlooked in current approaches. At the same time, large language models require substantial amounts of data to perform reliably, a requirement that low-resource languages often cannot meet. In this work, we address these gaps by compiling a meta-collection of hate speech datasets for European Spanish, standardised with unified labels and metadata. This collection is based on a systematic analysis and integration of existing resources, aiming to bridge the data gap and support more consistent and scalable hate speech detection. We extended this collection by translating it into European Portuguese and into a Galician standard that is more convergent with Spanish and another Galician variant that is more convergent with Portuguese, creating aligned multilingual corpora. Using these resources, we establish new benchmarks for hate speech detection in Iberian languages. We evaluate state-of-the-art large language models in zero-shot, few-shot, and fine-tuning settings, providing baseline results for future research. Moreover, we perform a cross-lingual analysis with our target languages. Our findings underscore the importance of multilingual and variety-aware approaches in hate speech detection and offer a foundation for improved benchmarking in underrepresented European languages. | |
| dc.description.sponsorship | Open Access funding provided thanks to the CRUE-CSIC agreement with Springer Nature. The authors thank the funding from the Horizon Europe research and innovation programme under the Marie Skłodowska-Curie Grant Agreement No. 101073351. The first and third authors also thank the financial support supplied by the grant PID2022-137061OB-C21 funded by MI-CIU/AEI/10.13039/501100011033 and by “ERDF/EU”. The authors also thank the funding supplied by the Consellería de Cultura, Educación, Formación Profesional e Universidades (accreditations ED431G 2023/01 and ED431C 2025/49) and the European Regional Development Fund, which acknowledges the CITIC, as a center accredited for excellence within the Galician University System and a member of the CIGUS Network, receives subsidies from the Department of Education, Science, Universities, and Vocational Training of the Xunta de Galicia. Additionally, it is co-financed by the EU through the FEDER Galicia 2021-27 operational program (Ref. ED431G 2023/01). | |
| dc.description.sponsorship | Xunta de Galicia; ED431G 2023/01 | |
| dc.description.sponsorship | Xunta de Galicia; ED431C 2025/49 | |
| dc.identifier.citation | Piot, P., Pichel Campos, J.R. & Parapar, J. Bridging gaps in hate speech detection: meta-collections and benchmarks for low-resource Iberian languages. Lang Resources & Evaluation 60, 41 (2026). https://doi-org.accedys.udc.es/10.1007/s10579-026-09921-z | |
| dc.identifier.doi | 10.1007/s10579-026-09921-z | |
| dc.identifier.issn | 1574-0218 | |
| dc.identifier.uri | https://hdl.handle.net/2183/48277 | |
| dc.language.iso | eng | |
| dc.publisher | Springer Nature | |
| dc.relation.isbasedon | https://huggingface.co/datasets/irlab-udc/MetaHateES | |
| dc.relation.isbasedon | https://github.com/palomapiot/mlhate | |
| dc.relation.projectID | info:eu-repo/grantAgreement/EC/HE/101073351 | |
| dc.relation.projectID | info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica, Técnica y de Innovación 2021-2023/PID2022-137061OB-C21/ES/BUSQUEDA, SELECCION Y ORGANIZACION DE CONTENIDOS PARA NECESIDADES DE INFORMACION RELACIONADAS CON LA SALUD - CONSTRUCCION DE RECURSOS Y PERSONALIZACION | |
| dc.relation.uri | https://doi.org/10.1007/s10579-026-09921-z | |
| dc.rights | Attribution 4.0 International | en |
| dc.rights.accessRights | open access | |
| dc.rights.uri | http://creativecommons.org/licenses/by/4.0/ | |
| dc.subject | Hate speech | |
| dc.subject | Meta-collection | |
| dc.subject | Low-resource | |
| dc.subject | LLMs | |
| dc.subject | Cross-lingual | |
| dc.title | Bridging gaps in hate speech detection: meta-collections and benchmarks for low-resource Iberian languages | |
| dc.type | journal article | |
| dc.type.hasVersion | VoR | |
| dspace.entity.type | Publication | |
| relation.isAuthorOfPublication | 0563c6c3-cd50-4d7d-b11f-127ee297dd6b | |
| relation.isAuthorOfPublication | fef1a9cb-e346-4e53-9811-192e144f09d0 | |
| relation.isAuthorOfPublication.latestForDiscovery | 0563c6c3-cd50-4d7d-b11f-127ee297dd6b |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Parapar_Javier_2026_Bridging_gaps_in_hate_speech_detection.pdf
- Size:
- 3.21 MB
- Format:
- Adobe Portable Document Format

