Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
| UDC.coleccion | Investigación | |
| UDC.conferenceTitle | LREC 2026 | |
| UDC.departamento | Ciencias da Computación e Tecnoloxías da Información | |
| UDC.endPage | 4370 | |
| UDC.grupoInv | Information Retrieval Lab (IRlab) | |
| UDC.institutoCentro | CITIC - Centro de Investigación de Tecnoloxías da Información e da Comunicación | |
| UDC.startPage | 4358 | |
| dc.contributor.author | Piot, Paloma | |
| dc.contributor.author | Otero, David | |
| dc.contributor.author | Martín-Rodilla, Patricia | |
| dc.contributor.author | Parapar, Javier | |
| dc.date.accessioned | 2026-07-21T09:25:37Z | |
| dc.date.available | 2026-07-21T09:25:37Z | |
| dc.date.issued | 2026 | |
| dc.description | Presented at: The Fifteenth Language Resources and Evaluation Conference (LREC 2026), 11 - 16 May 2026, Palma, Mallorca, Spain Material complementario: https://f003.backblazeb2.com/file/lrec-media/lrec2026/posters/385.pdf (poster) e https://f003.backblazeb2.com/file/lrec-media/lrec2026/videos/385.mp4 (video) Code is available at https://github.com/palomapiot/hate-eval-agreement/ | |
| dc.description.abstract | [Abstract]: Hate speech spreads widely online and harms both individuals and communities, making automatic detection essential for large-scale moderation. However, accurately detecting hate speech remains a difficult task. Part of the challenge lies in subjectivity: what one person flags as hate speech, another may see as benign. Traditional annotation agreement metrics, such as Cohen’s k, oversimplify this disagreement, treating it as an error rather than meaningful diversity. Meanwhile, Large Language Models (LLMs) promise scalable annotation, but prior studies demonstrate that they cannot fully replace human judgement, especially in subjective tasks. In this work, we reexamine LLM reliability using a subjectivity-aware framework, cross-Replication Reliability (xRR), revealing that even under fairer lens, LLMs still diverge from humans. Yet this limitation opens an opportunity: we find that LLM-generated annotations can reliably reflect performance trends across classification models, correlating with human evaluations. We test this by examining whether LLM-generated annotations preserve the relative ordering of model performance derived from human evaluation (i.e. whether models ranked as more reliable by human annotators preserve the same order when evaluated with LLM-generated labels). Our results show that, although LLMs differ from humans at the instance level, they reproduce similar ranking and classification patterns, suggesting their potential as proxy evaluators. While not a substitute for human annotators, they might serve as a scalable proxy for evaluation in subjective NLP tasks. | |
| dc.description.sponsorship | The authors thank the funding from the Horizon Europe research and innovation programme under the Marie Skłodowska-Curie Grant Agreement No. 101073351. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency (REA). Neither the European Union nor the granting authority can be held responsible for them. The authors thank the financial support supplied by the grant PID2022-137061OB-C21 funded by MI-CIU/AEI/10.13039/501100011033 and by “ERDF/EU”. The authors also thank the funding supplied by the Consellería de Cultura, Educación, Formación Profesional e Universidades (accreditations ED431G 2023/01 and ED431C 2025/49) and the European Regional Development Fund, which acknowledges the CITIC, as a center accredited for excellence within the Galician University System andamember of the CIGUS Network,receives subsidies from the Department of Education, Science, Universities, and Vocational Training of the Xunta de Galicia. Additionally, it is co-financed by the EU through the FEDER Galicia 2021-27 operational program (Ref. ED431G 2023/01). | |
| dc.description.sponsorship | Xunta de Galicia; ED431G 2023/01 | |
| dc.description.sponsorship | Xunta de Galicia; ED431C 2025/49 | |
| dc.description.uri | https://github.com/palomapiot/hate-eval-agreement/ | |
| dc.description.uri | https://f003.backblazeb2.com/file/lrec-media/lrec2026/videos/385.mp4 | |
| dc.identifier.citation | P. Piot, D. Otero, P. Martin-Rodilla, and J. Parapar, "Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection," in Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), Palma, Mallorca, Spain, 2026, pp. 4358-4370. doi: 10.63317/22n5hekovvrz | |
| dc.identifier.doi | 10.63317/22n5hekovvrz | |
| dc.identifier.isbn | 978-2-493814-49-4 | |
| dc.identifier.issn | 2522-2686 | |
| dc.identifier.uri | https://hdl.handle.net/2183/48903 | |
| dc.language.iso | eng | |
| dc.publisher | European Language Resources Association (ELRA) | |
| dc.relation.isversionof | https://f003.backblazeb2.com/file/lrec-media/lrec2026/posters/385.pdf | |
| dc.relation.projectID | info:eu-repo/grantAgreement/EC/HE/101073351 | |
| dc.relation.projectID | info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica, Técnica y de Innovación 2021-2023/PID2022-137061OB-C21/ES/BUSQUEDA, SELECCION Y ORGANIZACION DE CONTENIDOS PARA NECESIDADES DE INFORMACION RELACIONADAS CON LA SALUD - CONSTRUCCION DE RECURSOS Y PERSONALIZACION | |
| dc.relation.uri | https://doi.org/10.63317/22n5hekovvrz | |
| dc.rights | Attribution-NonCommercial 4.0 International | en |
| dc.rights.accessRights | open access | |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc/4.0/ | |
| dc.subject | Hate speech | |
| dc.subject | IAA | |
| dc.subject | xRR | |
| dc.subject | LLMs | |
| dc.subject | System evaluation | |
| dc.title | Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection | |
| dc.type | conference output | |
| dspace.entity.type | Publication | |
| relation.isAuthorOfPublication | 0563c6c3-cd50-4d7d-b11f-127ee297dd6b | |
| relation.isAuthorOfPublication | 00d04042-9b75-419e-9aab-33fd14b201af | |
| relation.isAuthorOfPublication | a1440782-cd8e-4634-b8f3-936cc0220cdb | |
| relation.isAuthorOfPublication | fef1a9cb-e346-4e53-9811-192e144f09d0 | |
| relation.isAuthorOfPublication.latestForDiscovery | 0563c6c3-cd50-4d7d-b11f-127ee297dd6b |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Otero_David_2026_Can_LLMs_Evaluate_What_They_Cannot_Annotate.pdf
- Size:
- 375.61 KB
- Format:
- Adobe Portable Document Format

