A linguistic approach for determining the topics of Spanish Twitter messages

UDC.coleccionInvestigaciónes_ES
UDC.departamentoLetrases_ES
UDC.endPage145es_ES
UDC.grupoInvLingua e Sociedade da Información (LYS)es_ES
UDC.issue2es_ES
UDC.journalTitleJournal of Information Sciencees_ES
UDC.startPage127es_ES
UDC.volume41es_ES
dc.contributor.authorVilares, David
dc.contributor.authorAlonso, Miguel A.
dc.contributor.authorGómez-Rodríguez, Carlos
dc.date.accessioned2024-01-18T15:53:17Z
dc.date.available2024-01-18T15:53:17Z
dc.date.issued2015
dc.descriptionThis manuscript version is made available under the CC-BY-NC-ND 4.0 license https://creativecommons.org/licenses/by-nc-nd/4.0/. This version of the article: Vilares, D., Alonso, M. A., & Gómez-Rodríguez, C. (2015). ‘A linguistic approach for determining the topics of Spanish Twitter messages’ has been accepted for publication in Journal of Information Science, 41(2), 127-145. Copyright © 2014 The Authors. DOI: https://doi.org/10.1177/0165551514561652.es_ES
dc.description.abstract[Abstract]: The vast number of opinions and reviews provided in Twitter is helpful in order to make interesting findings about a given industry, but given the huge number of messages published every day, it is important to detect the relevant ones. In this respect, the Twitter search functionality is not a practical tool when we want to poll messages dealing with a given set of general topics. This article presents an approach to classify Twitter messages into various topics. We tackle the problem from a linguistic angle, taking into account part-of-speech, syntactic and semantic information, showing how language processing techniques should be adapted to deal with the informal language present in Twitter messages. The TASS 2013 General corpus, a collection of tweets that has been specifically annotated to perform text analytics tasks, is used as the dataset in our evaluation framework. We carry out a wide range of experiments to determine which kinds of linguistic information have the greatest impact on this task and how they should be combined in order to obtain the best-performing system. The results lead us to conclude that relating features by means of contextual information adds complementary knowledge over pure lexical models, making it possible to outperform them on standard metrics for multilabel classification tasks.es_ES
dc.description.sponsorshipThe research reported in this article was partially funded by Ministerio de Economía y Competitividad and FEDER (grant TIN2010-18552-C03-02), Ministerio de Educación, Cultura y Deporte (FPU13/01180) and by Xunta de Galicia (Grants CN2012/008, CN2012/319).es_ES
dc.description.sponsorshipXunta de Galicia; CN2012/008es_ES
dc.description.sponsorshipXunta de Galicia; CN2012/319es_ES
dc.identifier.citationVilares, D., Alonso, M. A., & Gómez-Rodríguez, C. (2015). A linguistic approach for determining the topics of Spanish Twitter messages. Journal of Information Science, 41(2), 127-145. https://doi.org/10.1177/0165551514561652es_ES
dc.identifier.doi10.1177/0165551514561652
dc.identifier.issn0165-5515
dc.identifier.issn1741-6485
dc.identifier.urihttp://hdl.handle.net/2183/34987
dc.language.isoenges_ES
dc.publisherSAGE Publications & CILIPes_ES
dc.relation.isversionofhttps://doi.org/10.1177/0165551514561652
dc.relation.projectIDinfo:eu-repo/grantAgreement/MICINN/Plan Nacional de I+D+i 2008-2011/TIN2010-18552-C03-02/ES/ANALISIS DE TEXTOS Y RECUPERACION DE INFORMACION PARA LA MINERIA DE OPINIONES: ANALISIS DE ENUNCIADOS Y EXTRACCION DE RELACIONESes_ES
dc.relation.projectIDinfo:eu-repo/grantAgreement/MECD/Plan Estatal de Investigación Científica y Técnica y de Innovación 2013-2016/FPU13%2F01180/ES/es_ES
dc.relation.urihttps://doi.org/10.1177/0165551514561652es_ES
dc.rightsAtribución-NoComercial-SinDerivadas 4.0 Internacionales_ES
dc.rights.accessRightsopen accesses_ES
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/3.0/es/*
dc.subjectTwitteres_ES
dc.subjectNatural language processinges_ES
dc.subjectMulti-label topic classificationes_ES
dc.titleA linguistic approach for determining the topics of Spanish Twitter messageses_ES
dc.typejournal articlees_ES
dspace.entity.typePublication
relation.isAuthorOfPublication37dabbe9-f54f-43bb-960e-0bf3ac7e54eb
relation.isAuthorOfPublication1318edb8-3967-465c-a267-146624c05837
relation.isAuthorOfPublicatione70a3969-39f6-4458-9339-3b71756fa56e
relation.isAuthorOfPublication.latestForDiscovery37dabbe9-f54f-43bb-960e-0bf3ac7e54eb

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Vilares_David_2015_A_linguistic_approach_for_determining_the_topics_of_Spanish_Twitter_messages.pdf
Size:
660.34 KB
Format:
Adobe Portable Document Format
Description: