Use this link to cite:
https://hdl.handle.net/2183/49331 Modelado y análisis lingüístico de las intervenciones en el Congreso usando técnicas de PLN
Loading...
Identifiers
Publication date
Authors
González Rodríguez, Inés
Advisors
Other responsabilities
Universidade da Coruña. Facultade de Informática
Journal Title
Bibliographic citation
Type of academic work
Academic degree
Abstract
[Resumen] El debate parlamentario es la base de la actividad política en España, a partir de la cual se genera una gran cantidad de textos difícil de analizar manualmente. Para facilitar dicha tarea, este trabajo aplica técnicas de Procesamiento del Lenguaje Natural que permiten analizar de manera sistemática el alto volumen de datos. Este trabajo incluye la preparación de un corpus de intervenciones procedentes de la web del Congreso de los Diputados; donde se integra el texto del discurso con los metadatos asociados como el orador, el partido político y la legislatura. Sobre este conjunto de datos se realiza un análisis orientado a extraer la estructura morfosintáctica de las oraciones (las categorías gramaticales y las dependencias sintácticas). Asimismo, con el objetivo de profundizar en la semántica del discurso, se lleva a cabo un análisis de sentimiento enfocado en el tono, las emociones y la hostilidad del lenguaje. A su vez, el análisis de contenido permite identificar los temas esenciales tratados en el Congreso y evaluar la similitud entre los discursos de partidos y oradores. Los resultados se integran en una aplicación web para facilitar la exploración e interpretación de los mismos. Pese a la homogeneidad de los discursos, marcada por el contexto institucional, se identifican rasgos lingüísticos particulares en determinados partidos y oradores.
[Abstract] Parliamentary debate plays a central role in politics in Spain, which generates a large volume of textual data that is difficult to analyze manually. To address this issue, this work applies Natural Language Processing techniques, which enable the systematic analysis of large volumes of data. This study includes the construction of a corpus of speeches collected from the website of the Congress of Deputies, enriching each speech with associated metadata such as the speaker, political party and legislative term. Using this dataset, a linguistic analysis is performed with the aim of extracting the morphosyntactic structure of sentences, including grammatical categories and syntactic dependencies. Furthermore, in order to further examine the semantics of the discourse, a sentiment analysis is conducted focusing on tone, emotions and hostility in language use. In addition, content analysis is used to identify the main topics addressed in Congress and to assess the similarity of speeches at both the party and speaker levels. The results are integrated into a web application to facilitate their exploration and interpretation. Despite the homogeneity of the speeches, characterized by the institutional context, particular linguistic patterns emerge across certain political parties and speakers.
[Abstract] Parliamentary debate plays a central role in politics in Spain, which generates a large volume of textual data that is difficult to analyze manually. To address this issue, this work applies Natural Language Processing techniques, which enable the systematic analysis of large volumes of data. This study includes the construction of a corpus of speeches collected from the website of the Congress of Deputies, enriching each speech with associated metadata such as the speaker, political party and legislative term. Using this dataset, a linguistic analysis is performed with the aim of extracting the morphosyntactic structure of sentences, including grammatical categories and syntactic dependencies. Furthermore, in order to further examine the semantics of the discourse, a sentiment analysis is conducted focusing on tone, emotions and hostility in language use. In addition, content analysis is used to identify the main topics addressed in Congress and to assess the similarity of speeches at both the party and speaker levels. The results are integrated into a web application to facilitate their exploration and interpretation. Despite the homogeneity of the speeches, characterized by the institutional context, particular linguistic patterns emerge across certain political parties and speakers.
Description
Keywords
Procesamiento del Lenguaje Natural Modelos de lenguaje Congreso de los Diputados Análisis morfosintáctico Análisis de sentimiento Modelado de temas Similitud textual Informes interactivos Natural Language Processing Language models Congress of Deputies Morphosyntactic analysis Sentiment analysis Topic Modeling Textual similarity Interactive reports
Editor version
Rights
Attribution 4.0 International








