Integrating external dictionaries into Part-of-speech taggers
| UDC.coleccion | Investigación | es_ES |
| UDC.departamento | Ciencias da Computación e Tecnoloxías da Información | es_ES |
| dc.contributor.author | Graña Gil, Jorge | |
| dc.contributor.author | Chappelier, J. C. | |
| dc.contributor.author | Vilares Ferro, Manuel | |
| dc.date.accessioned | 2005-11-21T13:23:07Z | |
| dc.date.available | 2005-11-21T13:23:07Z | |
| dc.date.issued | 2001 | |
| dc.description.abstract | [Abstract] The highest performances in part-of-speech tagging have been obtained by using stochastic methods, such as hidden Markov models. The running parameters of a hidden Markov model for tagging can be estimated from tagged corpora. However, the current situation in the automatic processing of some languages is very short training texts, but very large dictionaries. These dictionaries can provide very useful information for improving the treatment of unknown words. In this paper we present new strategies for integrating external dictionaries into a stochastic tagging framework. Instead of the most intuitive Adding One method, we propose the use of the Good-Turing formulas, which produce less distortion of the model we are estimating. This technique guarantees good performances in the automatic processing of languages for which reference texts hardly exist. | es_ES |
| dc.description.sponsorship | European Commisision; 1FD97-0047-C04-02 | es_ES |
| dc.description.sponsorship | Ministerio de Educación y Ciencia; TIC2000-0370-C02-01 | |
| dc.description.sponsorship | Xunta de Galicia; PGIDT99XI10502B. | |
| dc.format.mimetype | application/postscript | |
| dc.format.mimetype | text/plain | |
| dc.identifier.citation | Angelova, G.; Bontcheva, K.; Mitkov, R.; Nicolov, N.; Nikolov, N. (eds.), Proceedings of the Euroconference on Recent Advances in Natural Language Processing (RANLP-2001), Tzigov Chark (Bulgaria), pp. 122-128. | es_ES |
| dc.identifier.isbn | 954-90906-1-2 | |
| dc.identifier.uri | http://hdl.handle.net/2183/163 | |
| dc.language.iso | eng | es_ES |
| dc.rights.accessRights | open access | es_ES |
| dc.title | Integrating external dictionaries into Part-of-speech taggers | es_ES |
| dc.type | journal article | es_ES |
| dspace.entity.type | Publication | |
| relation.isAuthorOfPublication | 42896d75-4435-4f99-82e4-48a60a48d799 | |
| relation.isAuthorOfPublication | 3d821e9c-de0b-47cc-a4e0-7c531569602e | |
| relation.isAuthorOfPublication.latestForDiscovery | 42896d75-4435-4f99-82e4-48a60a48d799 |
Files
Original bundle
1 - 1 of 1

