GraƱa Gil, JorgeChappelier, J. C.Vilares Ferro, Manuel2005-11-212005-11-212001Angelova, G.; Bontcheva, K.; Mitkov, R.; Nicolov, N.; Nikolov, N. (eds.), Proceedings of the Euroconference on Recent Advances in Natural Language Processing (RANLP-2001), Tzigov Chark (Bulgaria), pp. 122-128.954-90906-1-2http://hdl.handle.net/2183/163[Abstract] The highest performances in part-of-speech tagging have been obtained by using stochastic methods, such as hidden Markov models. The running parameters of a hidden Markov model for tagging can be estimated from tagged corpora. However, the current situation in the automatic processing of some languages is very short training texts, but very large dictionaries. These dictionaries can provide very useful information for improving the treatment of unknown words. In this paper we present new strategies for integrating external dictionaries into a stochastic tagging framework. Instead of the most intuitive Adding One method, we propose the use of the Good-Turing formulas, which produce less distortion of the model we are estimating. This technique guarantees good performances in the automatic processing of languages for which reference texts hardly exist.application/postscripttext/plainengIntegrating external dictionaries into Part-of-speech taggersjournal articleopen access