For the ongoing creation of a database of Spanish verbs, we developed a methodological proposal for the automatic tagging of semantic types in running text according to Hanks' Corpus Pattern Analysis (CPA) guidelines. In this task, a text document is the input and the output is the tagging of each noun, noun phrase or proper noun with one of the semantic types in the CPA Ontology. The present proposal is based on a combination of algorithms for automatic ontology population, named entity recognition and, most importantly, word sense disambiguation, to assign the appropriate type to a noun according to the context. The paper includes an evaluation of the method tagging a random sample of 200 Wikipedia pages in Spanish and English. Evaluation figures by a panel of three experts show 84% precision and 88% recall in Spanish and 83% precision and 93% recall in English. These are competitive results considering the simplicity and computational efficiency of the algorithm.
更多
查看译文
关键词
semantic typing,semantic tagging,word sense disambiguation,corpus pattern analysis,named entity recognition