One essential task in extracting information from biomedical literature is the bio Named Entity Recognition (NER) process, which basically defines the boundaries between typical words and biomedical terminology in particular text data, and assigns them based on domain knowledge. This paper presents a semi supervised integration of completely different classifiers to cover knowledge from unlabeled data to recognize bio named entities in text. We modified the original co-training, a semi supervised learning algorithm, with a scalable feature processing schema, which extracts the bio NER feature from a number of unlabeled data and converts different types of feature sets. Our base result shows that the classifiers of co-training achieve significant learning from unlabeled data.
更多
查看译文
关键词
unlabeled data,particular text data,bio NER feature,feature set,scalable feature processing schema,biomedical literature,biomedical terminology,different classifier,different type,domain knowledge,Co-training Algorithm,Entity Recognition