
Unsupervised parsing induction has attracted a significant amount of attention over the last few years. However, current systems exhibit a degree of complexity that can shy away newcomers to the field. We challenge the need for such complexity and present a straightforward weak-EM based system. The results we obtained are close to state-of-the-art ones while still making it extremely simple to experiment with different sub-components. We use a k-best parser, an inductor for Probabilistic Bilexical Grammars (PBGs) [1] and a simple treebank builder. Since our algorithm is independent of the PBG inductor, it overlaps with other models from the literature such as Dependency Model with Valence [2]. Our algorithms are fully fleshed and easily reproducible. We experiment in 8 languages that inform intuitions in training- size dependent parameterization.
This paper describes how the analysis of the lexical structure of the words can be used translate from Thai writing to Braille printing characters. The context free grammar for Thai Braille is defined to recognize the structure of Thai words. The rules are defined according to Thai grammatical structure based on the word pattern consisting of consonant and vowel structure that is how they are formed to become words. The experiments showed that the translation engine can handle more complex word patterns yielding high precision of the translation results.
In this paper we argue for a word-sense based formalization for collocation, and proposes a seed-based approach for collocation extraction for specific purposes. The approach uses RFR_SUM model to iteratively classify polysemous word sense in the corpus. The collocation strength is also obtained by RFR. To capture the syntactic relation inside collocations, this paper presents a frame-based collocation extraction method, which uses word-related frames to obtain collocation with structural information automatically from a large-scale corpus with an average accuracy rate of 89.69%.
Finding pages on the Web that are similar to a query page is an important component of modern search engines. Especially recognition method of content about Web pages is important role in search engine. However, if Web page include query words, it does not necessarily mean that Web page describe query. The main challenge here is identification factors that affect the relationship between query and text. In this paper we tried generation and assigning relationship names between query and text. We call descriptive element as relationship name. Experimental results shows low precision. But we found that if we want to see high precision value, assigning method should not constructed only keywords matching. It need to identify themes of text and extract part of text for that the theme corresponds to query.
Word Sense Disambiguation is a crucial problem in documents whose purpose is to serve as specifications for automatic systems. The combination of different techniques of Natural Language Processing can help in this task. In this paper, we show how to detect ambiguous terms in Software Requirements Specifications. And we propose a computer-aided method that signals the reader for possibly ambiguous usage of terms. The method uses compound term measure (C-value), WordNet semantic similarity (WordNet wup_similarity) and a proposed semantic similarity measure between sentences.
This paper proposes state-oriented approach for knowledge management system, especially for risk management of nuclear power plants. The knowledge model is a kind of state flow in which each state is changed by actions and each action is activated by actors or states. To construct such knowledge base, huge manuals have been analyzed by a unique natural language processing algorithm, called a functional model of language. The authors used this knowledge model and developed an open-source platform for risk management. The system enables to develop procedural knowledge management systems that can check incompleteness, ambiguity and inconsistency of the set of procedures.
A novel speaker verification method is proposed that utilizes pseudo speaker models in the speaker ranking selection (SRS) method. SRS is a recently proposed method that has been shown to give higher performance than the traditional T-norm method. However, the superior performance of SRS is based on utilizing a large number of background speaker models. When enough number of speakers is not available for the background models, the performance of SRS significantly degrades. To achieve higher performance with SRS even when only a small number of speakers are available, we propose to augment the set of background models by adding pseudo speaker models (PSMs). Text-independent speaker verification experiments are performed using a large scale corpus designed for speaker recognition constructed by National Research Institute of Police Science (NRIPS) in Japan. It is shown that the proposed SRS based system with PSMs gives 0.29% equal error rate, which is lower than 0.46% by the original SRS. The minimum DCF scores by the proposed and the original methods are 0.14 and 0.63, respectively.
The purpose of this article is to introduce PSYCHOMS® (Psychiatric Outcome Management System, registered trademark, Tanioka et al.), an electronic nursing management system to facilitate interdisciplinary communication and improve patient outcomes in psychiatric hospitals and report on the agenda for commercialization of the PSYCHOMS® system. Our team has been developing the PSYCHOMS® system since 2006. This system has four major components: (1) Clinical pathway and variance analysis system, (2) Nursing manager and staff's daily recording system, (3) nursing care planning system, and (4) nursing management support system. Any interdisciplinary team member using this system can access the patient's information. Therefore, each interdisciplinary team member's expertise can be maximally utilized to achieve improved patient outcomes. It was necessary to conduct a survey on what standard items were common in different hospitals to allow for the development of PSYCHOMS®'s data base. In order to improve psychiatric care, it is necessary to develop a database that shares common language with all psychiatric hospitals. This manuscript reports on the functions and agenda for the PSYCHOMS® system.
In Japan during recent years, percutaneous coronary intervention (PCI) has been a main therapeutic method for ischemic heart disease (IHD). However, there is little known about early outcomes such as Health-Related Quality of Life (QOL) and life rhythm of daily living (LRDL) post PCI. The purpose of this study is to examine the effectiveness of integrating different types of quantitative and qualitative assessment indicators for patients with ischemic heart disease who underwent PCI. We used the SF36 ver2 (Medical Outcomes study 36- Item Short-Form Health Survey) and Accelerometer (AMI, Actigraph). Patients (n=18) who underwent PCI at one public hospital were assessed on QOL and LRDL during hospitalization to discharge, approximately 10 days. Furthermore, 6 patients were selected for a semi-structure interview about life after the discharge. There were significant differences observed in the items of Sleep Minutes (Smin) and Percentage of Sleep (Pslp) during the test period. The item of Activity Index (Actx) decreased significantly during the period of 1 to 3 days after discharge and during the period 4 to 6 days after discharge, compared with the day before discharge (p <; 0.05). The scores of RE (Role Emotional) significantly decreased a week from discharge, compared with the score at the time of discharge (p <; 0.05). Patients' experiences were divided into five categories utilizing three different quantitative and qualitative assessment indicators for comprehensive evaluation. Basic results indicated: 1) Patient's sleep quality was improved following discharge (evaluated by Actigraph), 2) Patient's QOL was not improved (evaluated by SF36), and 3) The subjects were not fully functioning to their previous level due to decline in physical activity and attitude toward continuation of lifestyle improvement after PCI treatment (semi-structure interviews). The combined uses of three assessment indicators provide comprehensively more accurate evaluation of t- e patient conditions.
In this paper a semantic-tended approach for recombination in EBMT systems is presented. The proposed method uses the least structural information of a sentence. The shared verbs stem between the sentences is the key to find the differences of two sentences. The differences between the input sentence and the matched example sentence are minimized by substituting the translations of their different parts to produce the output. This method is implemented in an EBMT system for English-Farsi language pairs. The results of the proposed method are benchmarked with Google translator. The results are promising based on BLEU and manual evaluations.
Ontologies have served as a knowledge representation about the whole world or some part of it. Building ontologies is a challenging and active research area. Manually constructed Ontologies often have higher quality than the ones created by automatic or semi-automatic approaches but they tend to be more applicable to small domains. Automatic approaches are considered more suitable for building large scale Ontologies where time and efforts of human experts become a bottleneck. For both paradigms, approaches to building Ontologies from Vietnamese texts are still very limited. In this paper, we propose a system that automatically builds Ontology from Vietnamese texts using cascades of annotation-based grammars. Obtained experimental results on a university organizational structure domain are very promising.
In recent decades, with the continuous improvement of computer performance, in medicine, biology and other fields, its applications are paid more and more attention to by researchers and scholars. Aided medical diagnosis and treatment system based on computer has already become a hotspot. In this paper, the aided medical diagnosis system based on knowledge base and maximum entropy is put forward. And this system will become important aids of medical diagnosis and treatment.
In this paper a new method to detect cross-language plagiarism in electronic documents is presented. The method is based on natural language processing techniques and machine learning methods. Preliminary results show that the system is able to identify the Internet sources of plagiarism of texts written in Spanish that have been copied (by humans and machines) from English sources.
To distinguish facts from unreliable or uncertain information, hedges have to be identified. This paper presents an approach to hedges scope detection based on phrase structures and dependency structures. First, phrase structures and dependency structures are used for hedges scope detection respectively. Phrase structures are adapted as important features for hedges scope detection by a machining learning method. Dependency structures are used to detect hedges scope by a rule-based method. Then, the phrase-based system and the dependency-based system are combined by a Conditional Random Field (CRF)-based model, which simply extends the feature vectors with the scope tags generated by the two individual phrase-based and dependency-based systems. Experiments on the CoNLL-2010 biological corpus show that our model achieves F-scores of 55.47% on hedges scope detection based on phrase structures using machine learning and 55.67% based on dependency structures using manual rules, and 58.97% based on dependency structures and phrase structures using our combined method. The analysis results show that phrase structures and dependency structures are both effective for hedges scope detection and their combination can improve the scope detection performance further.
Dependency structural tree is widely used in the research of dependency grammar and parsing, and it is one of the important forms of the syntactic structure of natural language. The research of the theoretical forms and their probabilities are of great importance to probability-based disambiguation of parsing. In this paper we propose the formula and its algorithm to calculate the probability of any well-formed dependency structural tree.
A study of a communication robot that aims to ease stress and to heal is very important. This paper develops a dialogue robot with haptic interactions, and two experiments are carried out to estimate the effects of dialogue and haptic interactions. From experimental results, the haptic interactions have more effective than dialogue in the interest conversation and in the robustness.
This paper describes the structure and function of conceptual dictionary placing in our narrative generation system architecture to provide knowledge of verb concepts and noun concepts associated with an event concept that is a basic constitutional unit in narrative. Currently, verb concept dictionary has 5337 case frames and modified 1158 constraints, and noun concept dictionary contains 142168 noun concepts including 5770 intermediate concepts. After we explain the structure and development process, we introduce the role and function in narrative generation process based on two application systems.
This study examined the effects of hand massage on autonomic activity, anxiety, relaxation and sense of affinity by performing it to healthy people before applying the technic in actual clinical practice. Findings were showed below: 1) the significant increase in the pNN50 and the significant decrease in the heart rate meant the intervention of massage increased the autonomic nervous activity, improved the parasympathetic nerve activity and reduced the sympathetic nerve activity. This means the subjects were considered to be in a state of relaxation. 2) Salivary a amylase has been reported as a possible indicator for sympathetic nerve activity. In the results of this study, there was no significant difference in the salivary a amylase despite a decrease after massage. 3) State anxiety score is to measure temporal situational reactions while being in the state of anxiety. This score decreased significantly after massage, which was a result similar to those in a previous study. 4) The level of willingness to communicate with other person and the sense of affinity toward the massage-performer had a positive change of 70 percent. From this, it can be considered that a comfortable physical contact between a patient and a nursing profession, who are in a supported-supportive relationship, leads to an effect of shortening the gap in their psychological distance.
Protein-Protein Interaction (PPI) extraction from biomedicine literature can supply the biomedicine researcher with useful information rapidly. This paper presents a PPI extraction system based on the ensemble kernel model and active learning. Firstly, the ensemble kernel within SVM classifier combines the lexical feature-based kernel and the path-based kernel. Experimental results show that the F-score of PPI extraction using ensemble kernel model on AIMED, IEPA and BCPPI corpora are 64.50%, 69.74% and 60.38% respectively with 10-fold cross-validation, which are better than the lexical feature-based kernel and the path-based kernel separately. As the above ensemble kernel model based on SVM needs large labeled data and it is expensive to label data manually, we integrate active learning into the ensemble kernel model. The active learning method uses the uncertainty-based sampling strategy. The experimental results integrating the active learning show that the F-score on AIMED, IEPA and BCPPI corpora are 65.24%, 70.19% and 61.87% respectively, which are better than those using the ensemble kernel model with the passive learning, and meantime reduce the labeling data by 20%, 30% and 30%, respectively.
A new fuzzy reasoning method called the reverse universal triple I method of (1,1,2) type (reverse universal triple I method for short) is proposed and investigated by means of the famous Lukasiewicz implication, which contains the reverse triple I method as its special case. The reverse triple I principles are improved, and the optimal solutions of reverse universal triple I method are obtained for Fuzzy Modus Ponens (FMP) and Fuzzy Modus Tollens (FMT). Furthermore, it is found that the reverse universal triple I method has the reversibility property. Lastly, the reverse universal triple I method is further generalized to the alpha-reverse universal triple I method, and the related principles and optimal solutions are achieved.