Complex noun sequences in Hindi can be formed by the sequences of nouns and genitives. In Hindi, the genitive marker is “kā”, and its allomorphic variations are “ke” and “kī”. When two or more nouns occur without any intervening post-positions, it is known as compound noun. Following are some examples of complex noun sequences: (1) “jilā cunāva adhikārī” (district election officer), (2) “tila kī mit.hāī kī dukāna” (shop of sweets made with sesame) and (3) “upabhoktā adālata ke vakīla” (consumer court’s lawyer). The rightmost noun is the head of the whole construction. The inner structure of the sequence can be quite complex. In it, (a) nouns within the sequence can modify the rightmost head or (b) the local head can modify another local head or the head of the complex noun sequence. For example, in (1), “adhikārī” is the head and both “jilā” and “cunāva” are modifying “adhikārī” thus having a structure (jilā (cunāva adhikārī)). But, the complex sequence in (2) has a structure where “tila” modifies “mit.hāī” and “mit.hāī” in turn modifies “dukāna”. So the structure is ((tila kī mit.hāī) kī dukāna). More number of nouns within a sequence, more complex is the structure. From the Hindi Treebank data, we have obtained 85.37%, 12.54% and 1.80% of the sequences having three, four and five nouns respectively. In this thesis, we attempt to bracket the local sub-structure of a complex noun sequence which is termed as constituency parsing. Constituency parsing recursively builds the inner structure of the complex noun sequence. It is a very significant NLP task because the interpretation of sequence depends on the correct identification of its inner structure. We explore both syntactic and statistical method for predicting the bracketing of the complex noun sequences. In Hindi, the genitive marker agrees with the head of the sub-sequence modified by it. This clue has been used in our syntactic approach. In statistical approach, we have mainly exploited the affinity factor of a head and its modifier based on the frequency of occurring together in the corpus. The method has been augmented by introducing the semantic class information for the head and modifier nouns from Hindi WordNet. Finally, we combine the two methods and implement a hybrid approach for bracketing complex noun sequences. Using this, we have obtained 85.85% accuracy. In this thesis, we show that the identification of the inner structure of complex noun sequence helps in determining the translation. For this experiment, we take three-word noun compounds of English and translate them into Hindi. The strategy of the translation is determined by our observation of EnglishHindi parallel corpora where we observe (and others have reported also) that English licenses multiword noun compound more frequently than what Hindi does. Hindi prefers syntactic phrases where a genitive post-position is inserted between the head and the modifier. In the case of compounds with three
The paper investigates the use of semantic similarity scores as feature in the phrase based machine translation system. We propose the use of partial least square regression to learn the bilingual word embedding using compositional distributional semantics. The model outperforms the baseline system which is shown by an increase in BLEU score. We also show the effect of varying the vector dimension and context window for two different approaches of learning word vectors.
Morphologically rich languages generally require large amounts of parallel data to adequately estimate parameters in a statistical Machine Translation(SMT) system. However, it is time consuming and expensive to create large collections of parallel data. In this paper, we explore two strategies for circumventing sparsity caused by lack of large parallel corpora. First, we explore the use of distributed representations in an Recurrent Neural Network based language model with different morphological features and second, we explore the use of lexical resources such as WordNet to overcome sparsity of content words.
Recent studies in machine translation support the fact that multi-model systems perform better than the individual models. In this paper, we describe a Hindi to English statistical machine translation system and improve over the baseline using multiple translation models. We have considered phrase based as well as hierarchical models and enhanced over both these baselines using a regression model. The system is trained over textual as well as syntactic features extracted from source and target of the aforementioned translations. Our system shows significant improvement over the baseline systems for both automatic as well as human evaluations. The proposed methodology is quite generic and easily be extended to other language pairs as well.
Unfortunately, a wrong country is displayed in the authors’ affiliations. Instead of “Iran” it should read “India”.
We present our efforts on studying the effect of transliteration, on the human readability. We have tried to explore the effect by studying the changes in the eye-gaze patterns, which are recorded with an eye-tracker during experimentation. We have chosen Hindi and English languages, written in Devanagari and Latin scripts respectively. The participants of the experiments are subjected to transliterated words and asked to speak the word. During this, their eye movements are recorded. The eye-tracking data is later analyzed for eye-fixation trends. Quantitative analysis of fixation count and duration as well as visit count is performed over the areas of interest.
This paper deals with the question how to integrate smart devices in Java applications. It will outline how different smart devices can be used to enrich learning environments, we will point to some of the problems one has to face while dealing with smart devices, a differentiation of smart devices will be done and we will give an overview about existing Java Virtual Machines available for different smart devices. Furthermore we will tackle the question of the communication between different smart devices and also between different kinds of smart devices. An outlook to the future work will also be given at the end of this work.
Srinivas Bangalore合作论文数Interactions, LLC1