Dilin modellenmesi, dil çalışmalarında önemli bir temel olarak yer alır. Farklı modelleme yöntemleri, farklı diller için uyarlanabilir olsa da bu uyarlamalar, hedef dil için her zaman yeterli olmayabilir. Bu durumdan en çok biçimbirimsel açıdan zengin diller etkilenir. Böyle bir dil için hazırlanacak model kurgulanırken dilin evrensel olarak ortak olan özelliklerinin yanı sıra, dilin kendine özgü özelliklerine odaklanılmalıdır. Bu makalede, bağımlı biçimbirim bakımından zengin bir görünüm sunan Türkçe ele alınarak uyarlanan gramer sunulmuştur. Çalışmada açıklanan gramer temelleri geleneksel üretici gramer yönteminden uyarlanmıştır. Bununla birlikte, sunulan gramer, biçimbirimleri söz dizimi elemanı olarak geleneksel söz dizimi elemanlarıyla birlikte, söz dizimine olan etkilerini ele almasıyla ve kullanılan özel bileşen kümesiyle geleneksel üretici gramer yöntemden ayrılır. Geleneksel yöntemden farklı olarak önerilen gramerde, tümce çözümlemesine sözcüklerden değil, biçimbirim elemanları olan sözcük gövdeleri, ekler, biçimbirimler ve bu gibi elemanların oluşturduğu gruplardan başlanır. Buna ek olarak Türkçenin söz dizimsel ve birimbirimsel özelliklerine göre kurgulanan bir bileşen kümesi de sunulmuştur. Sunulan bileşen kümesi, tümce, ad öbeği, eylem öbeği, belirteç öbeği gibi geleneksel sözdizimsel bileşenleri, öbek gövdesi olarak adlandırılan ara bir yapıyı ve çoğul eki, durum eki, zaman çekimi eki gibi, biçimbirimleri veya biçimbirim gruplarını temsil eden bileşenleri içerir.
The observation that some Turkish nouns license accusative marking of their complements has puzzled linguistics for the last three decades. The first and the most common hypothesis is that such nouns are part of light-verb compounds in their infinitive nominal form where the verbal part is dropped. In this paper, I argue that light-verb hypothesis is incorrect by providing several constructions where it fails to account for. I propose an alternative, two-part hypothesis. The first part is independent of the light-verb constructions and claims that in Turkish, the accusative case is licensed by verb stems rather than finite verbs. This is a natural corollary of a yet larger hypothesis about the structure of syntax trees for languages like Turkish which claims that the syntax trees go deeper than the word level. The leaves of such trees are morphemes or morpheme groups and the whole of infection/derivation is considered within a single morphosyntactic structure. The implication is that as far as the case-licensing is concerned, only the verb stems license accusative case. In constructions where a deverbal noun seems to be licensing accusative case, it is actually the verb stem in the noun’s derivation tree that licenses the accusative case. The second part of the hypothesis is the claim that the case-licensing nouns are actually lexicalizations involving multiple adjacent leaves in the morphosyntactic tree. The most common manifestations of this are the lexicalizations of verb stem-participle pairs through replacement with templatic surface forms borrowed from Arabic. The new proposal naturally explains the observation that accusative licensing is most common for a small set of words borrowed from Arabic. I illustrate that these words are the templatic forms of their multi-morpheme counterparts in Turkish which correspond to deverbal nouns, adjectives or adverbs.
In this paper, we describe the steps in a visual modeling of Turkish morphology using diagramming tools. We aimed to make modeling easier and more maintainable while automating much of the code generation. We released the resulting analyzer, MorTur, and the diagram conversion tool, DiaMor as free, open-source utilities. MorTur analyzer is also publicly available on its web page as a web service. MorTur and DiaMor are part of our ongoing efforts in building a set of natural language processing tools for Turkic languages under a consistent framework.
In this study, a new content based image retrieval (CBIR) method, which uses HSV histogram data is proposed. The model uses the HSV histogram to find the background from the image by analyzing the peaks in the histogram data and performing a moving window algorithm to identify the region within the histogram that belongs to the background colors. After identifying the background information, the sections of the image that are part of the background are removed from the original image and the remaining foreground or content information is stored for comparison with other images. In order to verify the methodology, a graphical user interface is developed and 1000 different images from 10 different groups from the coral database are put into the image database for comparison. The analysis and preliminary tests show that comparing only the foreground information of the images pro- vided better results than comparing images themselves, especially when searching for particular objects within the images. This algorithm can also be used as a background elimination technique to reduce the storage requirements of images and the comparison time between images can be reduced significantly.
MorAz is an open-source morphological analyzer for Azerbaijani Turkish. The analyzer is available through both as a website for interactive exploration and as a RESTful web service for integration into a natural language processing pipeline. MorAz implements the morphology of Azerbaijani Turkish following a two-level approach using Helsinki finite-state transducer and wraps the analyzer with python scripts in a Django instance.
In this article, we summarize the methodology and the results of our 2-year-long efforts to construct a comprehensive WordNet for Turkish. In our approach, we mine a dictionary for synonym candidate pairs and manually mark the senses in which the candidates are synonymous. We marked every pair twice by different human annotators. We derive the synsets by finding the connected components of the graph whose edges are synonym senses. We also mined Turkish Wikipedia for hypernym relations among the senses. We analyzed the resulting WordNet to highlight the difficulties brought about by the dictionary construction methods of lexicographers. After splitting the unusually large synsets, we used random walk–based clustering that resulted in a Zipfian distribution of synset sizes. We compared our results to BalkaNet and automatic thesaurus construction methods using variation of information metric. Our Turkish WordNet is available online.
We present the results of an experiment to assess the validity of prior polarities available in sentiment lexicons. We designed a ranking task that was elicited through pairwise comparisons and compared the results to those predicted by two popular sentiment lexicons. We find that the experiment results show a moderate level of agreement between the lexicons and human judgments.
We give a FST description of nominal and finite verb morphology of Azarbaijani Turkish.We use a hybrid approach where nominal inflection is expressed as a slot-based paradigm and major parts of verb inflection are expressed as optional paths on the FST.We collapse adjective and noun categories in a single nominal category as they behave similarly as far as their paradigms are concerned.Thus, we defer a more precise identification of POS to further down the NLP pipeline.
In this study, a rule based application that collects news from public network sources and automatically tags these news gathered has been implemented. The sub-goal of the study is to measure which features are more appropriate for tagging. The success rate of each feature was measured on 100 manually tagged news.
In this paper, we report our tree based statistical translation study from English to Turkish. We describe our data generation process and report the initial results of tree-based translation under a simple model. For corpus construction, we used the Penn Treebank in the English side. We manually translated about 5K trees from English to Turkish under grammar constraints with adaptations to accommodate the agglutinative nature of Turkish morphology. We used a permutation model for subtrees together with a word to word mapping. We report BLEU scores under simple choices of inference algorithms.
In this paper, we describe our initial efforts for creating a Turkish constituency parse treebank by utilizing the English Penn Treebank. We employ a semi-automated approach for annotation. In our previous work [18], the English parse trees were manually translated to Turkish. In this paper, the words are semi-automatically annotated morphologically. As a second step, a rule-based approach is used for refining the parse trees based on the morphological analyses of the words. We generated Turkish phrase structure trees for 5143 sentences from Penn Treebank that contain fewer than 15 tokens. The annotated corpus can be used in statistical natural language processing studies for developing tools such as constituency parsers and statistical machine translation systems for Turkish.
In this paper, we report our work on chunking in Turkish. We used the data that we generated by manually translating a subset of the Penn Treebank. We exploited the already available tags in the trees to automatically identify and label chunks in their Turkish translations. We used conditional random fields (CRF) to train a model over the annotated data. We report our results on different levels of chunk resolution.
In this paper, we analyze the errors of NP attachments that occur when they combine with NP's ending in possessor clitic.We suggest a simple pattern based detection and correction solution.We illustrate the errors with examples from Penn Treebank.
In this paper, we report our preliminary efforts in building an English-Turkish parallel treebank corpus for statistical machine translation. In the corpus, we manually generated parallel trees for about 5,000 sentences from Penn Treebank. English sentences in our set have a maximum of 15 tokens, including punctuation. We constrained the translated trees to the reordering of the children and the replacement of the leaf nodes with appropriate glosses. We also report the tools that we built and used in our tree translation task.
In the software engineering world, creating and maintaining relationships between byproducts generated during the software lifecycle is crucial. A typical relation is the one that exists between an item in the requirements document and a block in the subsequent system design, i.e. class in the source code. In many software engineering projects, the requirement documentation is prepared in the language of the developers, whereas developers prefer to use the English language in the software development process. In this paper, we use the vector space model to extract traceability links between the requirements written in one language (Turkish) and the implementations of classes in another language (English). The experiments show that, by using a generic translator such as Google translate, we can obtain promising results, which can also be improved by using comment info in the source code.
Despite being an active research area for the last two decades, chaotic cryptography has not found a place in mainstream cryptography literature. This paper explores the reasons for this separation.
In this paper, we provide an algebraic cryptanalysis of a recently proposed chaotic image cipher. We show that the secret parameters of the algorithm can be revealed using chosen-plaintext attacks. Our attack uses the orbit properties of the permutation maps to deduce encryption values for a single round. Once a single round encryption is revealed, the secret parameters are obtained using simple assignments.
We report a break for a recently proposed class of cryptosystems. The cryptosystem uses constant points of a periodic secret orbit to encrypt the plaintext. In order to break the system, it suffices to sort the constant points and find the initial fixed point. We also report breaks for modified versions of the cryptosystem. In addition, we discuss some efficiency issues of the cryptosystem.
Cryptanalysis is an integral part of any serious effort in designing secure encryption algorithms. Indeed, a cryptosystem is only as secure as the most powerful known attack that failed to break it. The situation is not different for chaos-based ciphers. Before attempting to design a new chaotic cipher, it is essential that the designers have a thorough grasp of the existing attacks and cryptanalysis tools.
We cryptanalyze Fridrich’s chaotic image encryption algorithm. We show that the algebraic weaknesses of the algorithm make it vulnerable against chosen-ciphertext attacks. We propose an attack that reveals the secret permutation that is used to shuffle the pixels of a round input. We demonstrate the effectiveness of our attack with examples and simulation results. We also show that our proposed attack can be generalized to other well-known chaotic image encryption algorithms.
William E. Leithead合作论文数2