A critical edition takes into account various versions of the same text in order to show the differences between two distinct versions, in terms of words that have been missing, changed, omitted or displaced. Traditionally, Sanskrit is written without spaces between words, and the word order can be changed without altering the meaning of a sentence. This paper describes the characteristics which make Sanskrit text comparisons a specific matter. It presents two different methods for comparing Sanskrit texts, which can be used to develop a computer assisted critical edition. The first one method uses the L.C.S., while the second one uses the global alignment algorithm. Comparing them, we see that the second method provides better results, but that neither of these methods can detect when a word or a sentence fragment has been moved. We then present a method based on N-gram that can detect such a movement when it is not too far from its original location. We show how the method behaves on several examples. Index Terms—Sanskrit, text alignment.
Clustering is one of the most common operation in data analysis while constrained is not so common. We present here a clustering method in the framework of Symbolic Data Analysis (S.D.A) which allows to cluster Symbolic Data. Such data can be constrained relations between the variables, expressed by rules which express the domain knowledge. But such rules can induce a combinatorial increase of the computation time according to the number of rules. We present in this paper a way to cluster such data in a quadratic time. This method is based first on the decomposition of the data according to the rules, then we can apply to the data a clustering algorithm based on dissimilarities.
Traditionally Sanskrit is written without blank, sentences can make thousands of characters without any separation. A critical edition takes into account all the different known versions of the same text in order to show the differences between any two distinct versions, in term of words missing, changed or omitted. This paper describes the Sanskrit characteristics that make text comparisons different from other languages, and will present different methods of comparison of Sanskrit texts which can be used for the elaboration of computer assisted critical edition of Sanskrit texts. It describes two sets of methods used to obtain the alignments needed. The first set is using the L. C. S., the second one the global alignment algorithm. One of the methods of the second set uses a classical technique in the field of artificial intelligence, the A* algorithm to obtain the suitable alignment. We conclude by comparing our different results in term of adequacy as well as complexity.
Dealing with multi-valued data has become quite common in both the framework of databases as well as data analysis. Such data can be constrained by domain knowledge provided by relations between the variables and these relations are expressed by rules. However, such knowledge can introduce a combinatorial increase in the computation time depending on the number of rules. In this paper, we present a way to cluster such data in polynomial time. The method is based on the following: a decomposition of the data according to the rules, a suitable dissimilarity function and a clustering algorithm based on dissimilarities.
This paper shows a fuzzy relational clustering method in order to perform the clustering of symbolic data. The presented method yields a fuzzy partition and prototype for each cluster by optimizing an adequacy criterion based on suitable dissimilarity measures. This work considers two volume-based measures that may be applied to data described by set-valued, list-valued or interval-valued symbolic variables. Experiments with real and synthetic symbolic data sets show the usefulness of the proposed approach. The accuracy of the results were assessed by the corrected Rand index and the overall error rate of classification.
A critical edition takes into account all the different known versions of the same text in order to show the differences between any two distinct versions. The construction of a critical edition is a long and, sometimes, tedious work. Some software that help the philologist in such a task have been available for a long time for the European languages. However, such software does not exist yet for the Sanskrit language because of its complex graphical characteristics that imply computationally expensive solutions to problems occurring in text comparisons.This paper describes the Sanskrit characteristics that make text comparisons different from other languages, presents computationally feasible solutions for the elaboration of the computer assisted critical edition of Sanskrit texts, and provides, as a byproduct, a distance between two versions of the edited text. Such a distance can then be used to produce different kinds of classifications between the texts.
To exhibit the differences between different versions of the same text is quite an easy operation for a computer program if it concerns texts written in any modern European language such as English or French, where words are well separated by spaces. On the contrary, if it concerns a language such as Sanskrit, where traditionally no spaces occur between words, and some very long compound words exist, it becomes a difficult problem, especially if there is no valid and reliable computer oriented lexicon available. In this article we intend to show how we proceed to determine, in terms of words, the differences that can exist between the different versions of the same Sanskrit text, with the aim to elaborate a critical edition. Exhibiting these differences, it becomes possible to elaborate a distance between the different versions of the text. Once this step is achieved, the study and determination of relations, which can exist between the versions, becomes possible through the use of phylogenic trees.
The recording of symbolic data has become a common practice with the advances in database technologies. This paper shows hard and fuzzy relational clustering in order to partition symbolic data. These methods optimize objective functions based on a dissimilarity function. The distance used is a volume based measure and may be applied to data described by set-valued, list-valued or interval-valued symbolic variables. Experiments with real and synthetic symbolic data sets show the usefulness of the proposed approach.
A critical edition takes into account all the different known versions of the same text in order to show the differences related to any two distinct versions. The construction of a critical edition is a long and, sometimes, tedious work. In order to make it easier, softwares helping the philologist are nowadays available for the European languages. Because of its complex graphical characteristics, which involve computationally expensive solutions to problems occurring in text comparisons, such softwares do not yet exist for Sanskrit language. This paper describes the Sanskrit characteristics that make text comparisons different, presents computationally feasible solutions for the elaboration of the computer assisted critical edition of Sanskrit texts, and provides, as a byproduct, a distance between two versions of the edited text.
Face a une collection de manuscrits retranscrivant un meme texte, frequemment disperses dans le temps et l’espace, le chercheur a recours a l’edition critique afin de prendre en compte tous les aspects du texte. En vue de faciliter leurs creations et d'offrir de nouvelles possibilites en matiere d’analyse et d’interactivite, quelques editions critiques informatisees ont deja ete proposees. Dans cet article, nous nous interessons au cas de l’edition critique (electronique) de manuscrits ecrits en sanskrit. Nous commencons par decrire les specificites du sanskrit qui soulevent un ensemble de problemes touchant aussi bien a la typographie qu’a l'informatique et l'analyse de donnees. Puis, nous presentons brievement les solutions informatiques actuelles permettant d’adapter les traitements a la typographie et la syntaxe particuliere au sanskrit. Enfin, nous proposons une approche automatique pour identifier et evaluer les differences entre deux manuscrits en vue notamment de decrire les relations de filiation entre manuscrits connus d'un meme texte. D'un point de vue algorithmique, notre approche est basee sur des techniques habituellement employees pour comparer des chaines moleculaires.
Brigitte Trousse合作论文数INRIA Sophia Antipolis11
Doru Tanasa合作论文数AxIS Team-Project, INRIA Sophia Antipolis3