A systematic approach to the random generation of labelled combinatorial objects is presented. It applies to structures that are decomposable, i.e., formally specifiable by grammars involving union, product, set, sequence, and cycle constructions. A general strategy is developed for solving the random generation problem with two closely related types of methods: for structures of size n, the boustrophedonic algorithms exhibit a worst-case behaviour of the form 𝒪( nlog n) ; the sequential algorithms have worst case 𝒪( n^2 ) , while offering good potential for optimizations in the average case. (Both methods appeal to precomputed numerical tables of linear size). A companion calculus permits to systematically compute the average case cost of the sequential generation algorithm associated to a given specification. Using optimizations dictated by the cost calculus, several random generation algorithms are developed, based on the sequential principle; most of them have expected complexity 1/2 n log n, thus being only slightly superlinear. The approach is exemplified by the random generation of a number of classical combinatorial structures including Cayley trees, hierarchies, the cycle decomposition of permutations, binary trees, functional graphs, surjections, and set partitions.
This paper studies the random indexed dendograms produced by agglomerative hierarchical algorithms under the non-classifiability hypothesis of independent identically distributed (i.i.d.) dissimilarities. New tests for classifiability are deduced. The corresponding test statistics are random variables attached to the indexed dendrograms, such as the indices, the survival time of singletons, the value of the ultrametric between two given points, or the size of classes in the different levels of the dendogram. For an indexed dendogram produced by the Single Link method on i.i.d. dissimilarities, the distribution of these random variables is computed, thus leading to explicit tests. For the case of the Average and Complete Link methods, some asymptotic results are presented. The proofs rely essentially on the theory of random graphs.
We propose statistical tests to decide between an hypothesis of non classifiability of the data against the presence of a classification in some simple situations. We consider only the single link algorithm and two null hypotheses of non classifiability, according to whether the dissimilarities or the objects themselves are i.i.d. random variables. Each choice for the distribution of the input of the single link algorithm induces a different probability distribution on the output, which is a random indexed dendrogram. Certain characteristics of these random indexed dendrograms are studied, and their asymptotic distributions computed under each null hypothesis. All these random variables can be used to define statistical tests. Explicit examples of such tests are provided.
This paper studies the dendrograms produced by algorithms of classification such as the Single Link Algorithm. We introduce probability distributions on dendrograms corresponding to distinct non classifiability hypotheses. The distributions of the height of a random dendrogram under these hypotheses are studied and their asymptotics explicitly computed. This leads to statistical tests for non-classifiability.
Critchley and Van Cutsem [7] recently developed the properties of dissimilarities and ultrametrics with values in an ordered set. This theoretical definition allows consideration of a pair of ultrametrics as a two dimensional ultrametric with values in ℝ+ × ℝ+, and provides a good framework to study the dependence between two ultrametrics. In this paper, we just present the basic definitions for introducing the definition of some new indices of dependence or of comparison of two real-valued ultrametrics. Details can be found in Benkaraache’s thesis [4].
Classifying objects according to their likeness seems to have been a step in the human process of acquiring knowledge, and it is certainly a basic part of many of the sciences. Historically, the scien
The primary objective of this and the following chapter is to show that one simple results unifies and generalises a number of familiar bijections in the mathematical classification literature. These include the classical bijections concerning ultrametrics on a finite set, obtained in Benzécri (1965), Hartigan (1967), Jardine, Jardine and Sibson (1967) and Johnson (1967), as well as their several later extensions reported in Jardine and Sibson (1971), Janowitz (1978), Barthélémy, Leclerc and Monjardet (1984) and the unpublished research report Critchley and Van Cutsem (1989).
Some recent results of Frank Critchley and Bernard Van Cutsem (1991) give new ways of representing dissimilarities defined on a general set and having values in an ordered set. These representations generalise and unify the well known representations of the classical bijections concerning ultrametrics on a finite set obtained by Benzécri (1965), Hartigan (1967), Jardine, Jardine and Sibson (1967) and Johnson (1967), their extensions to bijections between dissimilarities on a finite set and Numerical Stratified Clusterings discussed by Jardine and Sibson (1971), Janowitz (1978), Barthélémy, Leclerc and Monjardet (1984) and Critchley and Van Cutsem in a unpublished research report (1989). We first present our unifying and general result using equivalent definition of a function by its level map. Then we discuss in each case, how the classical results on dendrograms, prefilters and hierarchies can be obtained. We conclude by considering various possibilities opened by these results and looking at, for an example, the definition of set-valued ultrametrics.
HAL is a multi-disciplinary open access archive for the deposit and dissemination of scientific research documents, whether they are published or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et à la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Eléments aléatoires à valeurs convexes compactes Bernard Van Cutsem