The well-known Bermond-Thomassen conjecture states that every digraph of minimum out-degree at least 2k-1 contains k vertex-disjoint directed cycles. Despite being posed in 1981, this conjecture remains unresolved for all k ≥ 4. We prove a relaxation of this conjecture: every digraph D of minimum out-degree at least 2k-1 contains k vertex-disjoint cycles, each of which either is directed or can be made directed by reversing one of its arcs. This bound is sharp and answers a question raised by Cames van Batenburg during the online workshop "Entropy Compression and Related Methods" in 2021.
The genome galaxy identified in bacteria is studied by expressing the reading frame retrieval (RFR) function according to the YZ-content (GC-, AG- and GT-content) of bacterial codons. We have developed a simple probabilistic model for ambiguous sequences in order to show that the RFR function is a measure of the gene reading frame retrieval. Indeed, the RFR function increases with the ratio of ambiguous sequences and the ratio of ambiguous sequences decreases when the codon usage dispersion increases. The classical GC-content is the best parameter for characterizing the upper arm, which is related to bacterial genes with a low GC-content, and the lower arm, which is related to bacterial genes with a high GC-content. The galaxy center has a GC-content around 0.5. Then, these results are confirmed by expressing the GC-content of bacterial codons as a function of the codon usage dispersion. Finally, the bacterial genome galaxy is better described with the GC3-content in the 3rd codon site compared to the GC1-content and GC2-content in the 1st and 2nd codons sites, respectively. Whereas the codon usage is used extensively by biologists, its dispersion, which is an important parameter to reveal this genome galaxy, is surprisingly little known and unused. Therefore, we have developed a mathematical theory of codon usage dispersion by deriving several formulæ. It shows three important parameters in codon usage: the minimum and maximum codon probabilities and the number of codons with high frequency, i.e. with a probability at least 1/64. By applying this theory to the evolution of the genetic code, we see that bacteria have optimised the number of codons with high frequency to maximise the codon dispersion, thus maximising the capacity to retrieve the reading frame in genes. The derived formulæ of dispersion can be easily extended to any weighted code over a finite alphabet.
We study algebraic properties of the Tutte polynomial of a matroid and its generalizations to other combinatorially defined bivariate polynomial invariants. Merino, de Mier and Noy showed that the Tutte polynomial of a connected matroid is irreducible, and Bohn, Cameron and M{ü}ller conjectured the stronger property that the Galois/monodromy group of the Tutte polynomial of a connected matroid of rank r is isomorphic to the full symmetric group on r letters. First, we generalize the result of Merino-de Mier-Noy to the context of general ranked sets by exploiting a recent translation of the Brylawski relations, satisfied by the coefficients of the Tutte polynomial, into a functional identity. Second, we give the first confirmation of the conjecture of Bohn-Cameron-M{ü}ller for infinite families of connected matroids, including the cycle graphs and the uniform matroids. Moreover, we apply the large sieve to obtain a probabilistic statement showing that suitable linear combinations of coprime Tutte polynomials generically satisfy the conjecture.
Extending an earlier work by Kostochka for subcubic graphs, we show that a connected graph $G$ with minimum degree $2$ and maximum degree $4$ has at least $75^{n_4/5+n_3/10+1/5}$ spanning trees, where $n_i$ is the number of vertices of degree $i$ in $G$, unless $G$ is the complete graph on $5$ vertices or obtained from the complete graph on $6$ vertices by deleting the edges of a perfect matching. This, in particular, allows us to determine the value of the inferior limit of the normalised number of spanning trees (introduced by Alon) over the class of connected $4$-regular graphs to be $75^{1/5}$.
We consider a classic rendezvous game in which two players try to meet each other on a set of n locations. In each round, every player visits one of the locations, and the game finishes when the players meet at the same location. The goal is to devise strategies for both players that minimize the expected waiting time till the rendezvous. In the asymmetric case, when the strategies of the players may differ, it is known that the optimum expected waiting time of [Formula: see text] is achieved by the wait-for-mommy pair of strategies, in which one of the players stays at one location for n rounds, while the other player searches through all the n locations in a random order. However, if we insist that the players are symmetric—they are expected to follow the same strategy—then the best known strategy, proposed by Anderson and Weber [Anderson EJ, Weber RR (1990) The rendezvous problem on discrete locations. J. Appl. Probab. 27(4):839–851], achieves an asymptotic expected waiting time of [Formula: see text]. We show that the symmetry requirement indeed implies that the expected waiting time needs to be asymptotically larger than in the asymmetric case. Precisely, we prove that for every [Formula: see text], if the players need to employ the same strategy, then the expected waiting time is at least [Formula: see text], where [Formula: see text]. We propose in addition a different proof for one our key lemmas, which relies on a result by Ahlswede and Katona [Ahlswede R, Katona GOH (1978) Graphs with maximal number of adjacent pairs of edges. Acta Mathematica Academiae Scientiarum Hungaricae 32(1–2):97–120]: the argument is slightly shorter and provides a constant larger than [Formula: see text], namely, [Formula: see text]. However, it requires that n be at least 16. Both approaches seem conceptually interesting to us. Funding: This work is a part of project TOTAL (Mi. Pilipczuk) that has received funding from the European Research Council under the European Union’s Horizon 2020 research and innovation program [Grant 677651]. It is also partially funded by French ANR projects [Grant ANR-16-CE40-0023] (DESCARTES) and [Grant ANR-17-CE40-0015] (DISTANCIA).
Based on the circular code theory, we define a new function f that quantifies the property of reading frame retrieval (RFR) of genes from their codon usage. This RFR function f is computed on a massive scale in genes of genomes of bacteria, eukaryotes and archaea. By expressing f as a function of the mean number n of codons per gene, a “universal” property is identified, whatever the kingdom: the reading frame retrieval is enhanced in large genes. By investigating this property according to the theory developed, a Spearman’s rank correlation with a strong negative coefficient is observed between the codon usage dispersion d (from the uniform codon distribution 1/64 ) and the RFR function f , whatever the kingdom ( p -values <10^-180 in bacteria, <10^-61 in eukaryotes and <10^-159 in archaea). Thus, the reading frame retrieval is enhanced with the codon usage dispersion. Furthermore, this approach identifies a “genome centre” from which emerge two distinct “genome arms”: an upper arm and a lower arm, respectively, above and below the linear regression. The RFR function by itself or combined with classical methods (alignment, phylogeny) could also be a new approach to classify the genomes in the future.
Comma-free codes have been widely studied in the last sixty years, from points of view as diverse as biology, information theory and combina-torics. We develop new methods to study comma-free codes achieving the maximum size, given the cardinality of the alphabet and the length of the words. Specifically, we are interested in counting the number of such codes when all words have length 2, or 3. We first explain how different properties combine to obtain a closed-formula. We next develop an approach to tackle well-known sub-families of comma-free codes, such as self-complementary and (generalisations of) non-overlapping codes, for which the aforementioned properties do not hold anymore. We also study codes that are not contained in strictly larger ones. For instance, we de-termine the maximal size of self-complementary comma-free codes (over an alphabet of arbitrary cardinality) and the number of codes reaching the bound. We also provide a characterisation of non-overlapping trilet-ter codes that are inclusion-wise maximal, which allows us to devise the number of such codes. We point out other applications of the method, notably to self-complementary codes, including the recently introduced mixed codes. Our approaches mix combinatorial and graph-theoretical arguments.
A code X is (>= k)-circular if every concatenation of words from X that admits, when read on a circle, more than one partition into words from X, must contain at least k + 1 words. In other words, the reading frame retrieval is guaranteed for any concatenation of up to k words from X. A code that is (>= k) -circular for all integers k is said to be circular. Any code is (>= 0)-circular and it turns out that a code of trinucleotides is circular as soon as it is (>= 4)-circular. A code is k-circular if it is (>= k)-circular and not (>= k + 1)-circular. The theoretical aspects of trinucleotide k-circular codes have been developed in a companion article (Michel et al., 2022).& nbsp;Trinucleotide circular codes always retrieve the reading frame, leaving no ambiguous sequences. On the contrary, trinucleotide k-circular codes, for k is an element of {0,1, 2, 3} all have ambiguous sequences, for which the reading frame cannot always be retrieved. However, such a trinucleotide k-circular code is still able to retrieve the reading frame for a number of sequences, thereby exhibiting a partial circularity property. We describe this combinatorial property for each class of trinucleotide k-circular codes with k is an element of {0,1,2,3}. The circularity, i.e. the reading frame retrieval, is an ordinary property in genes. In order to consider the different cases of ambiguous sequences, we derive a new and general formula to measure the reading frame loss, whatever the trinucleotide k-circular code. This formula allows us to study the evolution of any trinucleotide k-circular code of (maximal) cardinality 20 to the genetic code, based on the reading frame retrieval property. We apply this approach to analyse the evolution of the trinucleotide circular code X observed in genes to the genetic code.& nbsp;The (>= 1)-circular codes of maximal size 20 necessarily have the same number of each nucleotide, specifically 15 = 3.20/4. This balanceness property can also be achieved by trinucleotide codes of cardinality 4, 8,12 and 16. We call such trinucleotide codes balanced. We develop a general mathematical method to compute the number of balanced trinucleotide codes of each size, which also applies to self-complementary trinucleotide codes. We establish and quantify a relation between this balanceness property and the self-complementarity property.& nbsp;The combinatorial hierarchy of trinucleotide k-circular codes is updated with the growth function results. The numbers of amino acids coded by the trinucleotide k-circular codes are given for the cases maximal, minimal, self-complementary k-, (k, k, k)- and self-complementary (k, k, k)-circular.
A code X is (⩾k)-circular if every concatenation of words from X that admits, when read on a circle, more than one partition into words from X, must contain at least k+1 words. In other words, the reading frame retrieval is guaranteed for any concatenation of up to k words from X. A code that is (⩾k)-circular for all integers k is said to be circular. Any code is (⩾0)-circular and it turns out that a code of trinucleotides is circular as soon as it is (⩾4)-circular. A code is k-circular if it is (⩾k)-circular and not (⩾k+1)-circular. Due to the explosive combinatorics of trinucleotide k-circular codes, we developed three classes of algorithms based on: (i) the smallest directed cycles (directed girth) in graphs; (ii) the eigenvalues of matrices; and (iii) the files that incrementally save partial results. These different approaches also allow us to verify the computational results obtained. We determine here the growth functions of trinucleotide k-circular codes, k varying between 0 and 4, in the general case and in various particular cases: minimum, minimal, maximum, self-complementary k-, (k,k,k)- and self-complementary (k,k,k)-circular.
In a recent article, Bogdanowicz determines the minimum number of spanning trees a connected cubic multigraph on a fixed number of vertices can have and identifies the unique graph that attains this minimum value. He conjectures that a generalized form of this construction, which we here call a padded paddle graph, would be extremal for d-regular multigraphs where $d\geq 5$ is odd. We prove that, indeed, the padded paddle minimises the number of spanning trees, but this is true only when the number of vertices, $n$, is greater than $(9d+6)/8$. We show that a different graph, which we here call the padded cycle, is optimal for $n<(9d+6)/8$ . This fully determines the $d$-regular multi-graphs minimising the number of spanning trees for odd values of $d$. We employ the approach we develop to also consider and completely solve the even degree case. Here, the parity of $n$ plays a major role and we show that, apart from a handful of irregular cases when both $d$ and $n$ are small, the unique extremal graphs are padded cycles when $n$ is even and a different family, which we call fish graphs, when $n$ is odd.
We consider a natural, yet seemingly not much studied, extremal problem in bipartite graphs. A bi-hole of size $t$ in a bipartite graph $G$ is a copy of $K_{t, t}$ in the bipartite complement of $G$. Let $f(n, \Delta)$ be the largest $k$ for which every $n \times n$ bipartite graph with maximum degree $\Delta$ in one of the parts has a bi-hole of size $k$. Determining $f(n, \Delta)$ is thus the bipartite analogue of finding the largest independent set in graphs with a given number of vertices and bounded maximum degree. Our main result determines the asymptotic behavior of $f(n, \Delta)$. More precisely, we show that for large but fixed $\Delta$ and $n$ sufficiently large, $f(n, \Delta) = \Theta(\frac{\log \Delta}{\Delta} n)$. We further address more specific regimes of $\Delta$, especially when $\Delta$ is a small fixed constant. In particular, we determine $f(n, 2)$ exactly and obtain bounds for $f(n, 3)$, though determining the precise value of $f(n, 3)$ is still open.
We introduce a new method for computing bounds on the independence number and fractional chromatic number of classes of graphs with local constraints, and apply this method in various scenarios. We establish a formula that generates a general upper bound for the fractional chromatic number of triangle-free graphs of maximum degree~$\Delta \ge 3$. This upper bound matches that deduced from the fractional version of Reed's bound for small values of~$\Delta$, and improves it when~$\Delta\ge 17$, transitioning smoothly to the best possible asymptotic regime, barring a breakthrough in Ramsey theory. Focusing on smaller values of~$\Delta$, we also demonstrate that every graph of girth at least~$7$ and maximum degree~$\Delta$ has fractional chromatic number at most~$1+ \min_{k \in \mathbb{N}} \frac{2\Delta + 2^{k-3}}{k}$. In particular, the fractional chromatic number of a graph of girth~$7$ and maximum degree~$\Delta$ is at most~$\frac{2\Delta+9}{5}$ when~$\Delta \in [3,8]$, at most~$\frac{\Delta+7}{3}$ when~$\Delta \in [8,20]$, at most~$\frac{2\Delta+23}{7}$ when~$\Delta \in [20,48]$, and at most~$\frac{\Delta}{4}+5$ when~$\Delta \in [48,112]$. In addition, we also obtain new lower bounds on the independence ratio of graphs of maximum degree~$\Delta \in \{3,4,5\}$ and girth~$g\in \{6,\dotsc,12\}$, notably~$1/3$ when~$(\Delta,g)=(4,10)$ and~$2/7$ when~$(\Delta,g)=(5,8)$.
We develop an algorithmic framework for graph colouring that reduces the problem to verifying a local probabilistic property of the independent sets. With this we give, for any fixed $k\ge 3$ and $\varepsilon>0$, a randomised polynomial-time algorithm for colouring graphs of maximum degree $\Delta$ in which each vertex is contained in at most $t$ copies of a cycle of length $k$, where $1/2\le t\le \Delta^\frac{2\varepsilon}{1+2\varepsilon}/(\log\Delta)^2$, with $\lfloor(1+\varepsilon)\Delta/\log(\Delta/\sqrt t)\rfloor$ colours. This generalises and improves upon several notable results including those of Kim (1995) and Alon, Krivelevich and Sudakov (1999), and more recent ones of Molloy (2019) and Achlioptas, Iliopoulos and Sinclair (2019). This bound on the chromatic number is tight up to an asymptotic factor $2$ and it coincides with a famous algorithmic barrier to colouring random graphs.
We consider, for every positive integer $a$, probability distributions on subsets of vertices of a graph with the property that every vertex belongs to the random set sampled from this distribution with probability at most $1/a$. Among other results, we prove that for every positive integer~$a$ and every planar graph $G$, there exists such a probability distribution with the additional property that deleting the random set creates a graph with component-size at most $(\Delta(G)-1)^{a+O(\sqrt{a})}$, or a graph with treedepth at most $O(a^3\log_2(a))$. We also provide nearly-matching lower bounds.
We give a short proof of the following theorem due to Jon H. Folkman (1969): The chromatic number of any graph is at most 2 plus the maximum over all subgraphs of the difference between the number of vertices and twice the independence number.
A code X is k-circular if any concatenation of at most k words from X, when read on a circle, admits exactly one partition into words from X. It is circular if it is k-circular for every integer k. While it is not a priori clear from the definition, there exists, for every pair $$(n,\ell )$$, an integer k such that every k-circular $$\ell $$-letter code over an alphabet of cardinality n is circular, and we determine the least such integer k for all values of n and $$\ell $$. The k-circular codes may represent an important evolutionary step between the circular codes, such as the comma-free codes, and the genetic code.
The Petersen colouring conjecture states that every bridgeless cubic graph admits an edge-colouring with 5 colours such that for every edge e, the set of colours assigned to the edges adjacent to e has cardinality either 2 or 4, but not 3. We prove that every bridgeless cubic graph G admits an edge-colouring with 4 colours such that at most 8/15 . vertical bar E(G)vertical bar edges do not satisfy the above condition. This bound is tight and the Petersen graph is the only connected graph for which the bound cannot be decreased. We obtain such a 4-edge-colouring by using a carefully chosen subset of edges of a perfect matching, and the analysis relies on a simple discharging procedure with essentially no reductions and very few rules.
The first author together with Jenssen, Perkins and Roberts (2017) recently showed how local properties of the hard-core model on triangle-free graphs guarantee the existence of large independent sets, of size matching the best-known asymptotics due to Shearer (1983). The present work strengthens this in two ways: first, by guaranteeing stronger graph structure in terms of colourings through applications of the Lov\'asz local lemma; and second, by extending beyond triangle-free graphs in terms of local sparsity, treating for example graphs of bounded local edge density, of bounded local Hall ratio, and of bounded clique number. This generalises and improves upon much other earlier work, including that of Shearer (1995), Alon (1996) and Alon, Krivelevich and Sudakov (1999), and more recent results of Molloy (2019), Bernshteyn (2019) and Achlioptas, Iliopoulos and Sinclair (2019). Our results derive from a common framework built around the hard-core model. It pivots on a property we call local occupancy, giving a clean separation between the methods for deriving graph structure with probabilistic information and verifying the requisite probabilistic information itself.
In 1980, Erdős, Rubin and Taylor asked whether for all positive integers a, b, and m, every (a:b)-choosable graph is also (am:bm)-choosable. We provide a negative answer by exhibiting a 4-choosable graph that is not (8:2)-choosable.
This paper contributes to a programme initiated by the first author: "How much information about a graph is revealed in its Potts partition function?" We show that the W-polynomial distinguishes non-isomorphic weighted trees of a good family. The framework developed to do so also allows us to show that the W-polynomial distinguishes non-isomorphic caterpillars. This establishes Stanley's isomorphism conjecture for caterpillars, an extensively studied problem.
Petr Kolman合作论文数Charles University;Department of Applied Mathematics2
Jerrold R. Griggs合作论文数Department of Mathematics
University of South Carolina2