
In this article we study a numeration system previously used to prove combinatorial properties in discrete geometry: the Δ -numeration. Since this system, introduced via the fully subtractive algorithm, has been seen mainly as a tool, we propose here to study it from the point of view of numeration systems. In particular, we make the link with β -numeration and Cantor real bases. We reintroduce the rewriting system introduced to calculate in Δ -numeration. This systems is based on the properties of the fully subtractive algorithm and is normalising. Finally, we study the ultimately periodic case, a special case of alternate bases, and show that the ultimately periodic words represent exactly the elements of ℚ[β ] where β is the inverse of a Pisot number.
The Heinis spectrum is the set of all pairs (α _u,β _u) such that α _u=lim inf _n→∞p_u(n)/n and β _u=lim sup _n→∞p_u(n)/n for some infinite word u. In this paper, we demonstrate that there exists a closed connected set with non-empty interior contained in . Furthermore, every point in this set can be represented as the pair (α _u,β _u) for some recurrent word u. The construction is explicit, algorithmic in nature and is based on constructing certain “Cantor sets of integers”, whose “gaps” correspond to blocks of zeros.
In this paper we introduce the notion of generalized factorization for an arbitrary submonoid M⊆ A^* , where A^* is the free monoid generated by an alphabet A, generalizing, in this way, the notion of factorization of A^* . Then we give a characterization of the free product of two submonoids of A^* in terms of unambiguous products of monoids. To do this we make use of the notion of coding partition of a set X⊆ A^+ , where A^+ is the free semigroup generated by an alphabet A. Moreover, given a coding partition of a set X⊆ A^+ , we will show how to construct a generalized factorization of X^* .
Christol’s theorem [7] characterizes algebraic formal power series over finite fields in terms of automatic sequences, establishing a fundamental link between algebraicity and computability. A refinement of this result by Bell and the author [2] provides an algebraic characterization of formal power series whose support is a sparse automatic set, showing that sparseness can be characterized in terms of certain key transformations such as the Frobenius map, multiplicative scaling, and power transformations. In this paper, we extend these results to the multivariate setting, building on Salon’s [16] generalization of Christol’s theorem. By using a characterization of sparse regular languages, we look at algebraic multivariate formal power series whose supports are sparse subsets of ℕ^d and characterize them in terms of above-mentioned transformations, providing a structural understanding of the interplay between algebraicity and sparseness in multiple variables.
Numeration systems are maps between a set of numbers and a set of words that act as representations of these numbers. One desirable property is positionality: the ability to relate positions in the words to values of the numbers. In general, positionality is hard to decide. In this article, we obtain a criterion to decide the positionality of so-called Dumont–Thomas numeration systems, arising from substitutions. Then, we particularize this criterion to some well-behaved classes of substitutions, allowing us to link the related systems to existing literature.
The so-called MP-ratio is a kind of measure of how “packed with palindromes” a given word is. The lower bound on the MP-ratio for the set of all n-ary words is (trivially) 1, while the best possible upper bound is an open problem in the general case. It is solved for n=2 (where the optimal upper bound is 4) and for n=3 (where the optimal upper bound is 6). Also, it is known that in the n-ary case the optimal bound is between 2n and the order of the growth n2^n/2 . In this article we solve this problem for quaternary words, for which we show that the best possible upper bound on the MP-ratio equals 8. We believe that this is the last case in which the result is 2n, that is, we believe that for n⩾ 5 there are words whose MP-ratio is strictly larger than 2n.
Motivated by Parikh matrices of picture arrays introduced in combinatorial image analysis, we propose a generalization of binomial coefficients of words to multidimensional arrays. These coefficients recursively count prescribed patterns occurring in an array. The base case is the one of binomial coefficients of words. With our definition we extend Pascal’s rule, the Chu–Vandermonde identity and therefore, the concept of Parikh matrices, in a natural way. We further present some more binomial-related identities and introduce (q, t)-deformations, i.e., multivariate polynomials whose evaluation at (q,t)=(1,1) recovers the value of the classical coefficients. We explain the additional combinatorial information encoded in the coefficients of these (q, t)-polynomials compared to their integer-valued counterparts.
In this paper we introduce the notion of free product of formal series. Using this notion, we can characterize the free product of submonoids of A^* , where A^* is the free monoid generated by an alphabet A. Moreover given a set X⊆ A^+ , where A^+ is the free semigroup generated by an alphabet A, we can characterize a partition of X by the free product of the formal series associated to the classes of the partition.
A finite word w is called closed if it has length at most 1 or it contains a proper factor that occurs both as a prefix and as a suffix but does not have internal occurrences. An infinite word u is called closed-rich if the infimum of all possible ratios between the number of closed factors within any factor w of u and square of the length of w exists and is positive. We define this infimum as the closed-rich constant C_u of the infinite closed-rich word u. Puzynina and Parshina (2024) proved that infinite closed-rich words exist. In this paper, we estimate possible values of C_u for an infinite closed-rich word u, and apply these results to estimate the supremum C_sup of the closed-rich constants of infinite closed-rich words. We show that 0.0952 < C_sup≤ 0.165964 , where the lower bound comes from the Fibonacci word.
We state a conjecture on the repetition threshold of rich sequences over alphabet of any size. It is known to hold for binary and ternary alphabets. We provide two main contributions that may be helpful for the proof on larger alphabets. First we show that the ternary rich sequence with minimum critical exponent is a morphic image of a fixed point, i.e., an HD0L sequence. Second we draw attention to the fact that the rich sequences having the minimum critical exponent show a large degree of symmetry, i.e., they are G-rich with respect to a group G generated by more than one antimorphism. The notion of G-richness generalizes the notion of richness in palindromes which is based on one antimorphism, namely the reversal mapping.
An infinite sequence with finitely many distinct letters is said to have the uniform distribution property if all letters in the alphabet of the sequence have the same density in all arithmetical progressions. This property was first studied by Gelfond for sum of digits functions. In this note, we characterize a larger class of purely automatic sequences with the same property.
We say that a word w contains a half-flip if it contains nonempty factors u and vu where |u| = |v|. Fici reports a non-constructive proof of the existence of an infinite word over a finite alphabet avoiding half-flips and asks the size of the smallest alphabet over which half-flips may be avoided. Half-flips are unavoidable over a 4-letter alphabet. Over an 8-letter alphabet we conjecture a 3-uniform D0L avoiding half-flips; over a 5-letter alphabet we conjecture a messier HD0L construction.
Let d be an integer between 0 and 4, and W be a 2-dimensional word of dimensions h x w on the binary alphabet 0, 1, where h, w in Z > 0. Assume that each occurrence of the letter 1 in W is adjacent to at most d letters 1. We provide an exact formula for the maximum number of letters 1 that can occur in W for fixed (h, w). As a byproduct, we deduce an upper bound on the length of maximum snake polyominoes contained in a h x w rectangle.
A word over an ordered alphabet is said to be clustering if identical letters appear adjacently in its Burrows-Wheeler transform. Such words are strictly related to (discrete) interval exchange transformations. We use an extended version of the well-known Rauzy induction to show that every return word in the language generated by a regular interval exchange transformation is clustering, partially answering a question of Lapointe (2021).
We study circularity in DF0L systems, a generalization of D0L systems. We focus on two different types of circularity, called weak and strong circularity. When the morphism is injective on the language of the system, the two notions are equivalent, but they may differ otherwise. Our main result shows that failure of weak circularity implies unbounded repetitiveness, and that unbounded repetitiveness implies failure of strong circularity. This extends previous work by the second and third authors for injective systems. To help motivate this work, we also give examples of non-injective but strongly circular systems.
A shuffle square is a word consisting of two shuffled copies of the same word. For instance, the French word is a shuffle square, as it can be split into two copies of the word . An ordered graph is a graph with a fixed linear order of vertices. We propose a representation of shuffle squares in terms of special nest-free ordered graphs and demonstrate the usefulness of this approach by applying it to several problems. Among others, we prove that binary words of the type ()^n , n odd, are not shuffle squares and, moreover, they are the only such words among all binary words whose every -run has length one or two, while every -run has length two. We also provide a counterexample to a believable stipulation that binary words of the form 1^n0^n-21^n-4⋯ , n odd, are far from being shuffle squares (the distance measured by the minimum number of letters one has to delete in order to turn a word into a shuffle square).
An upward (resp. downward) digitally convex word is a binary word that best approximates from below (resp. from above) an upward (resp. downward) convex curve in the plane. We study these words from the combinatorial point of view, formalizing their geometric properties and highlighting connections with Christoffel words and finite Sturmian words. In particular, we study from the combinatorial perspective the operations of inflation and deflation on digitally convex words.
This paper concerns the avoidability of abelian and additive powers in infinite rich words. In particular, we construct an infinite additive 5-power-free rich word over {0,1} and an infinite additive 4-power-free rich word over {0, 1, 2}. The alphabet sizes are as small as possible in both cases, even for abelian powers.
Abstract numeration systems encode natural numbers using radix ordered words of an infinite regular language and linear recurrence sequences play a key role in their valuation. Sequence automata, which are deterministic finite automata with an additional linear recurrence sequence on each transition, are introduced to compute various ℤ -rational non commutative formal series in abstract numeration systems. Under certain Pisot conditions on the recurrence sequences, the support of these series is regular. This property can be leveraged to derive various synchronized relations including a deterministic finite automaton that computes the addition relation of various Dumont-Thomas numeration systems and deterministic finite automata converting between various numeration systems. A practical implementation for Walnut is provided.
Let $c>1$ be a real constant. We say that a language $L$ is $c$-\emph{constantly growing} if for every word $u\in L$ there is a word $v\in L$ with $\vert u\vert<\vert v\vert\leq c+\vert u\vert$. We say that a language $L$ is $c$-\emph{geometrically growing} if for every word $u\in L$ there is a word $v\in L$ with $\vert u\vert<\vert v\vert\leq c\vert u\vert$. Given a language $L$, we say that $L$ is $REG$-\emph{dissectible} if there is a regular language $R$ such that $\vert L\setminus R\vert=\infty$ and $\vert L\cap R\vert=\infty$. In 2013, it was shown that every $c$-constantly growing language $L$ is $REG$-dissectible. In 2023, the following open question has been presented: "Is the family of geometrically growing languages $REG$-dissectible?" We construct a $c$-geometrically growing language $L$ that is not $REG$-dissectible. Hence we answer negatively to the open question.