
Existential width and universal width respectively quantify the amount of nondeterminism and parallelism present in an alternating finite automaton (AFA). These measures can be defined as either worst case measures, known as maximal existential and universal widths, or as the best case measures, referred to as optimal existential and universal widths. We investigate upper bounds on the finite optimal widths of AFAs and construct several examples of AFAs that attain finite optimal widths strictly larger than any achievable finite maximal width. We establish an upper bound on the finite optimal universal width of unary universal finite automata (UFAs) and derive a partial upper bound on the finite optimal existential width of unary nondeterministic finite automata (NFAs).
We consider shuffle ideals and their finite bases, both represented by regular expressions. We introduce an operation that, given a regular expression α∈_fin representing a finite language, returns an expression for the shuffle ideal . This operation induces a new combinatorial class, which we call _SI . For both classes, we define the Antimirov automaton and provide results on their state complexity in the worst as well as the average case. For α∈_fin , in the worst case, the number of states is at most |α |_ Σ , the number of letters in α . The asymptotic average is, however, bounded by |α |_ Σ /2 . For τ∈_SI , in the worst case, the number of states is at most |τ |_ Σ ^⋆ , the number of occurrences of Σ ^⋆ in τ . On the other hand, the asymptotic average number of states is bounded by |τ |_ Σ ^⋆/2 .
We investigate the state complexity of the boundary operation on the classes of closed and ideal languages. We establish a tight upper bound of (n+2)2^n-3+2 for prefix-closed languages, n(n-1)/2+2 for suffix-closed languages, and n+2 for factor- and subword-closed languages. To describe witnesses, we use a binary alphabet for prefix-closed languages, and a ternary alphabet for the remaining classes. In the binary case, we derive a lower bound of n(n-1)/2 for suffix-closed languages, and a tight upper bound of n+1 for factor- and subword-closed languages. The state complexity of the boundary of unary closed languages is n. Since the boundary of a language coincides with that of its complement, and the complement of a closed language is an ideal language, all our results for closed languages extend to the classes of right, left, two-sided, and all-sided ideals.
Guo et al. proved that the number of palindromes occurring in the set of all conjugates of a word is bounded above by two. In this paper, we show that for an antimorphic involution θ , the number of θ -palindromes occurring in the set of all θ -conjugates of a word is bounded above by six.
Context-conditional grammars have been introduced by Masopust and Meduna about 20 years ago, when also first few descriptional complexity results have been obtained. These grammars generalize both semi-conditional and generalized forbidding grammars. We extend these studies considerably and derive some descriptional complexity results which are new characterizations of the class of recursively enumerable languages by parsimonious resources. As context-conditional grammars offer a broad variety of natural descriptional complexity parameters, we will focus on the number of nonterminals, the degree, the index, and the number of conditional rules, which can be refined by additionally counting the non-simple rules and the non-semi-conditional rules. Most parameters are known from the literature studying descriptional complexity aspects of semi-conditional grammars, apart from the index which is the pair of numbers formed by the maximum cardinality of any permitting context set and the maximum cardinality of any forbidden context set. We also introduce a novel way of illustrating the proofs that eases the reader to keep track of the various case distinctions.
In descriptional complexity, one often tries to achieve a goal when limiting the available resources. In the context of semi-conditional grammars, these could be the degree (the length of permitting and forbidden contexts) or the number of nonterminals, to name two prominent examples. Concerning the degree (and similar parameters), we now propose to quantify the “distance from simplicity” as a secondary parameter, counting the number of rules that actually make use of the admissible length of permitting or forbidden contexts. This way, we quantify the distance to the next level of simplicity of this grammar type. To better understand this approach, we apply it to a grammar formalism that allows for a richer parameterization by the distance from simplicity principle: semi-conditional matrix grammars. In passing, we will also improve several results known for the descriptional complexity of such grammars.
Multi-entry finite automata (MDFAs) are a generalization of DFAs that allow for an arbitrary number of initial states. This paper extends existing research on MDFAs by studying their operational state complexity and the computational complexity of their associated decision problems. In particular, we analyze the cost on the number of states of the standard language-theoretic operations when performed on MDFAs and compare these results with the well known complexities for DFAs and NFAs. Additionally, we also analyze the complexity of deciding the membership, emptiness, universality and inclusion problems for MDFAs. Our findings contribute to a deeper understanding of the role that nondeterminism plays in the computational difficulty of certain problems.
We continue the study of inductive inference of one- and two-way cellular automata ( ), where the goal is to infer a that is compatible with a finite amount of available data. Here, we consider this data to be in the form of a finite set of intervals, where each interval consists of two words w and w' over a state set alphabet, with a positive integer i, and the output should be a that is compatible with each interval (w,w',i) , meaning that it can derive w' from w in i steps, if one exists. We study inference where the state set is provided in advance as part of the input, and so the inferred s must only use those states and cannot create arbitrarily many states. We consider two main variations of this problem, where 1) the intervals, the state set, and some subset of the transitions of the cellular automaton are inputs, and the goal is to extend it to a compatible full cellular automaton using the given state set, 2) the intervals and the state set are inputs, but otherwise the is completely unknown and must be inferred. We determine that both problems are -complete. Both are also -complete if we only provide the number of states allowed.
A k-path nondeterministic input-driven pushdown automaton (NIDPDA) has at most k computation paths on any input. We present an improved determinization construction for k-path NIDPDAs and a lower bound for the size blow-up of determinization that is significantly better than the previous lower bounds.
This paper investigates the new notion of 2-word- π -representable graphs: the nodes of the graph correspond to the letters of the two words and there exists an edge between two nodes if the projections of any two letters of both words are equal. The benefit of not only using one word for a representation as introduced by Kitaev and Pyatkin is that every graph is 2-word- π -representable. We present an algorithm that returns two representing words for any graph. Aside, we show that every permutation graph is representable by two 1-uniform words and give constructions how graph operations on 2-word- π -representable graphs can be realised on their representing words which give further insights into the representation of cographs.
This paper resolves the open larger-alphabet quotient case in the accepting-state complexity theory of permutation automata. Rauch and Holzer showed that, in the unary setting, the attainable right-quotient accepting-state complexities are exactly [1,mn] . We prove that over arbitrary alphabets the exact spectrum is g^asc_-1,PFA(m,n)= {[ {0}, if m=0 or n=0,; ℕ_>0, if m,n≥ 1. ]. Thus, once both input languages are nonempty, every positive accepting-state complexity is attainable for right quotient, and 0 is the only unavoidable magic value. The proof has two parts. First, we show that if m,n≥ 1 , then the quotient language KL^-1 cannot be empty when K and L are accepted by permutation automata with asc(K)=m and asc(L)=n ; this follows from the bijectivity of the transition action. Second, for every m,n≥ 1 and every α≥ m , we construct a ternary witness pair (A^q_m,α,B^q_n,α) such that asc(L(A^q_m,α))=m , asc(L(B^q_n,α))=n , and asc (L(A^q_m,α)L(B^q_n,α)^-1 )=α . The high-range construction is group-theoretic: the words accepted by B^q_n,α induce exactly a point stabilizer in a symmetric group, and the standard quotient construction then saturates the original final set of A^q_m,α to a full orbit, yielding a minimal quotient automaton with exactly α final states. Combined with the known unary interval [1,mn] , this yields the complete spectrum and resolves the larger-alphabet right-quotient case for permutation automata.
We determine the accepting-state spectrum of reversal for permutation automata exactly, thereby proving the Rauch–Holzer conjecture on this operation. For every m≥ 2 and every α≥ 2 , we construct a binary permutation automaton A_m,α such that asc(L(A_m,α))=m and asc(L(A_m,α)^R)=α . Combined with the trivial cases m=0 and m=1 , and with the previously known fact that 1 is magic for every m≥ 2 , this yields the exact spectrum g^asc_R,PFA(m)= {[ {0}, if m=0,; {1}, if m=1,; ℕ_≥ 2, if m≥ 2. ]. Thus reversal has, for permutation automata, the simplest possible exact accepting-state spectrum compatible with the single nontrivial obstruction at value 1 . The proof uses a uniform group-theoretic witness family: the states of the forward automaton are the α -subsets of [n] , where n=m+α -1 , under the action generated by an n -cycle and a transposition, while the accepting states form a single star family. After reversal, the reachable subset-states are exactly the stars, which makes it possible to count the accepting reachable states precisely and to prove minimality of the reachable reverse automaton.
The Fibonacci infinite word f = (f_i)_i ≥ 0 = 01001010⋯ is one of the most celebrated objects in combinatorics on words. There is a simple 5-state automaton that, given i in lsd-first Zeckendorf representation, computes its i'th term f_i, and a 2-state automaton for msd-first. In this paper we consider the state complexity of the automaton generating the shifted sequence (f_i+c)_i ≥ 0, and show that it is O(log c) for both msd-first and lsd-first input. This is close to the information-theoretic minimum for an aperiodic sequence. The techniques involve a mixture of state complexity techniques and Diophantine approximation.
We study asymptotic densities and relative densities of formal languages. We show that the class of regular languages is, to some extent, the only class of formal languages for which those limit probabilities can be effectively computed.
In this paper we show a series of results related to the decidability and expressive power of quantifier-free first order logical formulas containing language membership predicates (for the classes of regular and visibly push-down languages), string concatenation, letter-count and string-length functions, as well as linear integer arithmetic.
Matrix grammars are one of the first approaches ever proposed in regulated rewriting, prescribing that rules have to be applied in a certain order. Semi-conditional grammars introduced the notion of permitting and forbidding contexts in context-free rules. In this paper, we introduce a new form of matrix grammars, called matrix forbidding grammars, where matrices of context-free rules are considered and each context-free rule is associated with a context so that a rule (in a matrix) can be applied only if this context is not a subword of the current sentential form. For matrix forbidding grammars, we study a range of descriptional complexity parameters, such as degree (length of forbidden contexts), number of nonterminals, number of conditional rules, number of matrices containing conditional rules, and matrix length, in order to explore the Pareto frontier (set of solutions that represents the best trade-off between all the parameters) of computational completeness.
Finite automata with translucent input letters can be obtained by equipping classical finite automata with a translucency function which, depending on the current state, establishes the set of invisible input symbols: such symbols are skipped in the current move and dealt with subsequently, while the first visible symbol from the current input head position is processed and consumed. These models can be returning, meaning that the input head is rewinded at each move after processing a visible symbol, or not. In their original definition by Mráz and Otto [13], translucent finite automata feature a one-way input head, i.e., the head always “looks to the right” to spot the next visible symbol. Here, we extend this definition in the deterministic case by allowing a two-way input head motion, meaning that the next visible symbol can be found by looking either to the left or to the right of the head. We also consider translucent sweeping finite automata, as the restricted version of two-way devices reverting input head direction at the endmarkers only. We prove that the sweeping restriction does not affect the recognition power in the returning mode, while the opposite holds for non-returning devices. We show that our devices are strictly more powerful than their one-way version, and that they are incomparable with translucent pushdown automata [9, 14, 17]. We investigate some closure properties for the families of languages recognized by translucent two-way and sweeping finite automata. In particular, while dealing with this latter topic, an open problem stated in [13] is addressed and solved.
We propose two parameters that should be incorporated into the measure for the descriptional complexity of deterministic and nondeterministic finite automata with translucent words: the maximal length of a translucent word and the maximal cardinality of a set of translucent words. We illustrate the influence of these two parameters on the expressive capacity of finite automata with translucent words, where we concentrate, in particular, on automata over a binary alphabet. We establish that the restrictions for the length of the longest word in any set of translucent words and the restriction for the cardinality of the admitted sets of translucent words yield a two-dimensional infinite hierarchy of classes of binary languages.
We explore small balanced vertex separators in NFA to regular expression conversion. We show that any class of finite automata whose underlying class of graphs admits "strongly sublinear separators" also admits shorter regular expressions. We propose a new state elimination ordering heuristic, called the "gate score" heuristic, which is a variant of the weight heuristic of Delgado and Morais. Empirically, we find that both the weight heuristic and the gate score heuristic perform better than directly exploiting vertex separators despite having no provable performance guarantee. We also provide some theoretical support for the weight heuristic of Delgado and Morais in finite automata with dense transition structures.
Considering regular expressions extended with synchronised shuffle on backbones, we present two equivalent automata: the location based position automaton and the partial derivative automaton. We show that the latter is a quotient of the former. Using the framework of analytic combinatorics, we study the average complexity of the partial derivative automaton. Surprisingly, for binary and ternary alphabets the average number of partial derivatives by a symbol is exponential on the size of the expression, while it is constant for larger alphabets which is what happens with the results for all other regular operators studied so far. Furthermore, we prove that the average number of states of the partial derivative automaton is bounded from above by (1.57708 + o(1))(m), while in the worst-case that value is O(3(m)), where m is the alphabetic size of the expression.