Fix a positive integer $n$ and consider the bipartite graph whose vertices are the $3$-element subsets and the $2$-element subsets of $[n]=\{1,2,\dots,n\}$, and there is an edge between $A$ and $B$ if $A\subset B$. We prove that the domination number of this graph is $\binom{n}{2}-\lfloor\frac{(n+1)^2}{8}\rfloor$, we characterize the dominating sets of minimum size, and we observe that the minimum size dominating set can e chosen as an independent set. This is an exact version of an asymptotic result from [Balogh2021]. For the corresponding bipartite graph between the $(k+1)$-element subsets and the $k$-elements subsets of $[n]$ ($k\geq 3$), we provide a new construction for small independent dominating sets. This improves on a construction from [Gerbner2012], where these independent dominating sets have been studied under the name saturating flat antichains.
This is the second of two papers investigating for which positive integers m there exists a maximal antichain of size m in the Boolean lattice B_n (the power set of [n]:={1,2,… ,n} , ordered by inclusion). In the first part, the sizes of maximal antichains have been characterized. Here we provide an alternative construction with the benefit of showing that almost all sizes of maximal antichains can be obtained using antichains containing only l-sets and (l+1) -sets for some l.
Object stores offer an increasingly popular choice for data management and analytics. As with every data model, managing the integrity of objects is fundamental for data quality but also important for the efficiency of update and query operations. In response to shortcomings of unique and existence constraints in object stores, we propose a new principled class of constraints that separates uniqueness from existence dimensions of data quality, and fully supports multiple labels and composite properties. We illustrate benefits of the constraints on real-world examples of property graphs where node integrity is enforced for better update and query performance. The benefits are quantified experimentally in terms of perfectly scaling the access to data through indices that result from the constraints. We establish axiomatic and algorithmic characterizations for the underlying implication problem. In addition, we fully characterize which non-redundant families of constraints attain maximum cardinality for any given finite sets of labels and properties. We exemplify further use cases of the constraints: elicitation of business rules, identification of data quality problems, and design for data quality. Finally, we propose extensions to managing the integrity of objects in object stores such as graph databases.
Extending a classical theorem of Sperner, we characterize the integers m such that there exists a maximal antichain of size m in the Boolean lattice Bn, that is, the power set of [n]:={1,2,… ,n} , ordered by inclusion. As an important ingredient in the proof, we initiate the study of an extension of the Kruskal-Katona theorem which is of independent interest. For given positive integers t and k, we ask which integers s have the property that there exists a family ℱ of k-sets with |ℱ|=t such that the shadow of ℱ has size s, where the shadow of ℱ is the collection of (k − 1)-sets that are contained in at least one member of ℱ . We provide a complete answer for t⩽ k+1 . Moreover, we prove that the largest integer which is not the shadow size of any family of k-sets is √(2)k^3/2+√(8)k^5/4+O(k) .
This is the second in a sequence of three papers investigating the question for which positive integers $m$ there exists a maximal antichain of size $m$ in the Boolean lattice $B_n$ (the power set of $[n]:=\{1,2,\dots,n\}$, ordered by inclusion). In the previous paper we characterized those $m$ between $\binom{n}{\lceil n/2\rceil}-\lceil n/2\rceil^2$ and the maximum size $\binom{n}{\lceil n/2 \rceil}$ that are not sizes of maximal antichains. In this paper we show that all smaller $m$ are sizes of maximal antichains.
Building on classical theorems of Sperner and Kruskal-Katona, we investigate antichains $\mathcal {F}$ in the Boolean lattice Bn of all subsets of $[n]:=\{1,2,\dots ,n\}$ , where $\mathcal {F}$ is flat, meaning that it contains sets of at most two consecutive sizes, say $\mathcal {F}=\mathcal {A}\cup {\mathscr{B}}$ , where $\mathcal {A}$ contains only k-subsets, while ${\mathscr{B}}$ contains only (k − 1)-subsets. Moreover, we assume $\mathcal {A}$ consists of the first m k-subsets in squashed (colexicographic) order, while ${\mathscr{B}}$ consists of all (k − 1)-subsets not contained in the subsets in $\mathcal {A}$ . Given reals α, β > 0, we say the weight of $\mathcal {F}$ is $\alpha \cdot |\mathcal {A}|+\beta \cdot |{\mathscr{B}}|$ . We characterize the minimum weight antichains $\mathcal {F}$ for any given n,k,α,β, and we do the same when in addition $\mathcal {F}$ is a maximal antichain. We can then derive asymptotic results on both the minimum size and the minimum Lubell function.
Extending a classical theorem of Sperner, we investigate the question for which positive integers m there exists a maximal antichain of sizem in the Boolean lattice Bn, that is, the power set of [n] := {1, 2, . . . , n}, ordered by inclusion. We characterize all such integers m in the range ( n ⌈n/2⌉ ) −⌈n/2⌉2 6 m 6 ( n ⌈n/2⌉ ) . As an important ingredient in the proof, we initiate the study of an extension of the Kruskal-Katona theorem which is of independent interest. For given positive integers t and k, we ask which integers s have the property that there exists a family F of k-sets with |F| = t such that the shadow of F has size s, where the shadow of F is the collection of (k − 1)-sets that are contained in at least one member of F . We provide a complete answer for the case t 6 k + 1. Moreover, we prove that the largest integer which is not the shadow size of any family of k-sets is √ 2k3/2 + 4 √ 8k5/4 +O(k).
Possibility theory is applied to introduce and reason about the fundamental notion of a key for uncertain data. Uncertainty is modeled qualitatively by assigning to tuples of data a degree of possibility with which they occur in a relation, and assigning to keys a degree of certainty which says to which tuples the key applies. The associated implication problem is characterized axiomatically and algorithmically. Using extremal combinatorics, we then characterize the families of non-redundant possibilistic keys that attain maximum cardinality. In addition, we show how to compute for any given set of possibilistic keys a possibilistic Armstrong relation, that is, a possibilistic relation that satisfies every key in the given set and violates every possibilistic key not implied by the given set. We also establish an algorithm for the discovery of all possibilistic keys that are satisfied by a given possibilistic relation. It is shown that the computational complexity of computing possibilistic Armstrong relations is precisely exponential in the input, and the decision variant of the discovery problem is NP-complete as well as W[2]-complete in the size of the possibilistic key. Further applications of possibilistic keys in constraint maintenance, data cleaning, and query processing are illustrated by examples. The computation of possibilistic Armstrong relations and discovery of possibilistic keys from possibilistic relations have been implemented as prototypes. Extensive experiments with these prototypes provide insight into the size of possibilistic Armstrong relations and the time to compute them, as well as the time it takes to compute a cover of the possibilistic keys that hold on a possibilistic relation, and the time it takes to remove any redundant possibilistic keys from this cover.
Embedded uniqueness constraints represent unique column combinations embedded in complete fragments of incomplete data. In contrast to SQL UNIQUE constraints, they offer a principled separation of completeness and uniqueness requirements and are capable of exploiting more resource-conscious index structures. The latter help relational database systems to be more efficient in enforcing entity and referential integrity, and in evaluating common types of queries.
Data profiling is an enabler for efficient data management and effective analytics. The discovery of data dependencies is at the core of data profiling. We conduct the first study on the discovery of embedded uniqueness constraints (eUCs). These constraints represents unique column combinations embedded in complete fragments of incomplete data. We showcase their implementation as filtered indexes, and their application in integrity management and query optimization. We show that the decision variant of discovering a minimal eUC is NP-complete and W[2]-complete. We characterize the maximum possible solution size, and show which families of eUCs attain that size. Despite the challenges, experiments with real-world and synthetic benchmark data show that our column(row)-efficient algorithms perform well with a large number of columns(rows), and our hybrid algorithm combines ideas from both. We show how to rank eUCs to help identify relevant eUCs.
Let $n\geqslant 3$ be a natural number. We study the problem to find the smallest $r$ such that there is a family $\mathcal{A}$ of 2-subsets and 3-subsets of $[n]=\{1,2,...,n\}$ with the following properties: (1) $\mathcal{A}$ is an antichain, i.e. no member of $\mathcal A$ is a subset of any other member of $\mathcal A$, (2) $\mathcal A$ is maximal, i.e. for every $X\in 2^{[n]}\setminus\mathcal A$ there is an $A\in\mathcal A$ with $X\subseteq A$ or $A\subseteq X$, and (3) $\mathcal A$ is $r$-regular, i.e. every point $x\in[n]$ is contained in exactly $r$ members of $\mathcal A$. We prove lower bounds on $r$, and we describe constructions for regular maximal antichains with small regularity.
Driven by the dominance of the relational model and the requirements of modern applications, we revisit the fundamental notion of a key in relational databases with NULL. In SQL, primary key columns are NOT NULL, and UNIQUE constraints guarantee uniqueness only for tuples without NULL. We investigate the notions of possible and certain keys, which are keys that hold in some or all possible worlds that originate from an SQL table, respectively. Possible keys coincide with UNIQUE, thus providing a semantics for their syntactic definition in the SQL standard. Certain keys extend primary keys to include NULL columns and can uniquely identify entities whenever feasible, while primary keys may not. In addition to basic characterization, axiomatization, discovery, and extremal combinatorics problems, we investigate the existence and construction of Armstrong tables, and describe an indexing scheme for enforcing certain keys. Our experiments show that certain keys with NULLs occur in real-world data, and related computational problems can be solved efficiently. Certain keys are therefore semantically well founded and able to meet Codd’s entity integrity rule while handling high volumes of incomplete data from different formats.
Integrity constraints capture relevant requirements of an application that should be satisfied by every state of the database. The theory of integrity constraints is largely a theory over relations. To make data processing more efficient, SQL permits database states to be partial bags that can accommodate incomplete and duplicate information. Integrity constraints, however, interact differently on partial bags than on the idealized special case of relations. In this current paper, we study the implication problem of the combined class of general cardinality constraints and not-null constraints on partial bags. We investigate structural properties of Armstrong tables for general cardinality constraints and not-null constraints, and prove exact conditions for their existence. For the fragment of general max-cardinality constraints, unary min-cardinality constraints and not-null constraints we show that the effort for constructing Armstrong tables is precisely exponential. For the same fragment we provide an axiomatic characterization of the implication problem. The major tool for establishing our results is the Hajnal and Szemerédi theorem on the equitable colorings of graphs.
Some inequalities for cross-unions of families of finite sets are proved that are related to the problem of minimizing the union-closure of a uniform family of given size. The cross-union of two families F and G of subsets of [n]={1,2,…,n} is the family F∨G={F∪G:F∈F,G∈G}. It is shown that |F∨G|/|F|≥|G∨Bn|/2n, where Bn denotes the power set of [n]. Besides, the problem of minimizing |F∨G| over all union-closed F and G generated by a given number r of singletons and a given number s>(r2) of two-sets, respectively, is solved.
Possibility theory is applied to introduce and reason about the fundamental notion of a key for uncertain data. Uncertainty is modeled qualitatively by assigning to tuples of data a degree of possibility with which they occur in a relation, and assigning to keys a degree of certainty which says to which tuples the key applies. The associated implication problem is characterized axiomatically and algorithmically. It is shown how sets of possibilistic keys can be visualized as possibilistic Armstrong relations, and how they can be discovered from given possibilistic relations. It is also shown how possibilistic keys can be used to clean dirty data by revising the belief in possibility degrees of tuples.
In standard SQL database management systems primary key columns are NOT NULL by default. While NULL columns may be included in unique constraints, such constraints only ensure uniqueness for tuples which do not feature any null marker occurrences in the columns involved, and do not fulfil the same function as primary keys. In this work we investigate the notions of possible and certain keys, which are intuitive and differ only in their treatment of null markers. It turns out that possible keys capture the unique constraint of SQL, while certain keys extend primary keys to include NULL columns, and can be used for similar purposes. In addition to basic characterization, axiomatization, and simple discovery approaches for possible and certain keys, we investigate the existence and construction of Armstrong tables, extremal set problems, and describe an indexing scheme for enforcing certain keys. Our experiments show that certain keys with NULLs do occur in real-world databases, and that related computational problems can be solved efficiently. Certain keys are semantically well-founded, achieve the goal of Codd’s entity integrity rule and offer more flexibility for data entry than primary keys.
We present an approach for data encoding and recovery of lost information in a distributed database system. The dependencies between the informational redundancy of the database and its recovery rate are investigated, and fast recovery algorithms are developed.
any n;k and positive real numbers ; , we determine all SFFA and all SMFA of minimum weight jAj + jBj . Based on this, asymptotic results on SMFA with minimum size and minimum BLYM value, respectively, are derived.
Jerrold R. Griggs合作论文数Department of Mathematics
University of South Carolina6