The coverage depth problem in DNA data storage is about computing the expected number of reads needed to recover all encoded strands. Given a generator matrix of a linear code, this quantity equals the expected number of randomly drawn columns required to obtain full rank. While MDS codes are optimal when they exist, i.e., over large fields, practical scenarios may rely on structured code families defined over small fields. In this work, we develop combinatorial tools to solve the DNA coverage depth problem for various linear codes, based on duality arguments and the notion of extended weight enumerator. Using these methods, we derive closed formulas for the simplex, Hamming, ternary Golay, extended ternary Golay, and first-order Reed-Muller codes. The centerpiece of this paper is a general expression for the coverage depth of a linear code in terms of the weight distributions of its higher-field extensions.
This paper tackles two problems that are relevant to coding for insertions and deletions. These problems are motivated by several applications, among them is reconstructing strands in DNA-based storage systems. Under this paradigm, a word is transmitted over some fixed number of identical independent channels and the goal of the decoder is to output the transmitted word or some close approximation of it. The first part of this paper studies the deletion channel that deletes a symbol with some fixed probability $p$, while focusing on two instances of this channel. Since operating the maximum likelihood (ML) decoder in this case is computationally unfeasible, we study a slightly degraded version of this decoder for two channels and its expected normalized distance. We identify the dominant error patterns and based on these observations, it is derived that the expected normalized distance of the degraded ML decoder is roughly $\frac{3q-1}{q-1}p^2$, when the transmitted word is any $q$-ary sequence and $p$ is the channel's deletion probability. We also study the cases when the transmitted word belongs to the Varshamov Tenengolts (VT) code or the shifted VT code. Additionally, the insertion channel is studied as well as the case of two insertion channels. These theoretical results are verified by corresponding simulations. The second part of the paper studies optimal decoding for a special case of the deletion channel, the $k$-deletion channel, which deletes exactly $k$ symbols of the transmitted word uniformly at random. In this part, the goal is to understand how an optimal decoder operates in order to minimize the expected normalized distance. A full characterization of an efficient optimal decoder for this setup, referred to as the maximum likelihood* (ML*) decoder, is given for a channel that deletes one or two symbols.
DNA has emerged as a promising medium for long-term digital data storage due to its exceptional density and durability. This work extends recent studies of the coverage depth problem by considering practical DNA storage settings and identifying error-correcting code constructions that minimize the required coverage depth. Focusing on random access and noisy channel models, we introduce a Markovian framework for analyzing coverage depth in DNA-based storage systems under noisy reading channels and random access queries and provide supporting simulations. This framework enables the computation of both the expectation and distribution of the sample size required to recover arbitrary subsets of files, and provides insight into the coverage efficiency of different coding strategies.
DNA labeling is an important tool in molecular biology and biotechnology to visualize, detect, and support the study of DNA at the molecular level. Existing literature has focused primarily on using standard labels, where each label is defined by a fixed q-ary pattern whose occurrences are detected along the sequence. Such schemes, however, impose limitations on the achievable labeling capacity and often require a large label set to approach the maximum possible capacity of log2q bits per symbol for q-ary sequences (and 2 bits per symbol in the quaternary DNA alphabet).In this work, we propose a novel approach for DNA labeling using composite labels, in which each composite symbol can store a mixture of the standard letters, consequently, allowing multiple patterns to be recognized by a single composite label. We show that composite labels strictly increase the achievable labeling capacity compared with standard labels, and substantially reduce the minimum number of labels required to attain the maximum labeling capacity.
We initiate the study of DNA-based distributed storage systems, where information is encoded across multiple DNA data storage containers to achieve robustness against container failures. In this setting, data are distributed over M containers, and the objective is to guarantee that the contents of any failed container can be reliably reconstructed from the surviving ones. Unlike classical distributed storage systems, DNA data storage containers are fundamentally constrained by sequencing technology, since each read operation yields the content of a uniformly random sampled strand from the container. Within this framework, we consider several erasure-correcting codes and analyze the expected recovery time of the data stored in a failed container. Our results are obtained by analyzing generalized versions of the classical Coupon Collector's Problem, which may be of independent interest.
Data storage in DNA has recently emerged as a promising archival solution, offering space-efficient and long-lasting digital storage. DNA’s ultra-high density and durability make it an attractive medium for long-term data storage. Among the various approaches, combinatorial DNA encoding further enhances this potential by increasing the logical density through the use of combinations of DNA shortmers, where each sequence position is represented by a set of predefined short DNA fragments. This approach allows for the encoding of a larger data volume using fewer synthesis cycles. However, this method introduces unique challenges, particularly in terms of synthesis and sequencing errors. In this study, we focus on the characterization of errors in combinatorial DNA-based storage systems. Our analysis revealed that asymmetric combinatorial erasure errors, defined as the omission of a single shortmer from the set defining the combinatorial letter, are a prevalent error type in combinatorial DNA-based storage, particularly in large-scale systems where read coverage is limited. We analyzed two previously published datasets, observing a high frequency of erasure errors, where missing sequences obstruct the reconstruction of specific combinatorial letters. To better understand these observations, we conducted a large-scale experimental proof-of-concept of a combinatorial DNA-based storage system and evaluated the error characteristics of the system. Our analysis confirmed that erasure errors become increasingly prominent with reduced sequencing depth. We demonstrated that below 50 reads per sequence, the frequency of erasure errors sharply increased. We developed an asymmetric error-correcting code specifically designed to address these errors. The code utilizes tensor-product (TP) codes to integrate standard erasure and substitution-correcting codes (such as Reed-Solomon (RS) codes) with Varshamov-Tenengolts (VT) codes, which are asymmetric error-correcting codes. We validated the performance of our new error correction code both in simulations and in a second large-scale experiment. The experimental comparison was designed to directly compare our suggested code with the more straightforward 2D Reed-Solomon (2D RS) scheme. Our method consistently outperformed the 2D RS scheme, particularly in scenarios dominated by erasure errors. Notably, in the second large-scale experiment, our method demonstrated superior decoding accuracy even under low coverage conditions, where the traditional 2D RS approach struggled to decode the data. Our findings demonstrate the importance of tailored error correction schemes in DNA-based data storage. By directly addressing the asymmetric nature of errors in combinatorial DNA, our method provides improved decoding accuracy under a wide range of conditions. The integration of such tailored error correction methods with existing DNA-based data storage technologies has the potential to deliver more reliable and scalable DNA-based data applications.
The sequence reconstruction problem, introduced by Levenshtein in 2001, considers a communication setting in which a sender transmits a codeword and the receiver observes K independent noisy versions of this codeword. In this work, we study the problem of efficient reconstruction when each of the K outputs is corrupted by a q-ary discrete memoryless symmetric (DMS) substitution channel with substitution probability p. Focusing on Reed-Solomon (RS) codes, we adapt the Koetter-Vardy soft-decision decoding algorithm to obtain an efficient reconstruction algorithm. For sufficiently large blocklength and alphabet size, we derive an explicit rate threshold, depending only on (p, K), such that the transmitted codeword can be reconstructed with arbitrarily small probability of error whenever the code rate R lies below this threshold.
In this paper, we study the Random Access Problem in DNA storage, which addresses the challenge of retrieving a specific information strand from a DNA-based storage system. In this framework, the data is represented by k information strands which represent the data and are encoded into n strands using a linear code. Then, each sequencing read returns one encoded strand which is chosen uniformly at random. The goal under this paradigm is to design codes that minimize the expected number of reads required to recover an arbitrary information strand. We fully solve the case when k=2, showing that the best possible code attains a random access expectation of 1+2/√(2)+1≈ 0.914· 2 for q large enough. Moreover, we generalize a construction from , specifically to k=3, for any value of k. Our construction uses B_k-1 sequences over ℤ_q-1, that always exist over large finite fields. We show that for every k≥ 4, this generalized construction outperforms all previous constructions in terms of reducing the random access expectation.
this work, we study linear error-correcting codes against adversarial insertion-deletion (indel) errors. While most constructions for the indel model are nonlinear, linear codes offer compact representations, efficient encoding, and decoding algorithms, making them highly desirable. A key challenge in this area is achieving rates close to the half-Singleton bound for efficient linear codes over finite fields. We improve upon previous results by constructing explicit codes over Fq2, linear over F-q, with rate 1/2-delta - epsilon that can efficiently correct a delta-fraction of indel errors, where q = O(epsilon(-4)). Additionally, we construct fully linear codes over Fq with rate 1/2-2 root delta - epsilon that can also efficiently correct delta-fraction of indels. These results significantly advance the study of linear codes for the indel model, bringing them closer to the theoretical half-Singleton bound. We also generalize the half-Singleton bound, for every code C subset of Fn linear over E subset of F a subfield of F, such that C has the ability to correct delta-fraction of indels, the rate is bounded by (1-delta)/2.
Constrained coding is a fundamental field in coding theory that tackles efficient communication through constrained channels. While channels with fixed constraints have a general optimal solution, there is increasing demand for parametric constraints that are dependent on the message length. Several works have tackled such parametric constraints through iterative algorithms, yet they require complex constructions specific to each constraint to guarantee convergence through monotonic progression. In this paper, we propose a universal framework for tackling any parametric constrained-channel problem through a novel simple iterative algorithm. By reducing an execution of this iterative algorithm to an acyclic graph traversal, we prove a surprising result that guarantees convergence with efficient average time complexity even without requiring any monotonic progression. We demonstrate the effectiveness of this universal framework by applying it to a variety of both local and global channel constraints. We begin by exploring the local constraints involving illegal substrings of variable length, where the universal construction essentially iteratively replaces forbidden windows. We apply this local algorithm to the minimal periodicity, minimal Hamming weight, local almost-balanced Hamming weight and the previously-unsolved minimal palindrome constraints. We then continue by exploring global constraints, and demonstrate the effectiveness of the proposed construction on the repeat-free encoding, reverse-complement encoding, and the open problem of global almost-balanced encoding. For reverse-complement, we also tackle a previously-unsolved version of the constraint that addresses overlapping windows. Overall, the proposed framework generates state-of-the-art constructions with significant ease while also enabling the simultaneous integration of multiple constraints for the first time.
SUMMARY DNA data storage allows sequences to be defined without biological constraints, yet readout workflows still depend on generic end-repair/dA-tailing chemistry. We developed NinjaSeq, a type IIS restriction endonuclease library-preparation strategy that incorporates recognition sites into primer flanks, enabling digestion to generate adapter-compatible overhangs and eliminating the need for conventional end preparation. By combining this chemistry with constrained coding that excludes internal recognition motifs, NinjaSeq produced sequencing quality and decoding performance consistent with standard protocols while reducing reagent burden and simplifying processing, including compatibility with one-pot restriction-ligation. The same sequence-directed design also enables physical random access during library preparation: targeting file-specific flanking sites enriched a desired file from a mixed pool by about sixteen-fold in a proof-of-concept experiment. These results position NinjaSeq as a practical ONT readout approach for DNA data storage. HIGHLIGHTS NinjaSeq replaces end-repair/dA-tailing with REases for nanopore sequencing Constrained encoding excludes recognition motifs to protect payloads from cleavage NinjaSeq achieves decoding accuracy comparable to standard library preparation Designing file-specific RRS enables random access during library preparation
Motivated by their applications in DNA-based storage systems, codes capable of correcting consecutive deletions have attracted significant attention. An important class of such codes consists of those that can correct multiple consecutive deletion errors, commonly referred to as multiple b-burst deletion-correcting codes. In this paper, we investigate the fundamental limits of multiple b-burst deletion-correcting codes. Specifically, we first characterize several structural properties of the associated deletion balls. Then, leveraging these properties, we derive several upper bounds and a combinatorial lower bound on the maximum size of such codes. As a consequence, our bounds improve upon the previously known results for general parameter regimes and are shown to be asymptotically optimal for certain cases.
We study coverage processes in which each draw reveals a subset of [n], and the goal is to determine the expected number of draws until all items are seen at least once. A classical example is the Coupon Collector's Problem, where each draw reveals exactly one item. Motivated by shotgun DNA sequencing, we introduce a model where each draw is a contiguous window of fixed length, in both cyclic and non-cyclic variants. We develop a unifying combinatorial tool that shifts the task of finding coverage time from probability, to a counting problem over families of subsets of [n] that together contain all items, enabling exact calculation. Using this result, we obtain exact expressions for the window models. We then leverage past results on a continuous analogue of the cyclic window model to analyze the asymptotic behavior of both models. We further study what we call uniform ℓ-regular models, where every draw has size ℓ and every item appears in the same number of admissible draws. We compare these to the batch sampling model, in which all ℓ-subsets are drawn uniformly at random and present upper and lower bounds, which were also obtained independently by Berend and Sher. We conjecture, and prove for special cases, that this model maximizes the coverage time among all uniform ℓ-regular models. Finally, we prove a universal upper bound on the entire class of uniform ℓ-regular models, which illuminates the fact that many sampling models share the same leading asymptotic order, while potentially differing significantly in lower-order terms.
We study the theoretical problem of synthesizing multiple DNA strands under spatial constraints, motivated by large-scale DNA synthesis technologies. In this setting, strands are arranged in an array and synthesized according to a fixed global synthesis sequence, with the restriction that at most one strand per row may be synthesized in any synthesis cycle. We focus on the basic case of two strands in a single row and analyze the expected completion time under this row-constrained model. By decomposing the process into a Markov chain, we derive analytical upper and lower bounds on the expected synthesis time. We show that a simple laggard-first policy achieves an asymptotic expected completion time of (q+3)L/2 for any alphabet of size q, and that no online policy without look-ahead can asymptotically outperform this bound. For the binary case, we show that allowing a single-symbol look-ahead strictly improves performance, yielding an asymptotic expected completion time of 7L/3. Finally, we present a dynamic programming algorithm that computes the optimal offline schedule for any fixed pair of sequences. Together, these results provide the first analytical bounds for synthesis under spatial constraints and lay the groundwork for future studies of optimal synthesis policies in such settings.
We study the combination of two recent coding approaches, in the context of DNA based data storage. Composite DNA alphabets leverage properties of the DNA synthesis and sequencing process. A composite symbol does not represent a single nucleotide, but rather a designed mixture of DNA nucleotides. Using the high multiplicity that is intrinsic to synthesis and sequencing a composite symbol consists of frequencies in the mixture. Rank modulation codes use permutations to represent information. Combining the two, we construct encoding that uses permutations of nucleotide frequencies rather than the exact frequency values. Codes for this approach were addressed in previous work, under Kendall's tau distances. In this work we study deletion and insertion codes. We present bounds and constructions of efficient codes defined over partial permutations.
DNA-based storage offers unprecedented density and durability, but its scalability is fundamentally limited by the efficiency of parallel strand synthesis. Existing methods either allow unconstrained nucleotide additions to individual strands, such as enzymatic synthesis, or enforce identical additions across many strands, such as photolithographic synthesis. We introduce and analyze a hybrid synthesis framework that generalizes both approaches: in each cycle, a nucleotide is selected from a restricted subset and incorporated in parallel. This model gives rise to a new notion of a complex synthesis sequence. Building on this framework, we extend the information rate definition of Lenz et al. and analyze an analog of the deletion ball, defined and studied in this setting, deriving tight expressions for the maximal information rate and its asymptotic behavior. These results bridge the theoretical gap between constrained models and the idealized setting in which every nucleotide is always available. For the case of known strands, we design a dynamic programming algorithm that computes an optimal complex synthesis sequence, highlighting structural similarities to the shortest common supersequence problem. We also define a distinct two-dimensional array model with synthesis constraints over the rows, which extends previous synthesis models in the literature and captures new structural limitations in large-scale strand arrays. Additionally, we develop a dynamic programming algorithm for this problem as well. Our results establish a new and comprehensive theoretical framework for constrained DNA, subsuming prior models and setting the stage for future advances in the field.
As DNA data storage moves closer to practical deployment, minimizing sequencing coverage depth is essential to reduce both operational costs and retrieval latency. This paper addresses the recently studied Random Access Problem, which evaluates the expected number of read samples required to recover a specific information strand from $n$ encoded strands. We propose a novel algorithm to compute the exact expected number of reads, achieving a computational complexity of $O(n)$ for fixed field size $q$ and information length $k$. Furthermore, we derive explicit formulas for the average and maximum expected number of reads, enabling an efficient search for optimal generator matrices under small parameters. Beyond theoretical analysis, we present new code constructions that improve the best-known upper bound from $0.8815k$ to $0.8811k$ for $k=3$, and achieve an upper bound of $0.8629k$ for $k=4$ for sufficiently large $q$. We also establish a tighter theoretical lower bound on the expected number of reads that improves upon state-of-the-art bounds. In particular, this bound establishes the optimality of the simple parity code for the case of $n=k+1$ across any alphabet $q$.
Motivated by DNA data storage, we study the expected number of coded symbols drawn from a linear code until a desired information symbol can be decoded - the random access expectation. We focus on generator matrices with a type of symmetry, conjectured in prior work to be optimal, which we call fully symmetric. We point out an equivalence between binary fully symmetric codes and LT codes. Using this observation, we analyze the random access expectation of binary fully symmetric codes under a peeling decoder, in the large blocklength limit. Under these assumptions, the random access expectation, normalized by the number of information symbols, is at least π/4 ≈ 0.7854, while a value of ≈ 0.7869 is achievable.
This paper introduces and studies staple codes, multi-purpose libraries of DNA oligonucleotides that are designed to self-assemble into long DNA strands via controlled ligation. By construction, staple codes enable flexible data encodings for DNA-based data storage as well as the implementation of logic gates for molecular computing. We analyze the conditions under which the DNA oligonucleotide reliably self-assemble into predetermined structures.
DNA labeling is a tool in molecular biology and biotechnology to visualize, detect, and study DNA at the molecular level. In this process, a DNA molecule is labeled by a set of specific patterns, referred to as labels, and is then imaged. The resulting image is modeled as an (ℓ+1)-ary sequence, where ℓ is the number of labels, in which any non-zero symbol indicates the appearance of the corresponding label in the DNA molecule. The labeling capacity refers to the maximum information rate that can be achieved by the labeling process for any given set of labels. The main goal of this paper is to study the minimum number of labels of the same length required to achieve the maximum labeling capacity of 2 for DNA sequences or log_2q for an arbitrary alphabet of size q. The solution to this problem requires the study of path unique subgraphs of the de Bruijn graph with the largest number of edges. We provide upper and lower bounds on this value. We draw new connections to existing literature that let us prove an asymptotic result as the label length tends to infinity.
P. H. Siegel合作论文数Center for Magnetic Recording Research11