Can AI systems discover genuinely new knowledge through iterative self improvement, and if so, at what cost? We introduce the NOVA framework, which models the common “generate, verify, accumulate, retrain” loop as an adaptive sampling process over a knowledge space. We identify sufficient conditions under which accumulated genuine knowledge eventually covers a finite domain, and show how their violations produce distinct failure modes: contamination, forgetting, exploration failure, and acceptance failure. We then analyze imperfect verification and identify a contamination trap: as easy-to-find knowledge is exhausted, the model mass assigned to new valid artifacts shrinks, so even small false-positive rates can cause invalid artifacts to enter the knowledge base faster than genuine discoveries. We clarify that Good–Turing estimation is a local batch-diversity diagnostic, not an estimator of the historically undiscovered valid mass that governs long-term discovery. Under a separate tail-equivalence assumption relating the model's effective discovery distribution to a Zipf law with exponent α>1, we prove that the cumulative generation cost required to obtain D distinct genuine discoveries satisfies R_cum(D)=Θ(c_genD^α), where c_gen is the per-candidate generation cost. This scaling law quantifies asymptotic diminishing returns as the discovery frontier advances. Finally, we formalize human amplification through guidance, generation, and verification, explaining why expert input is most valuable near autonomous exploration barriers.
We introduce and analyze a discrete soft-decision channel called the linear reliability channel (LRC) in which the soft information is the rank-ordering of the received symbol reliabilities. We prove that the LRC is an appropriate approximation to a general class of binary-input, continuous-output channels when the noise variance is high. The central feature of the LRC is that its combinatorial nature allows for an extensive mathematical analysis of the channel and its corresponding hard- and soft-decision maximum-likelihood (ML) decoders. In particular, we establish explicit error exponents for ML decoding in the LRC when using random codes under both hard- and soft-decision decoding. This analysis allows for a direct, quantitative evaluation of the relative advantage of soft-decision decoding. The discrete geometry of the LRC is distinct from that of the BSC, which is characterized by the Hamming weight, offering a new perspective on code construction for soft-decision settings.
Communications in highly dynamic channels relying on training-based channel estimation experience a trade-off between increasing channel measurement accuracy by sending more frequent training sequences and increasing data rate by sending fewer training sequences. Simultaneously, most communication systems use forward error correction to enable error detection and correction at the receiver. This paper presents decoder-provided pilots for time-varying channels by using decoded codewords as training sequences to update the channel estimate at the receiver. In contrast to approaches such as data-aided channel estimation, decision-feedback equalization, joint channel estimation and error correction, and turbo equalization, the decoder-provided pilots approach is non-iterative, which is ideal for low-latency requirements in highly dynamic scenarios. Furthermore, it is modulation-, code-, and decoder-agnostic, meaning it can be implemented on top of virtually any communication system that uses forward error correction. From an information-theoretic perspective, we derive the fundamental limits of decoder-provided pilots' ability to simultaneously sense the channel and transmit data. Simulation results demonstrate that decoder-provided pilots significantly improve performance, that when coding across frequency, soft-output can further enhance performance, and that when coding across time, short codes can outperform long codes of the same rate in fast-fading channels.
Proposals to reduce the guesswork of guessing random additive noise decoding (GRAND) by leveraging codebook structure can introduce an imprecise guessing order, which can degrade the block error rate (BLER). We establish one can preserve guesswork reduction while eliminating BLER degradation through dynamic list decoding terminated based on soft output GRAND's error probability estimate. We illustrate the approach with a method inspired by published literature and compare performance with guessing codeword decoding (GCD). We establish that it is possible to provide the same BLER performance as GCD while reducing guesswork by up to a factor of 32.
Guessing Random Additive Noise Decoding (GRAND) and its variants, known for their near-maximum likelihood performance, have been introduced in recent years. One such variant, Segmented GRAND, reduces decoding complexity by generating only noise patterns that meet specific constraints imposed by the linear code. In this paper, we introduce a new method to efficiently derive multiple constraints from the parity check matrix. By applying a random invertible linear transformation and reorganizing the matrix into a tree structure, we extract up to log2(n) constraints, reducing the number of decoding queries while maintaining the structure of the original code for a code length of n. We validate the method through theoretical analysis and experimental simulations.
Inter symbol interference (ISI), which occurs in a wide variety of channels, is a result of time dispersion. It can be mitigated by equalization, which results in noise coloring. Inspired by the development of Approximate Independence in statistical physics, for such colored noise we propose a decoder called Ordered Reliability Bits Guessing Random Additive Noise Decoding (ORBGRAND-AI) that operates without the need for turbo equalization or interleaving. By foregoing interleaving, ORBGRAND-AI can deliver the same, or lower, block error rate (BLER) for the same amount of energy per information bit in an ISI channel as a state-of-the-art soft input decoder, such as Cyclic Redundancy Check Assisted-Successive Cancellation List (CA-SCL) decoding, with an interleaver. To assess the decoding performance of ORBGRAND-AI, we consider delay tap models and their associated colored noise. In particular, we examine a two-tap dicode ISI channel as well as an ISI channel derived from data from RFView, a physics-informed modeling and simulation tool. We investigate the dicode and RFView channel under a variety of imperfect channel state information assumptions and show that a second order autoregressive model adequately represents the RFView channel effect.
Blockchain and other decentralized databases, known as distributed ledgers, are designed to store information online where all trusted network members can update the data with transparency. The dynamics of ledger's development can be mathematically represented by a directed acyclic graph (DAG). One essential property of a properly functioning shared ledger is that all network members holding a copy of the ledger agree on a sequence of information added to the ledger, which is referred to as consensus and is known to be related to a structural property of DAG called one-endedness. In this paper, we consider a model of distributed ledger with sequential stochastic arrivals that mimic attachment rules from the IOTA cryptocurrency. We first prove that the number of leaves in the random DAG is bounded by a constant infinitely often through the identification of a suitable martingale, and then prove that a sequence of specific events happens infinitely often. Combining those results we establish that, as time goes to infinity, the IOTA DAG is almost surely one-ended.
In cases for which there is no suspect, national forensic databases provide a mechanism by which to generate investigatory leads. National forensic DNA databases, however, have restrictions on what data to load. For example, uploading inferred alleles from DNA data that is a mixture of more than two contributors may be disallowed, leading to unresolved cases. A single-cell strategy has the potential to overcome the mixture gap by isolating each cell at the front-end of the pipeline. Once DNA signatures from each cell are obtained, they are clustered into groups. This is followed by asserting the probability we observe the data in the cluster had a person carrying genotype g donated. On applying Bayes' Rule, we obtain the probability of a genotype given the data in a cluster and model. If this probability is near one, it means that only one genotype reasonably explains the data and this genotype can be used in a national database query. Good clustering, therefore, is an invaluable step in single-cell forensic interpretation and it is for this reason we examine the fortitude of two clustering approaches - i.e., model-based clustering (MBC) and forensic-aware clustering (FAC) - within an end-to-end single-cell predictor named EESCIt™. Using proper scoring rules, we report the performance of our probabilistic single-cell evaluator and structure the analytics into categories of Salience, Legitimacy and Credibility (SLC). With Salience referring to the applicability of a technology to meet an actor's needs, we begin by discussing the relevance of single cell reports to forensic actors. Regarding Legitimacy, we determined the proportion of admixtures giving correct and incorrect cluster numbers and found that FAC returned correct cluster numbers for all admixtures tested. With improved clustering, 90 % of the loci returned only one credible genotype and it was the correct one, which improves on MBC's 84 %. We then examined the Brier Score and decomposed it into calibration and refinement. We show that the FAC-centered system returned better calibration scores than the MBC one, which was driven by its improved clustering performance. Regarding Credibility, we found that the FAC-based system also returned better refinement scores. With FAC being more Legitimate and Credible than an MBC system for single-cell forensics, we adopt it into EESCIt™, therein creating the first end-to-end single-cell probabilistic system able to address single-cell queries about how many donors there were, and who they were.
A fully integrated hardware design of the universal maximum likelihood Guessing Random Additive Noise Decoding (GRAND) algorithm implemented in 40 nm CMOS is presented. It is shown how this integrated hard-detection decoder, which is designed to process component codes of up to 128 bits in length, can be extended to efficiently decode product codes as long as 16,384 bits using the Iterative GRAND (IGRAND) algorithm. Pipelined stages provide throughput gain and dynamic energy savings when channel noise conditions improve. The chip allows for decoding product codes with two distinct component codes due to its ability to interleave between two codebooks without any switch-over time. Measurements demonstrate the decoder's accuracy and efficiency in decoding a broad selection of product codes, including the capacity-achieving random linear product codes. The chip consumes an average energy of 30.6 pJ/b with a latency of 1.04 mu s when decoding the BCH(127,106,7) component code at 68 MHz from 1.1 V at a bit flip probability of 10(-5). Using a single chip to decode a BCH(127,106,7)(2) product code which results in 16,129-bit code of rate 0.68, we demonstrate an average energy consumption of 61.2 pJ/b with an average latency of 265 mu s for the same operating conditions.
In addition to a proposed codeword, error correction decoders that provide blockwise soft output (SO) return an estimate of the likelihood that the decoding is correct. Following Forney, such estimates are traditionally only possible for list decoders where the soft output is the likelihood that a decoding is correct given it is assumed to be in the list. Recently, it has been established that Guessing Random Additive Noise Decoding (GRAND), Guessing Codeword Decoding (GCD), Ordered Statistics Decoding (OSD), and Successive Cancellation List (SCL) decoding can provide more accurate soft output, even without list decoding. Central to the improvement is a per-decoding estimate of the likelihood that a decoding has not been found that can be readily calculated during the decoding process. Here we explore how linear codebook constraints can be employed to further enhance the precision of such SO. We evaluate performance by adapting a forecasting statistic called the Brier Score. Results indicate that the SO generated by the approach is essentially as accurate as the maximum a posteriori estimate.
We establish that it is possible to extract accurate blockwise and bitwise soft output (SO) from Guessing Codeword Decoding (GCD) with minimal additional computational complexity by considering it through the lens of Guessing Random Additive Noise Decoding (GRAND). Blockwise SO can be used to control decoding misdetection rate, while bitwise SO results in a soft-input soft-output (SISO) decoder that can be used for efficient iterative decoding of long, high redundancy codes.
We introduce an algorithm for approximating the codebook probability that is compatible with all successive cancellation (SC)-based decoding algorithms, including SC list (SCL) decoding. This approximation is based on an auxiliary distribution that mimics the dynamics of decoding algorithms with an SC decoding schedule. Based on this codebook probability and SCL decoding, we introduce soft-output SCL (SO-SCL) to generate both blockwise and bitwise soft-output (SO). Using that blockwise SO, we first establish that, in terms of both block error rate (BLER) and undetected error rate (UER), SO-SCL decoding of dynamic Reed-Muller (RM) codes significantly outperforms the CRC-concatenated polar codes from 5G New Radio under SCL decoding. Moreover, using SO-SCL, the decoding misdetection rate (MDR) can be constrained to not exceed any predefined value, making it suitable for practical systems. Proposed bitwise SO can be readily generated from blockwise SO via a weighted sum of beliefs that includes a term where SO is weighted by the codebook probability, resulting in a soft-input soft-output (SISO) decoder. Simulation results for SO-SCL iterative decoding of product codes and generalized LDPC (GLDPC) codes, along with information-theoretical analysis, demonstrate significant superiority over existing list-max and list-sum approximations.
Guessing Codeword Decoding (GCD) is a recently proposed soft-input forward error correction decoder for arbitrary binary linear codes. Inspired by recent proposals that leverage binary linear codebook structure to reduce the number of queries made by Guessing Random Additive Noise Decoding (GRAND), for binary linear codes that include a full-message single parity-check (SPC) bit, we show that it is possible to reduce the number of queries made by GCD by a factor of up to 2 with the greatest guesswork reduction realized at lower SNRs, without impacting decoding precision. Codes without a full-message SPC can be modified to include one by changing a column of the generator matrix to obtain a decoding complexity advantage, and we demonstrate that this can often be done without losing decoding precision. To practically avail of the complexity advantage, a noise effect pattern generator capable of producing sequences for given Hamming weights, such as the landslide algorithm developed for ORBGRAND, is necessary.
We present a communication scheme using guessing random additive noise decoding (GRAND) to improve flexibility and reliability of the existing compressed error (CE) framework. The CE framework uses information feedback to construct follow-up transmissions by compressing previous noise realizations, offering high reliability at a code rate close to the forward channel capacity. The channel decoding algorithm GRAND allows us to efficiently maintain this performance in noisy feedback settings by shifting redundancy for forward message protection to the feedback channel. Our scheme, GRAND-CE, is therefore appropriate for cases where forward and feedback channel usage costs are asymmetric, e.g. uplink communications. GRAND-CE offers super-exponential error rate performance as channel use increases, with finite usage of a noisy feedback channel. Unlike the traditional forward error correction model, the receiver performs error correction encoding and the sender handles decoding. We also propose a technique for pipelining sequential transmissions to maintain fixed forward transmission length and good feedback channel coding performance.
Multiple input multiple output (MIMO) systems and nonorthogonal multiple access (NOMA) methods are both valuable techniques for modern and future communication systems. MIMO is commonly used in scenarios such as mobile communications and the Internet of Things in order to improve signal quality or increase channel capacity, while NOMA is important as its methods can be used to service the growing number of users. A combination of the two techniques should thus be investigated. We integrate MIMO methods such as the Alamouti space time block code (STBC) with Guessing Random Additive Noise Decoding Aided Macrosymbol (GRAND-AM), which was previously proposed for single input single output (SISO) systems. Originally, GRAND-AM handled the multiple access interference (MAI) using multiple access channel (MAC) codes assigned to each user, along with joint multi-user detection and decoding. We consider how using the Alamouti STBC as a replacement of the MAC codes performs with GRAND-AM. We show that GRAND-AM is able to outperform the multiuser Vertical-Bell Laboratories Layered Space-Time (V-BLAST) receiver by similar to 2.5dB when used to handle two users using the Alamouti STBC. We show how GRAND-AM with Alamouti STBCs outperforms the V-BLAST receiver, even when there is a larger number of users in the MIMO NOMA system. We additionally investigate methods of reducing the complexity of GRAND-AM, as the MIMO NOMA problem is more complex than the SISO NOMA problem. We then show that the complexity reducing measures do not lead to large losses and can still outperform V-BLAST receivers, giving incentive to use GRAND-AM for MIMO NOMA systems.
We propose a novel demodulation technique that leverages developments in guesswork-based forward error correction decoders and variable-length bit-to-symbol mappings. For most common channel models, the optimal modulation schemes are known to require nonuniform probability distributions over signal points, which presents practical challenges. An established way to map uniform binary sources to non-uniform symbol distributions is to assign a different number of bits to different constellation points. Doing so, however, means that erroneous demodulation at the receiver can lead to bit insertions or deletions, turning a channel with Hamming-type errors into an insertion-deletion channel. The demodulator we propose provides error detection and correction through the use of a low-overhead padding bit sequence. We evaluate the performance of the proposed demodulator in various channel models and various communication settings. We verify that the demodulator successfully corrects the insertion-deletion errors. Using the proposed demodulator, we study different constellation design schemes and how they behave in different channel conditions. Overall, we observe considerable gains that suggest, in some circumstances, one may improve the throughput while keeping the error rate the same.
We establish that a large, flexible class of long, high redundancy error correcting codes can be efficiently and accurately decoded with guessing random additive noise decoding (GRAND). Performance evaluation demonstrates that it is possible to construct simple product codes with lengths of approximately 200 to 4000 bits and rates between 0.2 and 0.8 that outperform low-density parity-check (LDPC) codes from the 5G New Radio standard in both AWGN and fading channels. The concatenated structure enables many desirable features, including: low-complexity hardware-friendly encoding and decoding; significant flexibility in length and rate through modularity; and high levels of parallelism in encoding and decoding that enable low latency. Central is the development of a method through which any soft-input (SI) GRAND algorithm can provide soft-output (SO) in the form of an accurate a-posteriori estimate of the likelihood that a decoding is correct or, in the case of list decoding, the likelihood that each element of the list is correct. The distinguishing feature of soft-output GRAND (SOGRAND) is the provision of an estimate that the correct decoding has not been found, even when providing a single decoding. Per-block SO can be converted into accurate per-bit SO by a weighted sum that includes a term for the SI. Implementing SOGRAND adds negligible computation and memory to the existing decoding process, and using it results in a practical, low-latency alternative to LDPC codes.
Long, powerful soft detection forward error correction codes are typically constructed by concatenation of shorter component codes that are decoded through iterative Soft-Input Soft-Output (SISO) procedures. The current gold-standard is Low Density Parity Check (LDPC) codes, which are built from weak single parity check component codes that are capable of producing accurate SO. Due to the recent development of SISO decoders that produce highly accurate SO with codes that have multiple redundant bits, square product code constructions that can avail of more powerful component codes have been shown to be competitive with the LDPC codes in the 5G New Radio standard in terms of decoding performance while requiring fewer iterations to converge. Motivated by applications that require more powerful low-rate codes, in the present paper we explore the possibility of extending this design space by considering the construction and decoding of cubic tensor codes.
We present a novel method for error correction in the presence of fading channel estimation errors (CEE). When such errors are significant, considerable performance losses can be observed if the wireless transceiver is not adapted. Instead of refining the estimate by increasing the pilot sequence length or improving the estimation algorithm, we propose two new approaches based on Guessing Random Additive Noise Decoding (GRAND) decoders. The first method involves testing multiple candidates for the channel estimate located in the complex neighborhood around the original pilot-based estimate. All these candidates are employed in parallel to compute log-likelihood ratios (LLR). These LLRs are used as soft input to Ordered Reliability Bits GRAND (ORBGRAND). Posterior likelihood formulas associated with ORBGRAND are then computed to determine which channel candidate leads to the most probable codeword. The second method is a refined version of the first approach accounting for the presence of residual CEE in the LLR computation. The performance of these two techniques is evaluated for [128,112] 5G NR CA-Polar and CRC codes. For the considered settings, block error rate (BLER) gains of several dBs are observed compared to cases where CEE is ignored.
We present a framework that can exploit the tradeoff between the undetected error rate (UER) and block error rate (BLER) of polar-like codes. It is compatible with all successive cancellation (SC)-based decoding methods and relies on a novel approximation that we call codebook probability. This approximation is based on an auxiliary distribution that mimics the dynamics of decoding algorithms following an SC decoding schedule. Simulation results demonstrates that, in the case of SC list (SCL) decoding, the proposed framework outperforms the state-of-art approximations from Forney's generalized decoding rule for polar-like codes with dynamic frozen bits. In addition, dynamic Reed-Muller (RM) codes using the proposed generalized decoding significantly outperform CRC-concatenated polar codes decoded using SCL in both BLER and UER.