
The sequence reconstruction problem has attracted considerable attention due to its applications in DNA storage. This problem can be modeled as transmitting a sequence over several identical noisy channels, each producing a distinct output, with the goal of reconstructing the original transmitted sequence from these outputs. The central question is to determine the minimum number of channels required to guarantee unique reconstruction of the transmitted sequence. In this paper, we consider channels that introduce k occurrences of (t, s)-bursts, where a (t, s)-burst replaces a substring of length t with another sequence of length s. For given integers t ≥ 1, s ≥ 1, k ≥ 1, n ≥ (k +1)t, we determine the minimum number of channels that guarantees unique reconstruction and design an efficient algorithm to reconstruct the transmitted sequence (of length n). Moreover, we prove that if n = (k + 1)t − 1, reconstruction is impossible in the worst-case scenario.
Repeat-free sequences were recently introduced to reduce errors in the reconstruction process of DNA strings. Such sequences have the property that every k-tuple appears at most once (for predefined k). In this paper, we consider the problem of encoding data into a k-repeat free sequence that also adheres to local constraints. First, we provide encoding and decoding schemes for primitive systems. Under an additional irreducibility assumption on a specific sub-system, we improve upon the best-known capacity result, and show that the capacity of such systems is determined only by the local constraints as long as k ≥ (1 + ε) log(n). With no additional assumptions, we provide encoding and decoding schemes for k ≥ (2 + ε) log(n). We then generalize the encoding and decoding scheme to irreducible systems.
Delegating large-scale computations to service providers is a common practice which raises privacy concerns. This paper studies information-theoretic privacy-preserving delegation of data to a service provider, who may further delegate the computation to auxiliary worker nodes, in order to compute a polynomial over that data at a later point in time. We study techniques which are compatible with robust management of distributed computation systems, an area known as coded computing. Privacy in coded computing, however, has traditionally addressed the problem of colluding workers, and assumed that the server that administrates the computation is trusted. This viewpoint of privacy does not accurately reflect real-world privacy concerns, since normally, the service provider as a whole (i.e., the administrator and the worker nodes) form one cohesive entity which itself poses a privacy risk. This paper explicitly shifts the privacy focus in coded computing from protecting user data only against colluding workers to protecting it against the service provider as a whole, by placing the privacy barrier directly at the user side before the data reaches the service provider. To this end, we leverage the recently defined notion of perfect subset privacy, which guarantees zero information leakage from all subsets of the data up to a certain size, and can be obtained via k-wise independence properties of linear codes. Using known techniques from Reed-Muller decoding, we provide a scheme which enables polynomial computation with perfect subset privacy in straggler-free systems. Most importantly, by studying information supersets in Reed-Muller codes, which were scantly studied and may be of independent interest, we extend the previous scheme to tolerate straggling worker nodes inside the service provider.
This paper establishes an information-theoretic framework for multi-user multi-target multiple-input multiple-output (MIMO) integrated sensing and communication systems under Gaussian signaling and treating-interference-as-noise (TIN) decoding. The analysis characterizes communication performance via weighted sum mutual information (MI) and sensing performance via weighted sum Kullback–Leibler divergence (KLD). To resolve the dimensional heterogeneity of these metrics, a normalization methodology based on single-functionality baselines is introduced. This approach maps the performance utilities onto a dimensionless unit square and defines the normalized achievable region as a compact and convex deterministic inner bound. Fixed-strategy low signal-to-noise ratio (SNR) trajectories are characterized through asymptotic analysis under the fixed-baseline convention. In the low-SNR regime, the linear scaling of MI and the quadratic scaling of KLD drive fixed-strategy tradeoff trajectories to become tangent to the communication axis near the origin. In the high-SNR regime, independent precoding that ignores cross-functional interference is shown to exhibit an asymptotic efficiency collapse in the normalized domain under interference saturation, implying that suppressing the projected leakage order is necessary within the specified persistent-overlap fixed-direction regime. For algorithmic construction, an alternating optimization algorithm based on the minorization-maximization principle is developed to compute scalarized operating points through an implemented rank-one multiple-input single-output block approximation with an objective safeguard. This implementation integrates a sensing-priced weighted minimum mean-square error update with a closed-form sensing covariance refinement and uses the objective safeguard to preserve monotonicity of the accepted objective sequence. Numerical results corroborate the theoretical scaling laws and illustrate improved normalized performance over the independent precoding baseline in the simulated interference-limited settings.
We investigate covert communication in a multi-stream setting where an adversary employs a cumulative sum (CUSUM) detector with a myopic sampling policy (MSP) to monitor multiple candidate streams. Within a sequential change-point detection framework, we characterize performance using metrics such as the average run length to false alarm (ARL2FA) and the average detection delay (ADD). Covertness is defined by ensuring that, at the adversary’s optimal detection threshold, the ratio ADD/γ remains at least 1–ε, where γ denotes the lower bound on ARL2FA. We derive exact expressions for ARL2FA and ADD under MSP and obtain a complete characterization of the ε-covert region of the power parameter θ for Gaussian and BPSK ensembles. We further establish sharp overshoot bounds, showing that the overshoot terms remain of lower order even as the optimal threshold tends to zero. Finally, we compare MSP with the single-stream setting and show that, in the covert regime, the performance gap depends explicitly on the number of monitored streams N. Numerical Fredholm calculations and Monte Carlo simulations validate the theoretical predictions.
Machine learning often involves structured optimization problems with an ℓ1-regularizer to encourage either sparsity or interpretability. While the stochastic regularized dual averaging (RDA) method addresses the limitations of stochastic gradient descent in structured optimization, its theoretical analysis focuses primarily on optimization errors, leaving the generalization behavior to testing examples unexplored. This work bridges this gap by initializing the stability and generalization analysis of stochastic RDA through the lens of algorithmic stability. We establish novel stability and generalization error bounds for three distinct scenarios: convex and smooth losses, strongly convex and smooth losses, and nonsmooth, Lipschitz and convex losses. Additionally, we derive new optimization error bounds for stochastic RDA by removing the bounded gradient assumption for smooth problems. We also develop optimal excess risk bounds for ℓ1-regularized problems, which can be fast under a low-noise condition.
Robust principal component analysis has been extensively studied over the past decade. However, integrated analysis under the simultaneous presence of dense heavy-tailed noise and sparse corruptions remains largely unexplored and challenging. This work addresses this gap by proposing an approach based on minimizing the absolute loss (quantile loss). While the absolute loss is known to be non-differentiable, we establish the convergence dynamics of subgradient descent in PCA under this loss. Notably, we identify a phase transition in the behavior of the loss function with respect to the distance from the iterate to the oracle solution: when the iterate is sufficiently close to the oracle, the loss behaves quadratically; when far, it behaves linearly. Based on this statistical regularity, we design a two-phase step-size scheme which leads to linear convergence. Compared with existing approaches, it doesn’t require prior knowledge of the corruption level, significantly enhancing practical applicability. Numerical experiments strongly support our theoretical findings and the practical performance is evaluated in the real dataset of food balance. Furthermore, we derive minimax optimal rates for noisy matrix and tensor PCA under slice-sparse corruption, which matches with our estimation error rate. In addition, the technical analyses on the properties of the absolute loss, concentration inequalities, and matrix/tensor perturbations may be of independent interest.
Consider a survey asking people for their favorite food. We would like to report the K most popular foods in the sample, along with their associated confidence intervals (CIs). The common practice is to construct K marginal (binomial) CIs at a desired confidence level and expect a matching coverage rate. This approach is genuinely wrong, as a result of selective inference. Specifically, the inferred parameters are not fixed but data-dependent, and may vary among different samples. This means that standard binomial CIs, which implicitly assume fixed parameters, fail to provide the desired coverage rate. In this work we introduce a novel selective inference scheme for the most frequent events in the sample. Our proposed scheme is closed-form, intuitive, simple to apply and guarantees the prescribed confidence level. Further, we introduce a matching lower bound which demonstrates the tightness of our results. We illustrate our proposed scheme in synthetic and real-world experiments, showing a significant improvement over currently known alternatives.
Sequences with low peak-to-average power ratio (PAPR) and low time-domain cross-correlation play a crucial role in the uplink of 4G LTE and 5G NR, serving as sounding reference signals for channel sounding and data demodulation. This paper, for the first time, transforms the time-domain cross-correlation problem into a characterization of sequence PAPR, thereby establishing a theoretical lower bound on the time-domain cross-correlation of sequences and defining Golay orthogonal sequence sets with low PAPR and low time-domain cross-correlation properties. Furthermore, by establishing a connection between mutually orthogonal Hamiltonian paths (H-paths) in the complete graph and generalized Boolean functions, this paper innovatively proposes a construction framework for designing Golay orthogonal sequence sets based on orthogonal H-paths. Simultaneously, this paper discusses the theoretical upper bound on the maximum number of orthogonal H-paths in the complete graph Kn, and proposes two specific construction schemes for orthogonal H-path sets for any odd number n. Based on this scheme, we designed two classes of Golay orthogonal sequence sets, which have better PAPR and time-domain cross-correlation performance compared to well-known Zadoff-Chu sequences used in 4G LTE and 5G NR.
This work studies sparse principal component analysis (PCA) in high dimensions. Given n independent p-dimensional Gaussian samples with covariance Σ := (λ − 1)vv⊤ + Ip, our goal is to estimate v under the assumption of sparsity. On the one hand, if the sparsity level m := ∥v∥0 satisfies m ≲ √ n, algorithms such as covariance thresholding (Krauthgamer et al., 2015) consistently outperform PCA. On the other hand, if m ≫ √ n, it is conjectured that no polynomial-time algorithm can recover v below the detection threshold of PCA. We investigate the “critical” high-dimensional regime, where n, p,m → ∞ with m/ √ n → β and p/n → γ, and study estimators based on kernel PCA, generalizing covariance thresholding. Within this framework, we achieve a fine-grained understanding of signal detection and recovery. Our main result establishes a detection phase transition, analogous to the Baik–Ben Arous–Péché (BBP) transition for PCA: above a signal strength threshold—depending on the kernel function, γ, and β—kernel PCA is informative. Conversely, below the threshold, kernel principal components are asymptotically orthogonal to the signal. Notably, (1) above this threshold, consistent support recovery is possible with high probability, (2) for all β ∈ (0,∞), kernel PCA strictly outperforms PCA, and (3) as β → ∞, kernel PCA and PCA coincide. We identify optimal kernel functions for detection and support recovery, and numerical calculations suggest that soft thresholding is nearly optimal. Our key technical contribution is approximation guarantees for deterministic equivalents of kernel random matrices, which enable sharp estimates of coordinate fluctuations of kernel principal components.
Group testing, a problem with diverse applications across multiple disciplines, traditionally assumes independence across nodes’ states. Recent research, however, focuses on real-world scenarios that often involve correlations among nodes, challenging the simplifying assumptions made in existing models. In this work, we consider a comprehensive model for arbitrary statistical correlation among nodes’ states. To capture and leverage these correlations effectively, we model the problem by hypergraphs, inspired by [GLS22], augmented by a probability mass function on the hyper-edges. Using this model, we first design a novel greedy adaptive algorithm capable of conducting informative tests and dynamically updating the distribution. Performance analysis provides upper bounds on the number of tests required, which depend solely on the entropy of the underlying probability distribution and the average number of infections. We demonstrate that the algorithm recovers or improves upon all previously known results for group testing settings with correlation. Additionally, we provide families of graphs where the algorithm is order-wise optimal and give examples where the algorithm or its analysis is not tight. We then generalize the proposed framework of group testing with general correlation in two directions, namely semi-non-adaptive group testing and noisy group testing. In both settings, we provide novel theoretical bounds on the number of tests required.
This paper characterizes the rate-distortion region for a quadratic Gaussian two-terminal source coding problem with direct and indirect reconstruction. In this problem, the sources consist of an indirect component and two observable components, which are jointly Gaussian. Two separate encoders observe and compress the observable components respectively, and a centralized decoder aims to jointly reconstruct the observable components as well as the indirect component subject to individual quadratic distortion constraints. Under the so-called μ-sum assumption, we provide a complete characterization of the rate region for this problem and prove that it is achieved by the Gaussian Berger-Tung coding scheme. The converse is established by proving that this coding scheme is optimal for every supporting hyperplane, or equivalently, for every weighted sum-rate, of the rate region. Leveraging an analysis of the corresponding Karush-Kuhn-Tucker (KKT) conditions, a novel splitting method is used to decompose the weighted sum-rate optimization problem into some classical problems, including the sum-rate problem, the CEO problem and the one-help-one problem, to demonstrate the optimality. This splitting method reveals the connection between the characterization of the entire rate region and those of its partial bounds.
High temperatures have dramatic negative effects on interconnect performance. In a bus, whenever the state transitions from “0” to “1”, or “1” to “0”, joule heating causes the temperature to rise. The low-power error-correcting cooling (LPECC) codes and constant-power error-correcting cooling (CPECC) codes, introduced in [IEEE Trans. Inf. Theory, 64 (2018), 3062–3085; 66 (2020), 4804–4818], are two coding schemes which can be used to control the peak temperature, the average power consumption of on-chip buses, and error-correction for the transmitted information, simultaneously. In this paper, we firstly show a new upper bound of (n, t, 4, 2)-CPECC codes and construct some families of optimal CPECC codes that attain this upper bound using some combinatorial configurations. Secondly, we present a new upper bound of (n, t, 4, 2)-LPECC codes, and construct some families of optimal LPECC codes whose sizes exceed the bound of the conjecture in [IEEE Trans. Inf. Theory, 71 (2025), 5215–,5225]. Finally, we completely settle the existence of optimal (n, 2, 3, 1)-LPECC codes.
Joint message and state transmission under arbitrarily varying jamming is investigated in this paper. The problem is modeled as the transmission over a channel with random states with a fixed distribution and jamming that varies in an unknown manner. We consider cases in which the channel state is known at the encoder strictly causally and noncausally. For each case, both the average error criterion and the maximal error criterion of the message transmission are adopted. The main results of this paper are lower bounds of the capacity–distortion function of the aforementioned scenarios. The proposed coding schemes are deterministic, and no correlated randomness is needed to achieve reliable communication and estimation. The gap of bounds caused by different error criteria is illustrated by a numerical example.
Divisible codes have been studied for many years, since the divisibility theorem on binary cyclic codes was proved by R. J. McEliece in 1972 and divisible codes were introduced by H. N. Ward in 1981. It is well-known that doubly even (4-divisible) binary codes can be used to construct stabilizer quantum codes. Triorthogonal binary codes and CSS-T pairs of binary codes were introduced to construct CSS-T quantum codes. In this paper, we show that 8-divisible (triply even) binary linear codes are triorthogonal codes and CSS-T pairs of codes can be obtained from 8-divisible codes directly. Infinitely many infinite families of 2r-divisible optimal binary codes are presented, for an arbitrarily large positive integer r. Seventeen doubly even (self-orthogonal) binary cyclic codes and five triply even binary cyclic codes with their minimum weights equal to the minimum weights of best-known linear codes are found. An infinite family of rate around 1/2 explicit doubly even (self-orthogonal) binary cyclic codes with the lower bound ni/log ni on their minimum distances and the lower bound √ni/2 on their dual distances is constructed. This is much stronger than the main results in two previous papers published in IEEE Transactions on Information Theory, 2022 and 2024 by C. Ding, C. Tang, Z. Sun and C. Li. Many CSS-T quantum codes with new parameters are constructed. Many stabilizer quantum codes given in this paper have distances close to the best-known ones documented in Grassl’s database.
Bent functions are maximally nonlinear Boolean functions with critical applications in cryptography, combinatorics, coding theory, and sequence theory. Despite decades of research, their structural characterization remains incomplete, and existing constructions cover only a small fraction of all bent functions. This paper proposes a general framework for constructing bent functions via 2-to-1 mappings. Specifically, we achieve this by modifying the image sets of 2-to-1 mappings and investigating the Boolean functions supported on such sets. This framework enables the reconstruction of classical primary classes (e.g., the Maiorana-McFarland classM, Dillon’s class PSap, and Carlet-Mesnager’s class H) via 2-to-1 mappings with a specific structure. Capitalizing on this unified structure, we utilize oval polynomials to design 2-to-1 mappings and derive a large number of bent functions distinct from the main primary classes.
Cyclic codes are widely used in data storage and communication systems. The construction of cyclic codes with optimal or best-known parameters holds considerable significance in academic research. In this paper, we focus on the constructions of repeated-root cyclic codes with optimal or best-known parameters. Our approach reduces the design of a repeated-root cyclic code to selecting a nested chain of simple-root cyclic codes. We control the degrees of the generator polynomials through cyclotomic-coset size arguments and guarantee their minimum distances via the Bose-Chaudhuri-Hocquenghem (BCH) bound and the Hartmann-Tzeng (HT) bound. For the binary case, by combining BCH codes with nested simple-root cyclic constituents, we construct repeated-root cyclic codes with minimum distances 4, 6, 8, and 10. Several families are proved distance-optimal with respect to the sphere packing bound, while additional families are dual-containing and achieve best-known parameters compared with existing code tables. For the ternary case, analogous constructions yield distance-optimal repeated-root cyclic codes with minimum distances 3 and 4, and further families with minimum distances 5, 6 and 8. Comprehensive parameter tables and representative examples are provided, including instances where the best-known linear codes were previously not known to be cyclic.
Reconfigurable Intelligent Surface (RIS) has emerged as a promising technology to enhance the wireless propagation environment for next-generation wireless communication systems. This paper introduces a new RIS-assisted multiple-antenna coded caching problem. Unlike the existing multi-antenna coded caching models, our considered model incorporates a passive RIS with a limited number of elements aimed at enhancing the multicast gain (i.e., Degrees of Freedom (DoF)). The system consists of a server equipped with multiple antennas and several single-antenna users. The RIS, which functions as a passive and configurable relay, improves communication by selectively ‘erasing’ certain transmission paths between transmit and receive antennas, thereby reducing interference. We first propose a new RIS-assisted interference nulling algorithm to determine the phase-shift coefficients of the RIS. This algorithm achieves faster convergence compared to the existing approach. By strategically nulling certain interference paths in each time slot, the transmission process is divided into multiple interference-free groups. Each group consists of a set of transmit antennas that serve a corresponding set of users without any interference from other groups. The optimal grouping strategy to maximize the DoF is formulated as a combinatorial optimization problem. To efficiently solve this, we design a low-complexity algorithm that identifies the optimal solution and develops a corresponding coded caching scheme to achieve the maximum DoF. Building on the optimal grouping strategy, we introduce a new framework, referred to as RIS-assisted Multiple-Antenna Placement Delivery Array (RMAPDA), to construct the cache placement and delivery phases. Then we propose a general RMAPDA design to achieve the maximum DoF under the optimal grouping strategy. In summary, this paper pioneers the integration of RIS into coded caching systems to boost multicast gain, offering a new direction for the future of wireless communications.