DNA molecules are so small that it might be practical to use their frequency vectors to encode messages. More precisely, a sender can inject M_X copies of the string X = CATCATCAT into a pool and the receiver can recover M_X by sequencing the pool. There are, however, two sources of uncertainty: (a) M_X is usually too big to be counted exactly, but is estimated by sampling. (b) The DNA sequencer could be noisy; it may have difficulty distinguishing CATCATCAT from CATGATCAT. Recently, Tamir, Weinberger, and Guillén i Fàbregas clarified the amount of information the frequency vector can carry under (a). They showed that each string can carry about log_4 R bits, where R is the average number of times each string is read. They also showed that log_4 R bits can be achieved by a low-complexity uncoded scheme under the condition that there are at least √(R) distinct strings. In this paper, we show that a low-complexity coded scheme can achieve the same log_4 R bits unconditionally. We then generalize the scheme to handle sequencing noise, (b), and show that the noise penalizes the total number of bits by log_2 W, together with a linear term due to the use of Fourier transforms in our proof. The former penalty log_2 W is asymptotically the same as that obtained by Gerzon, Shomorony, and Weinberger; our scheme trades a small amount of rate for practical complexity.
The Unique-Machine Precedence Scheduling (UMPS) problem, introduced by [DKRSTZ22], seeks a makespan-minimizing schedule of precedence-constrained jobs when each job has a unique eligible machine. On the one hand, UMPS generalizes job shop scheduling by allowing the precedence graph to be an arbitrary DAG rather than a disjoint union of chains. On the other hand, UMPS admits approximation-preserving reductions to scheduling problems with communication delays, including the job-job delay model [DKRSTZ22] and the job-machine delay model [RSY23]. Despite its central role, the approximability of UMPS has remained poorly understood: even for unit-length jobs, known scheduling techniques do not seem to yield a non-trivial approximation, and the existence of a polylogarithmic approximation was left open by [DKRSTZ22]. On the hardness side, the previous best lower bound for unit-length jobs was only the 5/4 inherited from job shop scheduling [WHHHLSS97]. We prove that unit-length UMPS is NP-hard to approximate within any constant factor. We further show that, assuming NP is not in quasi-polynomial time, unit-length UMPS admits no polynomial-time (log n)^γ-approximation for some constant γ>0. Via the known reductions from UMPS, these lower bounds also transfer to the corresponding unit-length communication-delay scheduling models. Our proof proceeds via a reduction from a hypergraph coloring promise problem. In the yes case, the input hypergraph admits a balanced coloring, while in the no case, the hypergraph has no large independent set. Instantiating this reduction with the hardness of [GL18] gives arbitrary constant-factor inapproximability, while combining the 4-colorable 4-uniform hypergraph coloring hardness of [GHHSV17] with a certain composition operation for hypergraphs yields the polylogarithmic factor inapproximability.
We consider the task of certifying that a random d-dimensional subspace X in R-n is well-spread-every vector x is an element of X satisfies c root n & Vert;x & Vert;2 <=& Vert;x & Vert;1 <=root n & Vert;x & Vert;2. In a seminal work, Barak et al. [Proceedings of the Forty-Fourth Annual ACM Symposium on Theory of Computing, ACM, New York, 2012, pp. 307-326] showed a polynomial-time certification algorithm when d <= O(root n). On the other hand, when d >>root n, the certification task is information-theoretically possible but there is evidence that it is computationally hard [C. Mao and A. S. Wein, Optimal Spectral Recovery of a Planted Vector in a Subspace, preprint, arXiv:2105.15081, 2021; H. Chen and T. d'Orsi, Proc. Mach. Learn, Res, (PMLR), 178 (2022), pp. 1-31], a phenomenon known as the information-computation gap. In this paper, we give subexponential-time certification algorithms in the d >>root n regime. Our algorithm runs in time exp((O)(n(epsilon))) when d <=(O)(n1+epsilon/2), establishing a smooth tradeoff between runtime and the dimension. Our techniques naturally extend to the related planted problem, where the task is to recover a sparse vector planted in a random subspace. Our algorithm achieves the same runtime and dimension tradeoff for this task.
In this paper, we continue the study of robust satisfiability of promise CSPs (PCSPs), initiated in (Brakensiek, Guruswami, Sandeep, STOC 2023 / Discrete Analysis 2025), and obtain the following results: For the PCSP 1-in-3-SAT vs NAE-SAT with negations, we prove that it is hard, under the Unique Games conjecture (UGC), to satisfy 1-Ω(1/log (1/ε)) constraints in a (1-ε)-satisfiable instance. This shows that the exponential loss incurred by the BGS algorithm for the case of Alternating-Threshold polymorphisms is necessary, in contrast to the polynomial loss achievable for Majority polymorphisms. For any Boolean PCSP that admits Majority polymorphisms, we give an algorithm satisfying 1-O() fraction of the weaker constraints when promised the existence of an assignment satisfying 1-ε fraction of the stronger constraints. This significantly generalizes the Charikar–Makarychev–Makarychev algorithm for 2-SAT, and matches the optimal trade-off possible under the UGC. The algorithm also extends, with the loss of an extra log (1/ε) factor, to PCSPs on larger domains with a certain structural condition, which is implied by, e.g., a family of Plurality polymorphisms. We prove that assuming the UGC, robust satisfiability is preserved under the addition of equality constraints. As a consequence, we can extend the rich algebraic techniques for decision/search PCSPs to robust PCSPs. The methods involve the development of a correlated and robust version of the general SDP rounding algorithm for CSPs due to (Brown-Cohen, Raghavendra, ICALP 2016), which might be of independent interest.
We derive the four principal asymptotic rate-distance tradeoffs for binary codes—Plotkin, Elias–Bassalygo, and the two McEliece–Rodemich–Rumsey–Welch (MRRW) bounds—from one theorem, the “pretty good criterion.” If the bit error rate under the pretty good measurement (PGM)—the quantum analog of posterior sampling—of a binary-input output-symmetric classical–quantum (cq) channel lies below δ, then every length-n binary code, linear or nonlinear, of relative distance δ has rate at most the channel's capacity, up to an O(n^-1/2) correction. Rate–distance bounds thereby reduce to a channel design problem, wherein the task is to minimize channel capacity subject to the posterior bit error rate constraint. Via the pretty good criterion, the binary erasure channel (BEC) yields Plotkin, the binary symmetric channel (BSC) yields Elias–Bassalygo, the pure-state channel (PSC) yields the first MRRW bound, and a masked PSC yields the second MRRW bound exactly. This framework is then instantiated with new channels to improve upon the MRRW bounds. Specifically, the mixed-qubit channel (MQC), a mixed-state version of PSC, strictly improves the first MRRW bound at every 0 < δ< 1/2, while the masked mixed-qubit channel (2MQC) strictly improves the second MRRW bound throughout the same interval.
We prove that for every prime power q and every p ∈ (0, 1-1/q), a random 𝔽_q-linear code of rate 1 - h_q(p) - ε is (p, C_p,q/ε)-average-radius list-decodable with probability at least 1 - q^-Ω(n), i.e., for every center y ∈𝔽_q^n, the C_p,q/ε codewords closest to y have average fractional Hamming distance at least p from y. This extends a similar result for (standard) list-decoding due to Guruswami, Håstad, and Kopparty (2010) to the stronger average-radius guarantee, with the same O(1/ε) list size. For average-radius list-decoding, such a result was previously known only for binary linear codes (Guruswami, Li, Mosheiff, Resch, Silas, and Wootters, 2021) and for general (non-linear) random codes over arbitrary alphabets (Elias, 1991).
Quantum locally recoverable codes (QLRCs) have recently gained attention as a framework for achieving efficient quantum storage with local recovery capabilities. Analogous to their classical counterparts, QLRCs allow a lost qudit to be reconstructed using only a small subset of other qudits, thereby reducing the resource and operational overhead in recovery. In this work, we extend the study of QLRCs by considering (r,δ) QLRCs characterized by locality parameter r and local distance δ≥ 2. We present constructions of both random and explicit (r,δ) QLRCs, including explicit families based on the quantum Tamo–Barg construction. We also present an efficient decoding algorithm for these quantum Tamo–Barg codes. Furthermore, we introduce quantum hierarchical locally recoverable codes (QHLRCs), which extend local recovery to multiple hierarchical levels. For any integer h≥ 2, we construct both random and explicit h-level QHLRCs, the latter being h-level quantum Tamo–Barg codes, and establish a Singleton-like bound for these codes using a CSS framework built from dual-containing classical codes. These results advance the theoretical foundations of quantum erasure recovery and contribute to the design of efficient quantum storage architectures.
The parameterized Minimum Monotone Satisfying Assignment (k-MMSA) problem asks whether a monotone Boolean circuit admits a satisfying assignment of Hamming weight at most k. The MMSA hierarchy is defined by allowing a bounded number of alternations between AND and OR gates in the circuit. While the polynomial-time approximability of the MMSA hierarchy has been studied extensively, much less is known in the parameterized setting. In particular, k-MMSA_2 is the well-known k-SetCover problem, whose parameterized inapproximability lies in the polylog(n) regime. In contrast, k-MMSA_4 captures k-MinLabel, for which known lower bounds give poly(n) inapproximability. Sandwiched by k-MMSA_2 and k-MMSA_4, the inapproximability of k-MMSA_3 remained comparatively unexplored. In this paper, we give an FPT-time O(2^k log n)-approximation algorithm for k-MMSA_3, suggesting that in the fixed-parameter regime, the third level of MMSA remains surprisingly close to the second level. Complementing this algorithm, we also give an FPT-time gap-preserving reduction from k-MMSA_3 to k-MMSA_2. Thus, stronger inapproximability for k-MMSA_3 would imply new hardness for k-MMSA_2, potentially offering a route around the current barriers for the latter problem. Revisiting Marx's reduction from k-MMSA_t to gap k-MMSA_t+2, we also show that k-MMSA_4 admits no n^o(1)-factor FPT approximation unless W[2]=FPT, and no n^O(1/k)-factor approximation running in n^o(k) time under ETH. These results separate the parameterized approximability behavior of the third and fourth levels and clarify where stronger inapproximability enters the k-MMSA hierarchy.
A recent work (Korten, Pitassi, and Impagliazzo, FOCS 2025) established an insightful connection between static data structure lower bounds, range avoidance of NC^0 circuits, and the refutation of pseudorandom CSP instances, leading to improvements to some longstanding lower bounds in the cell-probe/bit-probe models. Here, we improve these lower bounds in certain cases via a more streamlined reduction to XOR refutation, coupled with handling the odd-arity case. Our result can be viewed as a complete derandomization of the state-of-the-art semi-random k-XOR refutation analysis (Guruswami, Kothari and Manohar, STOC 2022, Hsieh, Kothari and Mohanty, SODA 2023), which complements the derandomization of the even-arity case obtained by Korten et al. As our main technical statement, we show that for any multi-output constant-depth circuit that substantially stretches its input, its output is very likely far from strings sampled from distributions with sufficient independence, and further this can be efficiently certified. Via suitable shifts in perspectives, this gives applications to cell-probe lower bounds and range avoidance algorithms for 𝖭𝖢^0 circuits.
Given a linear subspace of n × n matrices over 𝔽_2^r that is promised to contain a matrix of rank 1, we prove that it is hard to find a matrix of rank n^o(1/loglog n), assuming NP doesn't have sub-exponential algorithms. In addition to being a basic problem, the hardness of this problem, even for the exact version, drove recent PCP-free inapproximability results for minimum distance and shortest vector problems concerning codes and lattices. The proof combines the concept of superposition soundness introduced by Khot and Saket with moment matrices. To produce a rank-gap of 1 vs. k, the reduction runs in time n^O(log k). We also give another moment-matrix-based construction which runs in time n^O(k) but works for any finite field 𝔽_q.
The Parameterized Inapproximability Hypothesis (PIH) is the analog of the PCP theorem in the world of parameterized complexity. It asserts that no FPT algorithm can distinguish a satisfiable 2CSP instance from one which is only $(1-\varepsilon)$-satisfiable (where the parameter is the number of variables) for some constant $0<\varepsilon<1$. We consider a minimization version of CSPs (Min-CSP), where one may assign $r$ values to each variable, and the goal is to ensure that every constraint is satisfied by some choice among the $r \times r$ pairs of values assigned to its variables (call such a CSP instance $r$-list-satisfiable). We prove the following strong parameterized inapproximability for Min CSP: For every $r \ge 1$, it is W[1]-hard to tell if a 2CSP instance is satisfiable or is not even $r$-list-satisfiable. We refer to this statement as"Baby PIH", following the recently proved Baby PCP Theorem (Barto and Kozik, 2021). Our proof adapts the combinatorial arguments underlying the Baby PCP theorem, overcoming some basic obstacles that arise in the parameterized setting. Furthermore, our reduction runs in time polynomially bounded in both the number of variables and the alphabet size, and thus implies the Baby PCP theorem as well.
Proximity gaps are a property of error correcting codes that arise in the study of Interactive Oracle Proofs (IOPs) and Succinct Non-interactive Arguments of Zero Knowledge (SNARKs). Recent work of Goyal and Guruswami has established near-optimal proximity gaps for many families of codes, including subspace design codes, as well as random ensembles like random linear codes, Reed-Solomon codes with random evaluation points, and Gallager's ensemble of LDPC codes (Goyal Guruswami, 2025). However, the parameters for these latter randomized ensembles are worse than the parameters for subspace design codes, and degrade as the degree ell increases. In this work, we obtain improved proximity gaps for random ensembles of codes, including random linear codes, Reed-Solomon codes with random evaluation points, and Gallager's ensemble. Quantitatively, our results for these random ensembles match the results that Goyal and Guruswami attained for subspace design codes. In fact, our techniques are a black-box transference from subspace design codes: any progress on subspace design codes will automatically lead to analogous progress for these random ensembles. To obtain our results, we extend the Local Coordinate-wise Linear (LCL) property framework developed by Levi, Mosheiff, and Shagrithaya and by Brakensiek, Chen, Dhar, and Zhang to a row-span constrained version (Levi, Mosheiff Shagrithaya, 2025; Brakensiek, Chen, Dhar Zhang, 2025). This allows us to cast curve-decodability – a property that implies proximity gaps – directly as a row-span constrained LCL property, and make use of that machinery. In contrast, because curve-decodability is not obviously a vanilla LCL property, prior work had worked with a proxy property instead, leading to the aforementioned parameter losses.
The subspace design property for additive codes is a higher-dimensional generalization of the minimum distance property. As shown recently by Brakensiek, Chen, Dhar and Zhang, it implies that the code has similar performance as random linear codes with respect to all "local properties". Explicit algebraic codes, such as folded Reed-Solomon and multiplicity codes, are known to have the subspace design property, but they need alphabet sizes that grow as a large polynomial in the block length. Constructing explicit constant-alphabet subspace design codes was subsequently posed as an open question in Brakensiek, Chen, Dhar and Zhang. In this work, we answer their question and give explicit constructions of subspace design codes over constant-sized alphabets, using the expander-based Alon-Edmonds-Luby (AEL) framework. This generalizes the recent work of Jeronimo and Shagrithaya, which showed that such codes share local properties of random linear codes. Our work obtains this consequence in a unified manner via the subspace design property. In addition, our approach yields some improvements in parameters for list-recovery.
The non-redundancy (NRD) of a constraint satisfaction problem (CSP) is a combinatorial quantity closely tied to the behavior of CSPs in various computational models including their sparsification, kernelization, and streaming complexity. A primary open question in the study of non-redundancy is the identification of which CSP predicates have near-linear NRD. Recent works by Carbonnel [CP 2022], Khanna, Putterman and Sudan [STOC 2025], Brakensiek and Guruswami [STOC 2025] and Brakensiek, Guruswami, Jansen, Lagerkvist, and Wahlström [2025] have introduced various forms of gadget reductions between CSPs to relate their non-redundancy. The primary contribution of this work is to recontextualize many of these gadget reductions in a framework which we call hypergraph projections. By studying a quantity we call the shrinking factor of these hypergraph projections, we can more precisely predict when a gadget reduction between predicates can yield a super-linear NRD lower bound, greatly improving on the analysis of previous works. To illustrate the power of our framework, we identify some concrete CSP predicates whose non-redundancy is at the cusp of our understanding and show how our methods give lower bounds that could not have been achieved with these previous methods. We also demonstrate how these gadget reductions can be automatically deduced using SAT solvers, thereby opening up novel computational avenues for discovering further relationships between the non-redundancy of various CSPs.
Given a constraint satisfaction problem (CSP) predicate P ⊆ D^r, the non-redundancy (NRD) of P is maximum-sized instance on n variables such that for every clause of the instance, there is an assignment which satisfies all but that clause. The study of NRD for various CSPs is an active area of research which combines ideas from extremal combinatorics, logic, lattice theory, and other techniques. Complete classifications are known in the cases r=2 and (|D|=2, r=3). In this paper, we give a near-complete classification of the case (|D|=2, r=4). Of the 400 distinct non-trivial Boolean predicates of arity 4, we implement an algorithmic procedure which perfectly classifies 397 of them. Of the remaining three, we solve two by reducing to extremal combinatorics problems – leaving the last one as an open question. Along the way, we identify the first Boolean predicate whose non-redundancy asymptotics are non-polynomial.
The chain length of a set family 𝒮⊆ 2^[m] is the largest ascending sequence of sets in containment order in the union-closure of 𝒮. In this work, we provide a significantly simpler and more optimal characterization of the sparsifiability of set systems in terms of their chain length, improving on the work of Brakensiek and Guruswami [STOC 2025]. Our proof relies on a generalization of Karger's [SODA 1993] famous contraction algorithm and its recent linear algebraic extensions [Khanna-Putterman-Sudan SODA 2024], and our resulting bounds show that, just as VC dimension characterizes the additive sparsifiability of a set system, chain length governs the multiplicative sparsifiability. As a corollary, we obtain improved bounds for weighted CSP sparsification.
In order to communicate a message over a noisy channel, a sender (Alice) uses an error-correcting code to encode her message x into a codeword. The receiver (Bob) decodes it correctly whenever there is at most a small constant fraction of adversarial error in the transmitted codeword. This work investigates the setting where Bob is computationally bounded. Specifically, Bob receives the message as a stream and must process it and write x in order to a write-only tape while using low (say polylogarithmic) space. We show three basic results about this setting, which are informally as follows: (1) There is a stream decodable code of near-quadratic length. (2) There is no stream decodable code of sub-quadratic length. (3) If Bob need only compute a private linear function of the input bits, instead of writing them all to the output tape, there is a stream decodable code of near-linear length.
For quantum error-correcting codes to be realizable, it is important that the qubits subject to the code constraints exhibit some form of limited connectivity. The works of Bravyi & Terhal (BT) and Bravyi, Poulin & Terhal (BPT) established that geometric locality constrains code properties -- for instance $[[n,k,d]]$ quantum codes defined by local checks on the $D$-dimensional lattice must obey $k d^{2/(D-1)} \le O(n)$. Baspin and Krishna studied the more general question of how the connectivity graph associated with a quantum code constrains the code parameters. These trade-offs apply to a richer class of codes compared to the BPT and BT bounds, which only capture geometrically-local codes. We extend and improve this work, establishing a tighter dimension-distance trade-off as a function of the size of separators in the connectivity graph. We also obtain a distance bound that covers all stabilizer codes with a particular separation profile, rather than only LDPC codes.