In this work, we establish the strongest known lower bounds against QAC^0, while allowing its full power of polynomially many ancillae and gates. Our two main results show that: (1) Depth 3 QAC^0 circuits cannot compute PARITY regardless of size, and require at least Ω(exp(√(n))) many gates to compute MAJORITY. (2) Depth 2 circuits cannot approximate high-influence Boolean functions (e.g., PARITY) with non-negligible advantage in depth 2, regardless of size. We present new techniques for simulating certain QAC^0 circuits classically in AC^0 to obtain our depth 3 lower bounds. In these results, we relax the output requirement of the quantum circuit to a single bit (i.e., no restrictions on input preservation/reversible computation), making our depth 2 approximation bound stronger than the previous best bound of Rosenthal (2021). This also enables us to draw natural comparisons with classical AC^0 circuits, which can compute PARITY exactly in depth 2 using exponential size. Our proof techniques further suggest that, for inherently classical decision problems, constant-depth quantum circuits do not necessarily provide more power than their classical counterparts. Our third result shows that depth 2 QAC^0 circuits, regardless of size, cannot exactly synthesize an n-target nekomata state (a state whose synthesis is directly related to the computation of PARITY). This complements the depth 2 exponential size upper bound of Rosenthal (2021) for approximating nekomatas (which is used as a sub-circuit in the only known constant depth PARITY upper bound).
In the problem of quantum state tomography, one is given n copies of an unknown rank-r mixed state ρ∈ℂ^d × d and asked to produce an estimator of ρ. In this work, we present the debiased Keyl's algorithm, the first estimator for full state tomography which is both unbiased and sample-optimal. We derive an explicit formula for the second moment of our estimator, with which we show the following applications. (1) We give a new proof that n = O(rd/ε^2) copies are sufficient to learn a rank-r mixed state to trace distance error ε, which is optimal. (2) We further show that n = O(rd/ε^2) copies are sufficient to learn to error ε in the more challenging Bures distance, which is also optimal. (3) We consider full state tomography when one is only allowed to measure k copies at once. We show that n =O(max(d^3/√(k)ε^2, d^2/ε^2) ) copies suffice to learn in trace distance. This improves on the prior work of Chen et al. and matches their lower bound. (4) For shadow tomography, we show that O(log(m)/ε^2) copies are sufficient to learn m given observables O_1, …, O_m in the "high accuracy regime", when ε = O(1/d), improving on a result of Chen et al. More generally, we show that if tr(O_i^2) ≤ F for all i, then n = O(log(m) ·(min{√(r F)/ε, F^2/3/ε^4/3} + 1/ε^2)) copies suffice, improving on existing work. (5) For quantum metrology, we give a locally unbiased algorithm whose mean squared error matrix is upper bounded by twice the inverse of the quantum Fisher information matrix in the asymptotic limit of large n, which is optimal.
In Online Shadow Tomography, we are given copies of an unknown d-dimensional quantum state ρ, an adversary (adaptively) proposes a sequence of bounded observables A^(1),…,A^(m), and after each A^(t) is given we must estimate (A^(t)ρ) to within ± ε. This is the direct quantum generalization of the classical problem of Adaptive Data Analysis. Prior results for online Shadow Tomography were suboptimal in all three parameters m, d, ε, lagging behind the best known and classical rates , for which there is some evidence of optimality. In this work, we finally close this gap, giving a pair of algorithms matching the classical rates. The bound on the left is the first to achieve o(log^2 m)-dependence together with (log(d)/); moreover, it improves all three exponents even in the Offline Shadow Tomography setting. The bound on the right is known to be optimal among bounds independent of d, and improves the best prior result by a √(m)log m factor. The key to our proof is a new framework for quantifying post-measurement damage, based on the quantum Efron–Stein decomposition.
We design an algorithm for learning the coefficients of an n-qubit constant-local Lindbladian to ε error with O(g d^2 log(n) / ε^2) total evolution time, where g is the single-site energy and d is the (approximate) degree of the interaction graph. Though Lindbladians present new challenges not present in the special case of Hamiltonians, our algorithm achieves the suite of desiderata attained by state-of-the-art Hamiltonian learning algorithms: (1) it uses non-adaptive, ancilla-free randomized Pauli measurement circuits with a time resolution of only Θ(1/g); (2) it works without knowledge of the structure of the unknown Lindbladian; (3) it depends on a smooth form of degree, thereby supporting the learning of quasi-local and power-law Lindbladians. Our algorithm is a simple iterative method, where the objective function consists of Fourier coefficients of the Lindbladian restricted to few-site regions. Its analysis identifies the difficulty unique to open systems, which we call "confusing" terms. For settings where the "confusion" is limited, the performance of the algorithm improves. We demonstrate this for the case of structure learning of Hamiltonians from access to real-time evolution, where we obtain a new algorithm that is significantly simpler than previous work. In addition, using the same iterative method, we design the first efficient algorithm for structure learning Hamiltonians from high-temperature Gibbs states.
We study a generalization of entanglement testing which we call the "hidden cut problem." Taking as input copies of ann-qubit pure state which is product across an unknown bipartition, the goal is to learn precisely where the state is unentangled, i.e. to determine which of the exponentially many possible cuts separates the state. We give a polynomial-time quantum algorithm which can find the cut using O(n/epsilon(2)) many copies of the state, which is optimal up to logarithmic factors. Our algorithm also generalizes to learn the entanglement structure of arbitrary product states. In the special case of Haar-random states, we further show that our algorithm requires circuits of only constant depth. To develop our algorithm, we introduce a state generalization of the hidden subgroup problem (StateHSP) which might be of independent interest, in which one is given a quantum state invariant under an unknown subgroup action, with the goal of learning the hidden symmetry subgroup. We show how the hidden cut problem can be formulated as a StateHSP with a carefully chosen Abelian group action. We then prove that Fourier sampling on the hidden cut state produces similar outcomes as a variant of the well-known Simon's problem, allowing us to find the hidden cut efficiently. Therefore, our algorithm can be interpreted as an extension of Simon's algorithm to entanglement testing. We discuss possible applications of StateHSP and hidden cut problems to cryptography and pseudorandomness.
We show that n = Ω(rd/ε^2) copies are necessary to learn a rank r mixed state ρ∈ℂ^d × d up to error ε in trace distance. This matches the upper bound of n = O(rd/ε^2) from prior work, and therefore settles the sample complexity of mixed state tomography. We prove this lower bound by studying a special case of full state tomography that we refer to as projector tomography, in which ρ is promised to be of the form ρ= P/r, where P ∈ℂ^d × d is a rank r projector. A key technical ingredient in our proof, which may be of independent interest, is a reduction which converts any algorithm for projector tomography which learns to error ε in trace distance to an algorithm which learns to error O(ε) in the more stringent Bures distance.
A longstanding belief in quantum tomography is that estimating a mixed state is far harder than estimating a pure state. This is borne out in the mathematics, where mixed state algorithms have always required more sophisticated techniques to design and analyze than pure state algorithms. We present a new approach to tomography demonstrating that, contrary to this belief, state-of-the-art mixed state tomography follows easily and naturally from pure state algorithms. We analyze the following strategy: given n copies of an unknown state ρ, convert them into copies of a purification |ρ⟩; run a pure state tomography algorithm to produce an estimate of |ρ⟩; and output the resulting estimate of ρ. The purification subroutine was recently discovered via the "acorn trick" of Tang, Wright, and Zhandry. With this strategy, we obtain the first tomography algorithm which is sample-optimal in all parameters. For a rank-r d-dimensional state, it uses n = O((rd + log(1/δ))/ε) samples to output an estimate which is ε-close in fidelity with probability at least 1-δ. This algorithm also uses poly(n) gates, making it the first gate-efficient tomography algorithm which is sample-optimal even in terms of the dimension d alone. Moreover, with this method we recover essentially all results on mixed state tomography, including its applications to tomography with limited entanglement, classical shadows, and quantum metrology. Our proofs are simple, closing the gap in conceptual difficulty between mixed and pure tomography. Our results also clarify the role of entangled measurement in mixed state tomography: the only step of the algorithm which requires entanglement across copies is the purification step, suggesting that, for tomography, the reason entanglement is useful is for consistent purification.
We prove that the generic quantum speedups for brute-force search and counting only hold when the process we apply them to can be efficiently inverted. The algorithms speeding up these problems, amplitude amplification and amplitude estimation, assume the ability to apply a state preparation unitary U and its inverse U^†; we give problem instances based on trace estimation where no algorithm which uses only U beats the naive, quadratically slower approach. Our proof of this is simple and goes through the compressed oracle method introduced by Zhandry. Since these two subroutines are responsible for the ubiquity of the quadratic "Grover" speedup in quantum algorithms, our result explains why such speedups are far harder to come by in the settings of quantum learning, metrology, and sensing. In these settings, U models the evolution of an experimental system, so implementing U^† can be much harder – tantamount to reversing time within the system. Our result suggests a dichotomy: without inverse access, quantum speedups are scarce; with it, quantum speedups abound.
High-dimensional data are ubiquitous, with examples ranging from natural images to scientific datasets, and often reside near low-dimensional manifolds. Leveraging this geometric structure is vital for downstream tasks, including signal denoising, reconstruction, and generation. However, in practice, the manifold is typically unknown and only noisy samples are available. A fundamental approach to uncovering the manifold structure is local averaging, which is a cornerstone of state-of-the-art provable methods for manifold fitting and denoising. However, to the best of our knowledge, there are no works that rigorously analyze the accuracy of local averaging in a manifold setting in high-noise regimes. In this work, we provide theoretical analyses of a two-round mini-batch local averaging method applied to noisy samples drawn from a d-dimensional manifold ℳ⊂ℝ^D, under a relatively high-noise regime where the noise size is comparable to the reach τ. We show that with high probability, the averaged point 𝐪̂ achieves the bound d(𝐪̂, ℳ) ≤ σ√(d(1+κdiam(ℳ)/log(D))), where σ, diam(ℳ),κ denote the standard deviation of the Gaussian noise, manifold's diameter and a bound on its extrinsic curvature, respectively. This is the first analysis of local averaging accuracy over the manifold in the relatively high noise regime where σ√(D)≈ τ. The proposed method can serve as a preprocessing step for a wide range of provable methods designed for lower-noise regimes. Additionally, our framework can provide a theoretical foundation for a broad spectrum of denoising and dimensionality reduction methods that rely on local averaging techniques.
Many quantum algorithms, to compute some property of a unitary U, require access not just to U, but to cU, the unitary with a control qubit. We show that having access to cU does not help for a large class of quantum problems. For a quantum circuit which uses cU and cU^† and outputs |ψ(U)⟩, we show how to “decontrol” the circuit into one which uses only U and U^† and outputs |ψ(φU)⟩ for a uniformly random phase φ, with a small amount of time and space overhead. When we only care about the output state up to a global phase on U, then the decontrolled circuit suffices. Stated differently, cU is only helpful because it contains global phase information about U. A version of our procedure is described in an appendix of Sheridan, Maslov, and Mosca [SMM09]. Our goal with this work is to popularize this result by generalizing it and investigating its implications, in order to counter negative results in the literature which might lead one to believe that decontrolling is not possible. As an application, we give a simple proof for the existence of unitary ensembles which are pseudorandom under access to U, U^†, cU, and cU^†.
Learned denoisers play a fundamental role in various signal generation (e.g., diffusion models) and reconstruction (e.g., compressed sensing) architectures, whose success derives from their ability to leverage low-dimensional structure in data. Existing denoising methods, however, either rely on local approximations that require a linear scan of the entire dataset or treat denoising as generic function approximation problems, sacrificing efficiency and interpretability. We consider the problem of efficiently denoising a new noisy data point sampled from an unknown d-dimensional manifold M is an element of R-D, using only noisy samples. This work proposes a framework for test-time efficient manifold denoising, by framing the concept of "learning-to-denoise" as "learning-to-optimize". We have two technical innovations: (i) online learning methods which learn to optimize over the manifold of clean signals using only noisy data, effectively "growing" an optimizer one sample at a time. (ii) mixed-order methods which guarantee that the learned optimizers achieve global optimality, ensuring both efficiency and near-optimal denoising performance. We corroborate these claims with theoretical analyses of both the complexity and denoising performance of mixed-order traversal. Our experiments on scientific manifolds demonstrate significantly improved complexity-performance tradeoffs compared to nearest neighbor search, which underpins existing provable denoising approaches based on exhaustive search.
We give a natural problem over input quantum oracles U which cannot be solved with exponentially many black-box queries to U and U^†, but which can be solved with constant many queries to U and U^*, or U and U^T. We also demonstrate a quantum commitment scheme that is secure against adversaries that query only U and U^†, but is insecure if the adversary can query U^*. These results show that conjugate and transpose queries do give more power to quantum algorithms, lending credence to the idea put forth by Zhandry that cryptographic primitives should prove security against these forms of queries. Our key lemma is that any circuit using q forward and inverse queries to a state preparation unitary for a state σ can be simulated to ε error with n = 𝒪(q^2/ε) copies of σ. Consequently, for decision tasks, algorithms using (forward and inverse) state preparation queries only ever perform quadratically better than sample access. These results follow from straightforward combinations of existing techniques; our contribution is to state their consequences in their strongest, most counter-intuitive form. In doing so, we identify a motif where generically strengthening a quantum resource can be possible if the output is allowed to be random, bypassing no-go theorems for deterministic algorithms. We call this the acorn trick.
Inferring the exact parameters of a neural network with only query access is an NP-Hard problem, with few practical existing algorithms. Solutions would have major implications for security, verification, interpretability, and understanding biological networks. The key challenges are the massive parameter space, and complex non-linear relationships between neurons. We resolve these challenges using two insights. First, we observe that almost all networks used in practice are produced by random initialization and first order optimization, an inductive bias that drastically reduces the practical parameter space. Second, we present a novel query generation algorithm that produces maximally informative samples, letting us untangle the non-linear relationships efficiently. We demonstrate reconstruction of a hidden network containing over 1.5 million parameters, and of one 7 layers deep, the largest and deepest reconstructions to date, with max parameter difference less than 0.0001, and illustrate robustness and scalability across a variety of architectures, datasets, and training procedures.
Attention mechanisms play a crucial role in state-of-the-art vision architectures, enabling them to rapidly identify relationships between distant image patches. Conventional attention mechanisms do not incorporate other structural properties of images, such as invariance to geometric transformations, instead learning these properties from data. In this paper, we introduce a novel mechanism, Invariant Attention, which, like standard attention, captures image similarity, but with the additional guarantee of being agnostic to geometric transformations. We provide theoretical assurance and empirical verification that invariant attention is far more successful than standard kernel attention on multi-class, transformed vision data, and illustrate its potential to correctly cluster transformed data with intra-class variation.
Background: The association between fine particulate matter (PM2.5) and cardiovascular outcomes is well established. To evaluate whether source-specific PM2.5 is differentially associated with cardiovascular disease in New York City (NYC), we identified PM2.5 sources and examined the association between source-specific PM2.5 exposure and risk of hospitalization for myocardial infarction (MI). Methods: We adapted principal component pursuit (PCP), a dimensionality-reduction technique previously used in computer vision, as a novel pattern recognition method for environmental mixtures to apportion speciated PM2.5 to its sources. We used data from the NY Department of Health Statewide Planning and Research Cooperative System of daily city-wide counts of MI admissions (2007–2015). We examined associations between same-day, lag 1, and lag 2 source-specific PM2.5 exposure and MI admissions in a time-series analysis, using a quasi-Poisson regression model adjusting for potential confounders. Results: We identified four sources of PM2.5 pollution: crustal, salt, traffic, and regional and detected three single-species factors: cadmium, chromium, and barium. In adjusted models, we observed a 0.40% (95% confidence interval [CI]: –0.21, 1.01%) increase in MI admission rates per 1 μg/m3 increase in traffic PM2.5, a 0.44% (95% CI: –0.04, 0.93%) increase per 1 μg/m3 increase in crustal PM2.5, and a 1.34% (95% CI: –0.46, 3.17%) increase per 1 μg/m3 increase in chromium-related PM2.5, on average. Conclusions: In our NYC study, we identified traffic, crustal dust, and chromium PM2.5 as potentially relevant sources for cardiovascular disease. We also demonstrated the potential utility of PCP as a pattern recognition method for environmental mixtures.
Gravitational wave astronomy is a vibrant field that leverages both classic and modern data processing techniques for the understanding of the universe. Various approaches have been proposed for improving the efficiency of the detection scheme, with hierarchical matched filtering being an important strategy. Meanwhile, deep learning methods have recently demonstrated both consistency with matched filtering methods and remarkable statistical performance. In this work, we propose Hierarchical Detection Network (HDN), a novel approach to efficient detection that combines ideas from hierarchical matching and deep learning. The network is trained using a novel loss function, which encodes simultaneously the goals of statistical accuracy and efficiency. We discuss the source of complexity reduction of the proposed model, and describe a general recipe for initialization with each layer specializing in different regions. We demonstrate the performance of HDN with experiments using open LIGO data and synthetic injections, and observe with two-layer models a $79\%$ efficiency gain compared with matched filtering at an equal error rate of $0.2\%$. Furthermore, we show how training a three-layer HDN initialized using two-layer model can further boost both accuracy and efficiency, highlighting the power of multiple simple layers in efficient detection.
Gravitational wave science is a pioneering field with rapidly evolving data analysis methodology currently assimilating and inventing deep learning techniques. The bulk of the sophisticated flagship searches of the field rely on the time-tested matched filtering principle within their core. In this paper, we make a key observation on the relationship between the emerging deep learning and the traditional techniques: matched filtering is formally equivalent to a particular neural network. This means that a neural network can be constructed analytically to exactly implement matched filtering, and can be further trained on data or boosted with additional complexity for improved performance. Moreover, we show that the proposed neural network architecture can outperform matched filtering, both with or without knowledge of a prior on the parameter distribution. When a prior is given, the proposed neural network can approach the statistically optimal performance. We also propose and investigate two different neural network architectures MNet-Shallow and MNet-Deep, both of which implement matched filtering at initialization and can be trained on data. MNet-Shallow has simpler structure, while MNet-Deep is more flexible and can deal with a wider range of distributions. Our theoretical findings are corroborated by experiments using real LIGO data and synthetic injections, where our proposed methods significantly outperform matched filtering at false positive rates above $5\times 10^{-3}\%$. The fundamental equivalence between matched filtering and neural networks allows us to define a "complexity standard candle" to characterize the relative complexity of the different approaches to gravitational wave signal searches in a common framework. Finally, our results suggest new perspectives on the role of deep learning in gravitational wave detection.
BACKGROUND AND AIM: It is often of interest to identify sources of environmental exposures from imperfect data that suffer from block missingness, in which observations for multiple pollutants are missing across large portions of the study period. METHODS: We adapted Principal Component Pursuit (PCP), a robust dimensionality reduction algorithm, to identify air pollution sources from incomplete data. PCP decomposes the pollutant matrix into consistent patterns while separately isolating unique or outlying pollution events. PCP handles structural missingness by reconstructing missing blocks using the information from observed blocks of the exposure matrix. We applied PCP to apportion 26 PM2.5 constituents to their sources in New York City, using data from three monitors (2001 – 2020). Two constituents, elemental (EC) and organic carbon (OC), were missing all measurements from 2001 to 2007, comprising 2.6% of the overall pollution matrix. RESULTS: PCP reconstructed six years of EC and OC data consistent with existing literature, and identified five sources of PM2.5 pollution: crustal dust, road dust, salt, secondary/regional sulfate, and traffic, as well as three single-constituent components: arsenic, barium, and chromium. Traffic contributed most to total PM2.5 concentrations (20.2%) across the study period, peaking during the winter months and on weekdays. PCP also identified interpretable outlying pollution events, most notably spikes in potassium ion concentrations around each fourth of July. CONCLUSIONS: PCP can serve as a useful and robust technique capable of handling data suffering from block missingness to identify exposure patterns and sparse events amenable to public health messaging and research. KEYWORDS: Mixtures, Air Pollution, Source Apportionment, Missing Data, Fine Particles
Connecting theory with practice, this systematic and rigorous introduction covers the fundamental principles, algorithms and applications of key mathematical models for high-dimensional data analysis. Comprehensive in its approach, it provides unified coverage of many different low-dimensional models and analytical techniques, including sparse and low-rank models, and both convex and non-convex formulations. Readers will learn how to develop efficient and scalable algorithms for solving real-world problems, supported by numerous examples and exercises throughout, and how to use the computational tools learnt in several application contexts. Applications presented include scientific imaging, communication, face recognition, 3D vision, and deep networks for classification. With code available online, this is an ideal textbook for senior and graduate students in computer science, data science, and electrical engineering, as well as for those taking courses on sparsity, low-dimensional structures, and high-dimensional data. Foreword by Emmanuel Candès.
Donald Goldfarb合作论文数Department of Industrial Engineering and Operations Research, Columbia University3