Quantum low-density parity-check (qLDPC) codes offer a promising route to scalable fault-tolerant quantum computing due to their substantially reduced footprint. However, these gains can be diluted at utility scale if we cannot also realize space-time efficient logical operations for relevant quantum applications. We present RASCqL, a reaction-time-limited architecture for space-time efficient complex-instruction-set quantum computation with qLDPC logic. RASCqL supports key algorithmic subroutines such as quantum arithmetic and state preparation directly within co-designed qLDPC codes, achieving 2× to 7× reductions in qubit footprint while maintaining space-time volume comparable to state-of-the-art transversal surface-code architectures. Unlike prior approaches that aim for versatile logical instruction sets for arbitrary circuits, RASCqL adopts an application-tailored code modification that embeds specific complex Clifford transformations useful for common subroutines as virtually implementable operations arising from code automorphisms. RASCqL further leverages parallel physical operations available in reconfigurable neutral-atom arrays to enable fast QEC cycles and high-fidelity transversal operations. At the cost of increased design complexity and specialization, RASCqL can improve end-to-end resource estimates for applications such as factoring and quantum chemistry simulation in both footprint and space-time volume under realistic physical error rates of approximately 2×10^-3 to 5×10^-4, without requiring additional hardware capabilities. These results demonstrate that qLDPC codes can serve as complex quantum logic units for useful quantum algorithms, extending their practical utility in fault-tolerant quantum computing architectures.
Academic quantum computing platforms often face unique challenges in executing quantum workloads due to fragmented software environments and limited engineering support. Unlike commercial ecosystems, academic devices typically evolve without full-stack integration in mind, making it difficult to run complex applications—such as variational quantum algorithms (VQA)—reliably and efficiently. Issues such as incompatible software layers and lack of automated job management significantly increase the overhead of theory-experiment collaboration. To address these challenges, we develop a modular, end-to-end workflow that decouples application-layer code from low-level hardware control, automates circuit submission and result collection, and supports fine-grained circuit-level job scheduling and recovery. The architecture employs a dual-end application programming interface (API) design, enabling robust operation across unstable or resource-constrained hardware backends. For practical use, the framework is lightweight and user-friendly, allowing rapid prototyping of full-stack workflows using basic Python tools. We validate this workflow on a high-fidelity trapped-ion quantum computer by demonstrating a variational quantum eigensolver (VQE) experiment with a classically bootstrapped ansatz initialization technique. The system successfully executed over 60,000 circuits across multiple molecular test cases with minimal human intervention, highlighting the framework’s effectiveness in enabling reproducible, resilient quantum experimentation in academic settings.
As quantum machines have scaled up in their number of qubits, significant research has turned towards increasing their fidelity with quantum error correction codes. Although promising results have been shown with the surface code, which only requires near-neighbor connections between qubits, the high qubit overhead of such local codes promises to be problematic. Consequently, recent work has explored non-local quantum LDPC (qLDPC) codes, which have good asymptotic encoding rates. Despite theoretical progress, hardware implementations of these codes have been a longstanding challenge. At the experimental level, demonstrations of movement based communication on atom arrays suggest this is a powerful new primitive to achieve non-local connectivity. Leveraging this, we present a protocol for implementing non-local qLDPC codes in hardware. Our protocol, qSIEVE, is a co-design of such codes with movement in atom arrays. qSIEVE defines a restricted family of qLDPC codes that can be implemented efficiently with systolic movement. We then quantify the utility of qSIEVE in the context of a complete fault tolerant architecture. We compare the cost of implementing benchmark programs in a standard, surface code only architecture and a mixed architecture where data is stored in qLDPC memory with qSIEVE and loaded to surface codes for computation. CCS Concepts: center dot Computer systems organization -> Quantum computing;
Practical applications of quantum computing depend on fault-tolerant devices that employ error correction. A promising quantum error-correcting code for large-scale quantum computing is the surface code. For this code, Fault-Tolerant Quantum Computing (FTQC) can be performed via lattice surgery, i.e. merging and splitting of encoded qubit patches on a 2D grid. Lattice surgery operations result in space-time patterns of activity that are defined in this work as access traces. This work demonstrates that the access traces reveal when, where, and how logical qubits interact. Leveraging this formulation, this work further introduces TraceQ, a tracebased reconstruction framework that is able to reconstruct the quantum circuit dataflow just by observing the patch activity at each trace entry. The framework is supported by heuristics for handling inherent ambiguity in the traces, and demonstrates its effectiveness on a range of synthetic fault-tolerant quantum benchmarks. The access traces can have applications in a wide range of scenarios, enabling analysis and profiling of execution of quantum programs and the hardware they run on. As one example use of TraceQ, this work investigates whether such traces can act as a side channel through which an observer can recover the circuit's structure and identify known subroutines in a larger program or even whole programs. The findings show that indeed the minimal access traces can be used to recover subroutines or even whole quantum programs with very high accuracy. Only a single trace per program execution is needed and the processing can be done fully offline. Along with the custom heuristics, advanced subgraph matching algorithms used in this work enable a high rate of locating the subroutines while executing in minimal time.
Quantum computers have improved in size and quality in recent years, enabling the execution of complex circuits. However, for most researchers, access to compute time is limited. This necessitates the development of simulators that mimic noisy quantum hardware accurately and scalably. The ideal way to simulate noisy systems is via Density Matrix Simulation (DMS). However, its high memory footprint limits its scalability. Consequently, noisy simulations are performed in two steps: (a) sampling multiple circuits with fixed noisy gates from the stochastic noise channels, (b) performing the State Vector Simulations (SVS) of these circuits and averaging their output to obtain the effective noisy simulation result. This often leads to a substantial increase in compute overhead, slowing down the simulation. Existing methods solve this problem by caching critical intermediate results in memory and reusing them. However, when a simulation task is both compute and memory-intensive, we need to eliminate computational overheads without incurring extra memory overheads. To enable fast simulation in the compute and memory bound regime, we propose TUSQ - Tracking, Uncomputation, and Sampling for Noisy Quantum Simulation. TUSQ is composed of two modules: the Error Characterization Module (ECM), and Depth First Tree Traversal (DFTT). The ECM characterizes errors so that the simulator can eliminate redundant circuit instances (via ER Tallying and ER Commutation), followed by importance sampling (in the Pruning stage), significantly reducing the number of circuits to be simulated relative to the baseline strategy of simulating all circuits. This is followed by DFTT, which computes the statevectors for these sampled circuits efficiently by taking advantage of circuit similarity, representing similar circuits in a tree and using computation and uncomputation to traverse the tree efficiently. TUSQ is evaluated for a total of 198 benchmarks, executed for 1 million shots and reports an average speedup of $59.06 \times$ and $13.38 \times$ over Qiskit and CUDA-Q, with a maximum speedup of $7878.03 \times$ and $439.38 \times$ respectively. We also compare TUSQ against TQSim in the time and memory critical regime. We observe an average and maximum speedup of $39.32 \times$ and 3134.31×, respectively.
Fault-Tolerant Quantum Computing (FTQC) relies on Quantum Error Correction (QEC) codes to reach error rates necessary for large scale quantum applications. At a physical level, QEC codes perform parity checks on data qubits, producing syndrome information, through Syndrome Measurement (SM) circuits. These circuits define a code's logical error rate and must be run repeatedly throughout the entire program. The performance of SM circuits is therefore critical to the success of a FTQC system. While ultimately implemented as physical circuits, SM circuits have challenges that are not addressed by existing circuit optimization tools. Importantly, inside SM circuits themselves errors are expected to occur, and how errors propagate through SM circuits directly impacts which errors are detectable and correctable, defining the code's logical error rate. This is not modeled in NISQ-era tools, which instead optimize for targets such as gate depth or gate count to mitigate the chance that any error occurs. This gap leaves key questions unanswered about the expected real-world effectiveness of QEC codes. In this work we address this gap and present PropHunt, an automated tool for optimizing SM circuits for CSS codes. We evaluate PropHunt on a suite of relevant QEC codes and demonstrate PropHunt's ability to iteratively improve performance and recover existing hand-designed circuits automatically. We also propose a near-term QEC application, Hook-ZNE, which leverages PropHunt's fine-grained control over logical error rate to improve Zero-Noise Extrapolation (ZNE), a promising error mitigation strategy.
Programming a quantum device describes the usage of quantum logic gates, agnostic of hardware specifics, to perform a sequence of operations with (typically) a computing or sensing task in mind. Such programs have been executed on gate-based quantum computers, which despite their noisy character, have shown the ability to optimize metrological functions, for example in the generation of spin squeezing and optimization of quantum Fisher information for signals manifesting as spin rotations in a quantum register. However, the qubits of these programmable quantum sensors are tightly spatially confined and therefore suboptimal for enclosing the kinds of large spacetime areas required for performing inertial sensing. In this work, we derive a set of quantum logic gates for a cold atom optical lattice interferometer that manipulates the momentum of atoms. Here, the operations are framed in terms of single qubit operations and mappings between qubit subspaces with internal levels given by the Bloch (crystal) eigenstates of the lattice. We describe how the quantum optimal control method of direct collocation is well suited for obtaining modulation waveforms of an optical lattice to achieve these operations in existing experimental setups.
Presence of harmful noise is inevitable in quantum sensing systems, requiring careful allocation of resources to optimize sensing performance in practical scenarios. We advocate a simple but effective strategy to improve quantum sensing performance in the presence of noise. Given a fixed number of quantum sensors, we partition the preparation of GHZ states by preparing smaller, independent sub-ensembles of GHZ states instead of a GHZ state across all sensors. We perform extensive analytical studies of the phase estimation performance when using partitioned GHZ states under realistic noise - including state preparation error, particle loss during parameter encoding, and sensor dephasing during parameter encoding. We derive simple, closed-form expressions that quantify the optimal number of sub-ensembles for partitioned GHZ states. The results offer quantitative insights into the sensing performance impact of different noise sources and reinforce the importance of resource allocation optimization in realistic quantum applications.
We report on the fault-tolerant operation of logical qubits on a neutral atom quantum computer, with logical performance surpassing physical performance for multiple circuits including Bell state preparation (12x error reduction), random circuits (15x), and a prototype Anderson Impurity Model ground state solver for materials science applications (up to 6x, non-fault-tolerantly). The logical qubits are implemented via the [[4, 2, 2]] code (C4). Our work constitutes the first complete realization of the benchmarking protocol proposed by Gottesman 2016 demonstrating results consistent with fault tolerance. In light of recent advances on applying concatenated C4/C6 detection codes to achieve error correction with high code rates and thresholds, our work can be regarded as a building block towards a practical scheme for fault tolerant quantum computation. Our demonstration of a materials science application with logical qubits particularly demonstrates the immediate value of these techniques on current experiments.
Quantum computers have grown in size and qubit quality in recent years, enabling the execution of complex quantum circuits. However, for most researchers, access to compute time on quantum hardware is limited. This necessitates the need to build simulators that mimic the execution of quantum circuits on noisy quantum hardware accurately and scalably. In this work, we propose TUSQ - Tracking, Uncomputation, and Sampling for Noisy Quantum Simulation. To represent the stochastic noisy channels accurately, we average the output of multiple quantum circuits with fixed noisy gates sampled from the channels. However, this leads to a substantial increase in circuit overhead, which slows down the simulation. To eliminate this overhead, TUSQ uses two modules: the Error Characterization Module (ECM), and the Tree-based Execution Module (TEM). The ECM tracks the number of unique circuit executions needed to accurately represent the noise. That is, if initially we needed n_1 circuit executions, ECM reduces that number to n_2 by eliminating redundancies so that n_2 < n_1. This is followed by the TEM, which reuses computation across these n_2 circuits. This computational reuse is facilitated by representing all n_2 circuits as a tree. We sample the significant leaf nodes of this tree and prune the remaining ones. We traverse this tree using depth-first search. We use uncomputation to perform rollback-recovery at several stages which reduces simulation time. We evaluate TUSQ for a total of 186 benchmarks and report an average speedup of 52.5× and 12.53× over Qiskit and CUDA-Q, which goes up to 7878.03× and 439.38× respectively. For larger benchmarks (more than than 15 qubits), the average speedup is 55.42× and 23.03× over Qiskit and CUDA-Q respectively
As quantum machines have scaled up in their number of qubits, significant research has turned towards increasing their fidelity with quantum error correction codes. Although promising results have been shown with the surface code, which only requires near-neighbor connections between qubits, the high qubit overhead of such local codes promises to be problematic. Consequently, recent work has explored non-local quantum LDPC (qLDPC) codes, which have good asymptotic encoding rates. Despite theoretical progress, hardware implementations of these codes has been a longstanding challenge. At the experimental level, demonstrations of movement based communication on atom arrays suggest this is a powerful new primitive to achieve non-local connectivity. Leveraging this, we present a protocol for implementing non-local qLDPC codes in hardware. Our protocol, qSIEVE, is a co-design of such codes with movement in atom arrays. qSIEVE defines a restricted family of qLDPC codes that can be implemented efficiently with systolic movement. We then quantify the utility of qSIEVE in the context of a complete fault tolerant architecture. We compare the cost of implementing benchmark programs in a standard, surface code only architecture and a mixed architecture where data is stored in qLDPC memory with qSIEVE and loaded to surface codes for computation.
Quantum computing roadmaps predict the availability of 10,000-qubit devices within the next 3-5 years. With projected two-qubit error rates of 0.1%, these systems will enable certain operations under quantum error correction (QEC) using lightweight codes, offering significantly improved fidelities compared to the NISQ era. However, the high qubit cost of QEC codes like the surface code (especially at near-threshold physical error rates) limits the error correction capabilities of these devices. In this emerging era of Early Fault Tolerance (EFT), it will be essential to use QEC resources efficiently and focus on applications that derive the greatest benefit. In this work, we investigate the implementation of Variational Quantum Algorithms in the EFT regime (EFT-VQA). We explore the ideas of partial quantum error correction (pQEC), a strategy that error-corrects Clifford operations while performing R-Z (theta) rotations via magic state injection instead of the more expensive T-state distillation, and adapt it to VQAs. Our results show that pQEC can improve VQA fidelities by 9.27x over standard approaches. Furthermore, we propose architectural optimizations that reduce circuit latency by similar to 2x, and achieve qubit packing efficiency of 66% in the EFT regime. The source code can be accessed here https://github.com/siddharthdangwal/EFT-VQA.
A core challenge for superconducting quantum computers is to scale up the number of qubits in each processor without increasing noise or cross-talk. Distributed quantum computing across small qubit arrays, known as chiplets, can address these challenges in a scalable manner. We propose a chiplet architecture over microwave links with potential to exceed monolithic performance on near-term hardware. Our methods of modeling and evaluating the chiplet architecture bridge the physical and network layers in these processors. We find evidence that distributing computation across chiplets may reduce the overall error rates associated with moving data across the device, despite higher error figures for transfers across links. Preliminary analyses suggest that latency is not substantially impacted, and that at least some applications and architectures may avoid bottlenecks around chiplet boundaries. In the long-term, short-range networks may underlie quantum computers just as local area networks underlie classical datacenters and supercomputers today.
We present and open source Extraferm, a quantum circuit simulator tailored to chemistry applications. More specifically, our simulator can compute the Born-rule probabilities of samples obtained from circuits containing particle number-conserving matchgates and controlled-phase gates. We support both approximate and exact calculation of probabilities, and for approximate probability calculation, our simulator's runtime is exponential only in the magnitudes of the circuit's controlled-phase gate angles. This makes our simulator useful for simulating certain systems that are beyond the reach of conventional state vector methods. We demonstrate our simulator's utility by simulating the local cluster unitary Jastrow (LUCJ) ansatz and integrating it with sample-based quantum diagonalization (SQD) to improve the accuracy of molecular ground-state energy estimates with negligible computational overhead. More generally, we highlight a regime in which our simulator achieves substantially superior latency scaling and exponentially superior memory scaling over a tensor network simulator and a state vector simulator. As an efficient and flexible tool for simulating quantum chemistry circuits, our simulator enables new opportunities for enhancing near-term quantum algorithms in chemistry and related domains.
Presence of harmful noise is inevitable in entanglement-enhanced sensing systems, requiring careful allocation of resources to optimize sensing performance in practical scenarios. We advocate a simple but effective strategy to improve sensing performance in the presence of noise. Given a fixed number of quantum sensors, we partition the preparation of GHZ states by preparing smaller, independent sub-ensembles of GHZ states instead of a GHZ state across all sensors. We perform extensive analytical studies of the phase estimation performance when using partitioned GHZ states under realistic noise – including state preparation error, particle loss during parameter encoding, and sensor dephasing during parameter encoding. We derive simple, closed-form expressions that quantify the optimal number of sub-ensembles for partitioned GHZ states. We also examine the explicit noisy quantum sensing dynamics under dephasing and loss, where we demonstrate the advantage from partitioning for maximal QFI, short-time QFI increase, and the sensing performance in the sequential scheme. The results offer quantitative insights into the sensing performance impact of different noise sources and reinforce the importance of resource allocation optimization in realistic quantum applications.
Practical applications of quantum computing depend on fault-tolerant devices that employ error correction. A promising quantum error-correcting code for large-scale quantum computing is the surface code. For this code, Fault-Tolerant Quantum Computing (FTQC) can be performed via lattice surgery, i.e. merging and splitting of encoded qubit patches on a 2D grid. Lattice surgery operations result in space-time patterns of activity that are defined in this work as access traces. This work demonstrates that the access traces reveal when, where, and how logical qubits interact. Leveraging this formulation, this work further introduces TraceQ, a trace-based reconstruction framework that is able to reconstruct the quantum circuit dataflow just by observing the patch activity at each trace entry. The framework is supported by heuristics for handling inherent ambiguity in the traces, and demonstrates its effectiveness on a range of synthetic fault-tolerant quantum benchmarks. The access traces can have applications in a wide range of scenarios, enabling analysis and profiling of execution of quantum programs and the hardware they run on. As one example use of TraceQ, this work investigates whether such traces can act as a side channel through which an observer can recover the circuit's structure and identify known subroutines in a larger program or even whole programs. The findings show that indeed the minimal access traces can be used to recover subroutines or even whole quantum programs with very high accuracy. Only a single trace per program execution is needed and the processing can be done fully offline. Along with the custom heuristics, advanced subgraph matching algorithms used in this work enable a high rate of locating the subroutines while executing in minimal time.
We demonstrate a logical neutral atom architecture that integrates atom motion with in-place entanglement to achieve lower overheads than entangling-zone approaches. Using a 114-qubit device, we perform three proof-of-principle logical-qubit experiments. First, we implement a pre-compiled, non-scalable variant of Shor's algorithm, observing improved logical-over-physical performance, including with loss correction and leakage detection, achieving up to a 2x reduction in TVD. Second, we construct constant-depth logical CX ladders; on current hardware these execute with serial entangling operations, yet still yield 2-4x lower error for 8 and 12 logical qubits. Third, we prepare the [[16,4,4]] code and perform single-round decoding with post-processed error correction, achieving 8x improvement on logical vs physical. These results demonstrate how combining motion with in-place entanglement offers lower overhead than entangling-zone approaches.
Neutral atom arrays have seen exciting progress as a platform for quantum computation. However, as we move towards the regime of fault-tolerance, the large-scale impact of fundamental features in these systems is not well-studied. In this work we point out that the use of movement in neutral atom arrays may set an unavoidable constraint on the speed of computation, erasing potential quantum advantage. As one solution, we propose a movement-free QEC architecture based on groups of interleaved surface codes. Our architecture enables fast, high-fidelity transversal CNOTs on surface codes in the same group. We also introduce interleaved lattice surgery to create high-capacity routing channels between groups. We validate our architecture through detailed numerical simulations of the underlying circuits and we evaluate its scalability through compilation of key benchmark applications. In regimes of high parallelism, we find our architecture leads to a similar to 3x reduction in compute time. Our architecture leverages experimentally demonstrated dual-species atom arrays which exhibit asymmetric interaction strengths that scale with 1/r(3) for interspecies interactions and with 1/r(6) for standard, intraspecies interactions. We examine how such scalings enable interleaving with high fidelity and propose how error rates required for QEC could be achieved. We also evaluate the tolerance of our architecture to two-qubit gate fidelities. We find the advantage of interleaving admits sizable tolerances of similar to 1x to 3x increase in error rates. We conclude the benefits of our proposed interleaved architecture grants strong motivation for future experimental efforts targeting longer range dual-species gates.
We present a case study and forward-looking perspective on co-design for hybrid quantum-classical algorithms, centered on the goal of empirical quantum advantage (EQA), which we define as a measurable performance gain using quantum hardware over state-of-the-art classical methods on the same task. Because classical algorithms continue to improve, the EQA crossover point is a moving target; nevertheless, we argue that a persistent advantage is possible for our application class even if the crossover point shifts. Specifically, our team examines the task of biomarker discovery in precision oncology. We push the limitations of the best classical algorithms, improving them as best as we can, and then augment them with a quantum subroutine for the task where we are most likely to see performance gains. We discuss the implementation of a quantum subroutine for feature selection on current devices, where hardware constraints necessitate further co-design between algorithm and physical device capabilities. Looking ahead, we perform resource analysis to explore a plausible EQA region on near/intermediate-term hardware, considering the impacts of advances in classical and quantum computing on this regime. Finally, we outline potential clinical impact and broader applications of this hybrid pipeline beyond oncology.
Emergence of fault-tolerant quantum computers (FTQC) brings about promise of harnessing the power of quantum computing at larger scale. At the same time, as quantum computers are expected to process more sensitive information, there is a need to understand the security issues in fault-tolerant quantum computers, and develop defenses for attacks that may compromise confidentiality or integrity of the data processed by FTQC. While noisy intermediate-scale quantum (NISQ) computers have already been studied from the security perspective, understanding security issues with FTQC is still an open research question. To address the missing research gap, this work presents the first exploration of possible security vulnerabilities of FTQC. The work presents analysis of possible threat models and outlines potential vulnerabilities of FTQC. Understanding the landscape of the threats can help lead to development of safer FTQC design at both software and hardware levels.
Susmit Biswas合作论文数Advanced Micro Devices, Inc14