We provide practical simulation methods for scalar field theories on a quantum computer that yield improved asymptotics as well as concrete gate estimates for the simulation and physical qubit estimates using the surface code. We achieve these improvements through two optimizations. First, we consider a finite volume approach for estimating the elements of the S-matrix. This approach is appropriate in general for 1+1D and for certain low-energy elastic collisions in higher dimensions. Second, we implement our approach using a series of different fault-tolerant simulation algorithms for Hamiltonians formulated both in the field occupation basis and field amplitude basis. Our algorithms are based on either second-order Trotterization or qubitization. The cost of Trotterization in occupation basis scales as O(lambda N7Q3/(M5/2F3/2)) where ) is the coupling strength, N is the occupation cutoff, Qis the volume of the spatial lattice, M is the mass of the particles and F is the uncertainty in the energy calculation used for the S-matrix determination. Qubitization in the field basis scales as O(Q2(k2A + kM2)/E), where k is the cutoff in the field and A is a scaled coupling constant. We find in both cases that the bounds suggest physically meaningful simulations can be performed using on the order of 4 & times; 106 physical qubits and 1012 T-gates which corresponds to roughly one day on a superconducting quantum computer with surface code and a cycle time of 100 ns. This places the simulation of scalar field theory within striking distance of the gate counts for the best available chemistry simulation results.
Quantum algorithms claim significant speedup over their classical counterparts for solving many problems. An important aspect of many of these algorithms is the existence of a quantum oracle, which needs to be implemented efficiently in order to realize the claimed advantages in practice. A quantum random access memory (QRAM) is a promising architecture for realizing these oracles. In this paper we develop a new design for QRAM and implement it with Clifford+T circuit. We focus on optimizing the T-count and T-depth since non-Clifford gates are the most expensive to implement fault-tolerantly in most error correction schemes. Integral to our design is a polynomial encoding of bit strings and so we refer to this design as $$\text {QRAM}_{poly}$$ . Compared to the previous state-of-the-art bucket brigade architecture for QRAM, we achieve an exponential improvement in T-depth, while reducing T-count and keeping the qubit-count same. Specifically, if N is the number of memory locations to be queried, then $$\text {QRAM}_{poly}$$ has T-depth $$O(\log \log N)$$ , T-count $$O(N-\log N)$$ and uses O(N) logical qubits, while the bucket brigade circuit has T-depth $$O(\log N)$$ , T-count O(N) and uses O(N) qubits. Combining two $$\text {QRAM}_{poly}$$ we design a quantum look-up-table, $$\text {qLUT}_{poly}$$ , that has T-depth $$O(\log \log N)$$ , T-count $$O(\sqrt{N})$$ and qubit count $$O(\sqrt{N})$$ . A quantum look-up table (qLUT) or quantum read-only memory (QROM) has restricted functionality than a QRAM. For example, it cannot write into a memory location and the circuit needs to be compiled each time the contents of the memory change. The previous state-of-the-art CSWAP architecture has T-depth $$O(\sqrt{N})$$ , T-count $$O(\sqrt{N})$$ and qubit count $$O(\sqrt{N})$$ . Thus we achieve a double exponential improvement in T-depth while keeping the T-count and qubit-count asymptotically same. Additionally, with our polynomial encoding of bit strings, we develop a method to optimize the Toffoli-count of circuits, specially those consisting of multi-controlled-NOT gates.
An efficient implementation of unitary operators is important in order to practically realize the computational advantages claimed by quantum algorithms over their classical counterparts. In this paper we study the potential of using reinforcement learning (RL) in order to synthesize quantum circuits, while optimizing the T-count and CS-count, of unitaries that are exactly implementable by the Clifford+T and Clifford+CS gate sets, respectively. In general, the complexity of existing algorithms depend exponentially on the number of qubits and the non-Clifford-count of unitaries. We have designed our RL framework to work with channel representation of unitaries, that enables us to perform matrix operations efficiently, using integers only. We have also incorporated pruning heuristics and a canonicalization of operators, in order to reduce the search complexity. As a result, compared to previous works, we are able to implement significantly larger unitaries, in less time, with much better success rate and improvement factor. Our results for Clifford+T synthesis on two qubits achieve close-to-optimal decompositions for up to 100 T gates, 5 times more than previous RL algorithms and to the best of our knowledge, the largest instances achieved with any method to date. Our RL algorithm is able to recover previously-known optimal linear complexity algorithm for T-count-optimal decomposition of 1 qubit unitaries. For 2-qubit Clifford+CS unitaries, our algorithm achieves a linear complexity, something that could only be accomplished by a previous algorithm using SO(6) representation.
In this paper we study the universal V-basis gate sets, which have also been shown to be fault tolerant. Our methods and results can be applied to arbitrary dimensional basis gates, but we explicitly state results for basis gates for SU(2) and SU(4). We also include Cliffords, as done in earlier works. We introduce generating sets in order to represent any unitary implementable by these gate sets and with these we derive a bound on the V count of arbitrary multiqubit unitaries. We analyze the channel representation of the generating set elements, with the help of which we infer that none of these basis gates can be implemented exactly with the other universal fault-tolerant gate sets like Clifford+T, Clifford+CS, Clifford+Toffoli. Also T, CS, and Toffoli cannot be exactly implemented in the V basis. In fact, basis unitaries in dimension four cannot be implemented exactly with those in dimension two. We develop V-count-optimal synthesis algorithms for both approximately and exactly implementable multiqubit unitaries. With the help of these we show that two basis unitaries in dimension four are required to implement each basis unitary in dimension two. The space and time complexities of our provable algorithms are exponential in V count. But for the special case of one-qubit unitaries we achieve a complexity that is linear in V count. Both space and time complexities of our heuristic algorithms are polynomial in V count.
We provide an explicit recursive divide and conquer approach for simulating quantum dynamics and derive a discrete first quantized non-relativistic QED Hamiltonian based on the many-particle Pauli Fierz Hamiltonian. We apply this recursive divide and conquer algorithm to this Hamiltonian and compare it to a concrete simulation algorithm that uses qubitization. Our divide and conquer algorithm, using lowest order Trotterization, scales for fixed grid spacing as $\widetilde{O}(\Lambda N^2\eta^2 t^2 /\epsilon)$ for grid size $N$, $\eta$ particles, simulation time $t$, field cutoff $\Lambda$ and error $\epsilon$. Our qubitization algorithm scales as $\widetilde{O}(N(\eta+N)(\eta +\Lambda^2) t\log(1/\epsilon)) $. This shows that even a na\"ive partitioning and low-order splitting formula can yield, through our divide and conquer formalism, superior scaling to qubitization for large $\Lambda$. We compare the relative costs of these two algorithms on systems that are relevant for applications such as the spontaneous emission of photons, and the photoionization of electrons. We observe that for different parameter regimes, one method can be favored over the other. Finally, we give new algorithmic and circuit level techniques for gate optimization including a new way of implementing a group of multi-controlled-X gates that can be used for better analysis of circuit cost.
AbstractIn quantum computing there are quite a few universal gate sets, each having their own characteristics. In this paper we study the Clifford+CS universal fault-tolerant gate set. The CS gate is used is many applications and this gate set is an important alternative to Clifford+T. We introduce a generating set in order to represent any unitary implementable by this gate set and with this we derive a bound on the CS-count of arbitrary multi-qubit unitaries. Analysing the channel representation of the generating set elements, we infer $${\mathcal {J}}_n^{CS}\subset {\mathcal {J}}_n^T$$ J n CS ⊂ J n T , where $${\mathcal {J}}_n^{CS}$$ J n CS and $${\mathcal {J}}_n^T$$ J n T are the set of unitaries exactly implementable by the Clifford+CS and Clifford+T gate sets, respectively. We develop CS-count optimal synthesis algorithms for both approximately and exactly implementable multi-qubit unitaries. With the help of these we derive a CS-count-optimal circuit for Toffoli, implying $${\mathcal {J}}_n^{Tof}={\mathcal {J}}_n^{CS}$$ J n Tof = J n CS , where $${\mathcal {J}}_n^{Tof}$$ J n Tof is the set of unitaries exactly implementable by the Clifford+Toffoli gate set. Such conclusions can have an important impact on resource estimates of quantum algorithms.
In this paper we study the Clifford+Toffoli universal fault-tolerant gate set. We introduce a generating set in order to represent any unitary implementable by this gate set and with this we derive a bound on the Toffoli-count of arbitrary multi-qubit unitaries. We analyse the channel representation of the generating set elements, with the help of which we infer |𝒥_n^Tof|<|𝒥_n^T|, where 𝒥_n^Tof and 𝒥_n^T are the set of unitaries exactly implementable by the Clifford+Toffoli and Clifford+T gate set, respectively. We develop Toffoli-count optimal synthesis algorithms for both approximately and exactly implementable multi-qubit unitaries. With the help of these we prove |𝒥_n^Tof|=|𝒥_n^CS|, where 𝒥_n^CS is the set of unitaries exactly implementable by the Clifford+CS gate set.
Let $R_\epsilon(\cdot)$ stand for the bounded-error randomized query complexity with error $\epsilon > 0$. For any relation $f \subseteq \{0,1\}^n \times S$ and partial Boolean function $g \subseteq \{0,1\}^m \times \{0,1\}$, we show that $R_{1/3}(f \circ g^n) \in \Omega(R_{4/9}(f) \cdot \sqrt{R_{1/3}(g)})$, where $f \circ g^n \subseteq (\{0,1\}^m)^n \times S$ is the composition of $f$ and $g$. We give an example of a relation $f$ and partial Boolean function $g$ for which this lower bound is tight. We prove our composition theorem by introducing a new complexity measure, the max conflict complexity $\bar \chi(g)$ of a partial Boolean function $g$. We show $\bar \chi(g) \in \Omega(\sqrt{R_{1/3}(g)})$ for any (partial) function $g$ and $R_{1/3}(f \circ g^n) \in \Omega(R_{4/9}(f) \cdot \bar \chi(g))$; these two bounds imply our composition result. We further show that $\bar \chi(g)$ is always at least as large as the sabotage complexity of $g$, introduced by Ben-David and Kothari.
We provide a new approach for compiling quantum simulation circuits that appear in Trotter, qDRIFT and multi-product formulas to Clifford and non-Clifford operations that can reduce the number of non-Clifford operations by a factor of up to $4$. In fact, the total number of gates reduce in many cases. We show that it is possible to implement an exponentiated sum of commuting Paulis with at most $m$ (controlled)-rotation gates, where $m$ is the number of distinct non-zero eigenvalues (ignoring sign). Thus we can collect mutually commuting Hamiltonian terms into groups that satisfy one of several symmetries identified in this work which allow an inexpensive simulation of the entire group of terms. We further show that the cost can in some cases be reduced by partially allocating Hamiltonian terms to several groups and provide a polynomial time classical algorithm that can greedily allocate the terms to appropriate groupings. We further specifically discuss these optimizations for the case of fermionic dynamics and provide extensive numerical simulations for qDRIFT of our grouping strategy to 6 and 4-qubit Heisenberg models, $LiH$, $H_2$ and observe a factor of 1.8-3.2 reduction in the number of non-Clifford gates. This suggests Trotter-based simulation of chemistry in second quantization may be even more practical than previously believed.
The accurate estimation of quantum observables is a critical task in science. With progress on the hardware, measur-ing a quantum system will become increas-ingly demanding, particularly for vari-ational protocols that require extensive sampling. Here, we introduce a mea-surement scheme that adaptively modi-fies the estimator based on previously ob-tained data. Our algorithm, which we call AEQuO, continuously monitors both the estimated average and the associated er-ror of the considered observable, and de-termines the next measurement step based on this information. We allow both for overlap and non-bitwise commutation re-lations in the subsets of Pauli operators that are simultaneously probed, thereby maximizing the amount of gathered infor-mation. AEQuO comes in two variants: a greedy bucket-filling algorithm with good performance for small problem instances, and a machine learning-based algorithm with more favorable scaling for larger in-stances. The measurement configuration determined by these subroutines is further post-processed in order to lower the er-ror on the estimator. We test our proto-col on chemistry Hamiltonians, for which AEQuO provides error estimates that im-prove on all state-of-the-art methods based on various grouping techniques or random-ized measurements, thus greatly lowering the toll of measurements in current and future quantum applications.
While mapping a quantum circuit to the physical layer one has to consider the numerous constraints imposed by the underlying hardware architecture. Connectivity of the physical qubits is one such constraint that restricts two-qubit operations, such as CNOT, to “connected” qubits. SWAP gates can be used to place the logical qubits on admissible physical qubits, but they entail a significant increase in CNOT-count. In this article, we consider the problem of reducing the CNOT-count in Clifford+T circuits on connectivity-constrained architectures, like noisy intermediate-scale quantum (NISQ) computing devices. We “slice” the circuit at the position of Hadamard gates and “build” the intermediate $\{\text {CNOT},{T}\}$ subcircuits using Steiner trees, significantly improving on previous methods. We compared the performance of our algorithms while mapping different benchmark and random circuits to some well-known architectures, such as 9-qubit square grid, 16-qubit square grid, Rigetti 16-qubit Aspen, 16-qubit IBM QX5, and 20-qubit IBM Tokyo. Our methods give less CNOT-count compared to Qiskit and TKET transpiler as well as using SWAP gates. Assuming most of the errors in an NISQ circuit implementation are due to CNOT errors, then our method would allow circuits with a few times more CNOT gates be reliably implemented than the previous methods would permit.
An important part of reaping computational advantage from a quantum computer is to reduce the quantum resources needed to implement a desired quantum algorithm. Quantum algorithms that are too large to be practical on noisy intermediate scale quantum devices will require fault-tolerant error correction. This work focuses on reducing the physical cost of implementing quantum algorithms when using the state-of-the-art fault-tolerant quantum error correcting codes, in particular, those for which implementing the T gate consumes vastly more resources than the other gates in the gate set. More specifically, in this paper we consider the group of unitaries that can be exactly implemented by a quantum circuit consisting of the Clifford + T gate set. The Clifford + T gate set is a universal gate set and in this group, using state-of-the-art surface codes, the T gate is by far the most expensive component to implement fault-tolerantly. So it is important to minimize the number of T gates necessary for a fault-tolerant implementation. Our primary interest is to compute a circuit for a given n-qubit unitary U, using the minimum possible number of T gates (called the T-count of unitary U). We consider the problem COUNT-T, the optimization version of which aims to find the T-count of U. In its decision version the goal is to decide if the T-count is at most some positive integer m. Given an oracle for COUNT-T, we can compute a T-count-optimal circuit in time polynomial in the T-count and dimension of U. We give a provable classical algorithm that solves COUNT-T (decision) in time ON2(c-1) left ceiling mc right ceiling poly(m,N) ON2 left ceiling mc right ceiling poly(m,N) , where N = 2 (n) and c > 2. This gives a space-time trade-off for solving this problem with variants of meet-in-the-middle techniques. We also introduce an asymptotically faster multiplication method that shaves a factor of N (0.7457) off of the overall complexity. Lastly, beyond our improvements to the rigorous algorithm, we give a heuristic algorithm that outputs a T-count-optimal circuit and has space and time complexity poly(m, N), under some assumptions. In our heuristic algorithm we developed a novel way of pruning the search space. While our heuristic method still scales exponentially with the number of qubits (though with a lower exponent), there is a large improvement by going from exponential to polynomial scaling with m. We implemented our heuristic algorithm with up to 4 qubit unitaries and obtained a significant improvement in time. For all benchmark and random unitaries we studied, the T-count returned by our algorithm is at most the T-count of their circuits shown in previous papers.
The Super-SAT (SSAT) problem was introduced in [1], [2] to prove the NP-hardness of approximation of two popular lattice problems - Shortest Vector Problem and Closest Vector Problem. SSAT is conjectured to be NP-hard to approximate to within a factor of nc (c is positive constant, n is the size of the SSAT instance). In this paper we prove this conjecture assuming the Projection Games Conjecture (PGC) [3]. This implies hardness of approximation of these lattice problems within polynomial factors, assuming PGC. We also reduce SSAT to the Nearest Codeword Problem and Learning Halfspace Problem [4]. This proves that both these problems are NP-hard to approximate within a factor of Nc′/loglogn (c′ is positive constant, N is the size of the instances of the respective problems). Assuming PGC these problems are proved to be NP-hard to approximate within polynomial factors.
We investigate the problem of synthesizing T-depth optimal quantum circuits over the Clifford+T gate set. First we construct a special subset of T-depth 1 unitaries, such that it is possible to express the T-depth-optimal decomposition of any unitary as product of unitaries from this subset and a Clifford (up to global phase). The cardinality of this subset is at most $n\cdot 2^{5.6n}$. We use nested meet-in-the-middle (MITM) technique to develop algorithms for synthesizing provably \emph{depth-optimal} and \emph{T-depth-optimal} circuits for exactly implementable unitaries. Specifically, for synthesizing T-depth-optimal circuits, we get an algorithm with space and time complexity $O\left(\left(4^{n^2}\right)^{\lceil d/c\rceil}\right)$ and $O\left(\left(4^{n^2}\right)^{(c-1)\lceil d/c\rceil}\right)$ respectively, where $d$ is the minimum T-depth and $c\geq 2$ is a constant. This is much better than the complexity of the algorithm by Amy et al.(2013), the previous best with a complexity $O\left(\left(3^n\cdot 2^{kn^2}\right)^{\lceil \frac{d}{2}\rceil}\cdot 2^{kn^2}\right)$, where $k>2.5$ is a constant. We design an even more efficient algorithm for synthesizing T-depth-optimal circuits. The claimed efficiency and optimality depends on some conjectures, which have been inspired from the work of Mosca and Mukhopadhyay (2020). To the best of our knowledge, the conjectures are not related to the previous work. Our algorithm has space and time complexity $poly(n,2^{5.6n},d)$ (or $poly(n^{\log n},2^{5.6n},d)$ under some weaker assumptions).
Abstract We design an algorithm to determine the (minimum) T-count of any n-qubit (n ≥ 1) unitary W of size 2 n × 2 n , over the Clifford+T gate set. The space and time complexity of our algorithm are $$O\left({2}^{2n}\right)$$ O 2 2 n and $$O\left({2}^{2n{{{{\mathcal{T}}}}}_{\epsilon }(W)+4n}\right)$$ O 2 2 n T ϵ ( W ) + 4 n , respectively. $${{{{\mathcal{T}}}}}_{\epsilon }(W)$$ T ϵ ( W ) (ϵ-T-count) is the (minimum) T-count of an exactly implementable unitary U ( $${{{\mathcal{T}}}}(U)$$ T ( U ) ), such that d(U,W) ≤ ϵ and $${{{\mathcal{T}}}}(U)\le {{{\mathcal{T}}}}({U}^{{\prime} })$$ T ( U ) ≤ T ( U ′ ) where $${U}^{{\prime} }$$ U ′ is any exactly implementable unitary with $$d({U}^{{\prime} },W)\le \epsilon$$ d ( U ′ , W ) ≤ ϵ . d(. , .) is the global phase invariant distance. Our algorithm can also be used to determine the (minimum) T-depth as well as the minimum non-Clifford-gate count or depth required to implement any multi-qubit unitary with a finite universal gate set like Clifford+CS, Clifford+V, etc. For small enough ϵ, we can synthesize the optimal circuits.
In this work, we give provable sieving algorithms for the Shortest Vector Problem (SVP) and the Closest Vector Problem (CVP) on lattices in ℓp norm (1≤p≤∞). The running time we obtain is better than existing provable sieving algorithms. We give a new linear sieving procedure that works for all ℓp norm (1≤p≤∞). The main idea is to divide the space into hypercubes such that each vector can be mapped efficiently to a sub-region. We achieve a time complexity of 22.751n+o(n), which is much less than the 23.849n+o(n) complexity of the previous best algorithm. We also introduce a mixed sieving procedure, where a point is mapped to a hypercube within a ball and then a quadratic sieve is performed within each hypercube. This improves the running time, especially in the ℓ2 norm, where we achieve a time complexity of 22.25n+o(n), while the List Sieve Birthday algorithm has a running time of 22.465n+o(n). We adopt our sieving techniques to approximation algorithms for SVP and CVP in ℓp norm (1≤p≤∞) and show that our algorithm has a running time of 22.001n+o(n), while previous algorithms have a time complexity of 23.169n+o(n).
Many quantum algorithms can be written as a composition of unitaries, some of which can be exactly synthesized by a universal fault-tolerant gate set like Clifford+T, while others can be approximately synthesized. One task of a quantum compiler is to synthesize each approximately synthesizable unitary up to some approximation error, such that the error of the overall unitary remains bounded by a certain amount. In this paper we consider the case when the errors are measured in the global phase invariant distance. Apart from deriving a relation between this distance and the Frobenius norm, we show that this distance composes. If a unitary is written as a composition (product and tensor product) of other unitaries, we derive bounds on the error of the overall unitary as a function of the errors of the composed unitaries. Our bound is better than the sum-of-error bound, derived by Bernstein- Vazirani(1997), for the operator norm. This builds the intuition that working with the global phase invariant distance might give us a lower resource count while synthesizing quantum circuits. Next we consider the following problem. Suppose we are given a decomposition of a unitary, that is, the unitary is expressed as a composition of other unitaries. We want to distribute the errors in each component such that the resource-count (specifically T-count) is optimized. We consider the specific case when the unitary can be decomposed such that the R z (θ) gates are the only approximately synthesizable component. We prove analytically that for both the operator norm and global phase invariant distance, the error should be distributed equally among these components (given some approximations). The optimal number of T-gates obtained by using the global phase invariant distance is less than what is obtained using the operator norm. Furthermore, we show that in case of approximate Quantum Fourier Transform, the error obtained by pruning rotation gates is less when measured in this distance, rather than the operator norm.
The Super-SAT or SSAT problem was introduced by Dinur, Kindler, Raz and Safra [DKRS03, Din02] to prove the NP-hardness of approximation of two popular lattice problems Shortest Vector Problem (SVP) and Closest Vector Problem (CVP). They conjectured that SSAT is NP-hard to approximate to within factor n for some constant c > 0, where n is the size of the SSAT instance. In this paper we prove this conjecture assuming the Projection Games Conjecture (PGC), given by Moshkovitz [Mos12]. This implies hardness of approximation of SVP and CVP within polynomial factors, assuming the Projection Games Conjecture. We also reduce SSAT to the Nearest Codeword Problem (NCP) and Learning Halfspace Problem (LHP), as considered by Arora, Babai, Stern and Sweedyk [ABSS97]. This proves that both these problems are NP-hard to approximate within factor N c ′/ log logn for some constant c > 0 where N is the size of the instances of the respective problems. Assuming the Projection Games Conjecture these problems are proved to be NP-hard to approximate within polynomial factors. mukhopadhyay.priyanka@gmail.com Much of this work was done while the author was in Centre for Quantum Technologies, National University of Singapore.
Blomer and Naewe[BN09] modified the randomized sieving algorithm of Ajtai, Kumar and Sivakumar[AKS01] to solve the shortest vector problem (SVP). The algorithm starts with $N = 2^{O(n)}$ randomly chosen vectors in the lattice and employs a sieving procedure to iteratively obtain shorter vectors in the lattice. The running time of the sieving procedure is quadratic in $N$. We study this problem for the special but important case of the $\ell_\infty$ norm. We give a new sieving procedure that runs in time linear in $N$, thereby significantly improving the running time of the algorithm for SVP in the $\ell_\infty$ norm. As in [AKS02,BN09], we also extend this algorithm to obtain significantly faster algorithms for approximate versions of the shortest vector problem and the closest vector problem (CVP) in the $\ell_\infty$ norm. We also show that the heuristic sieving algorithms of Nguyen and Vidick[NV08] and Wang et al.[WLTB11] can also be analyzed in the $\ell_{\infty}$ norm. The main technical contribution in this part is to calculate the expected volume of intersection of a unit ball centred at origin and another ball of a different radius centred at a uniformly random point on the boundary of the unit ball. This might be of independent interest.
Blomer and Naewe[BN09] modified the randomized sieving algorithm of Ajtai, Kumar and Sivakumar[AKS01] to solve the shortest vector problem (SVP). The algorithm starts with $N = 2^{O(n)}$ randomly chosen vectors in the lattice and employs a sieving procedure to iteratively obtain shorter vectors in the lattice. The running time of the sieving procedure is quadratic in $N$. study this problem for the special but important case of the $ell_infty$ norm. We give a new sieving procedure that runs in time linear in $N$, thereby significantly improving the running time of the algorithm for SVP in the $ell_infty$ norm. As in [AKS02],[BN09], we also extend this algorithm to obtain significantly faster algorithms for approximate versions of the shortest vector problem and the closest vector problem (CVP) in the $ell_infty$ norm. also show that the heuristic sieving algorithms of Nguyen and Vidick [NV08] and Wang et.al.[WLTB11] can also be analyzed in the $ell_{infty}$ norm. The main technical contribution in this part is to calculate the expected volume of intersection of a unit ball centred at origin and another ball of a different radius centred at a uniformly random point on the boundary of the unit ball. This might be of independent interest.