Various computational problems as, e.g., equations with fractional diffusion operators, evaluation of high-dimensional integrals, the Møller–Plesset approach in quantum chemistry, etc., are easily solved by using approximations by rational functions or by exponential sums. In the case of Cauchy–Stieltjes or, respectively, Lebesgue–Stieltjes functions we provide a uniform proof of upper bounds of the convergence rates of their best approximations by rational functions or exponential sums. It turns out that the convergence rate by rational approximation is better than for exponential sums. We extend the analysis also to the approximation on infinite intervals and to the best approximation of the relative error. Instead of looking for the best approximation one can use the computationally cheaper quadrature method, in particular the sinc quadrature. The corresponding sharp error estimates are determined. The theoretical results are supported by numerical results.
In the paper ‘On the Dirac–Frenkel Variational Principle on Tensor Banach Spaces’, we provided a geometrical description of manifolds of tensors in Tucker format with fixed multilinear (or Tucker) rank in tensor Banach spaces, that allowed to extend the Dirac–Frenkel variational principle in the framework of topological tensor spaces. The purpose of this note is to extend these results to more general tensor formats. More precisely, we provide a new geometrical description of manifolds of tensors in tree-based (or hierarchical) format, also known as tree tensor networks, which are intersections of manifolds of tensors in Tucker format associated with different partitions of the set of dimensions. The proposed geometrical description of tensors in tree-based format is compatible with the one of manifolds of tensors in Tucker format.
A tensor v is the sum of at least rank(v) elementary tensors. In addition, a 'border rank' is defined: (rank) under bar (w) = r holds if r is the minimum integer such thatwis a limit of rank-r tensors. Usually, the set of rank-r tensors is not closed, i.e. tensors with r = (rank) under bar (w) < rank(w) may exist. It is easy to see that in such a case the representation of rank-r tensors v contains diverging elementary tensors as v approaches w. In a first part, werecall results about the uniform strength of the divergence in the case of general nonclosed tensor formats (restricted to finite dimensions). The second part discusses the r-term format for infinite-dimensional tensor spaces. It is shown that the general situation is very similar to the behaviour of finite-dimensional model spaces. The third part contains the main result: it is proved that in the case of <(rank)under bar>(w) = 2 < rank(w) the divergence strength is greater than or similar to epsilon(-1/2), i.e. if parallel to v - w parallel to < e and rank(v) <= 2, the parameters of v increase at least proportionally to epsilon(-1/2).
In this paper, we consider the Poisson equation on a “long” domain which is the Cartesian product of a one-dimensional long interval with a ( d − 1)-dimensional domain. The right-hand side is assumed to have a rank-1 tensor structure. We will present and compare methods to construct approximations of the solution which have tensor structure and the computational effort is governed by only solving elliptic problems on lower-dimensional domains. A zero-th order tensor approximation is derived by using tools from asymptotic analysis (method 1). The resulting approximation is an elementary tensor and, hence has a fixed error which turns out to be very close to the best possible approximation of zero-th order. This approximation can be used as a starting guess for the derivation of higher-order tensor approximations by a greedy-type method (method 2). Numerical experiments show that this method is converging towards the exact solution. Method 3 is based on the derivation of a tensor approximation via exponential sums applied to discretized differential operators and their inverses. It can be proved that this method converges exponentially with respect to the tensor rank. We present numerical experiments which compare the performance and sensitivity of these three methods.
A modification of standard linear iterative methods for the solution of linear equations is investigated aiming at improved data-sparsity with respect to a rank function. The convergence speed of the modified method is compared to the rank growth of its iterates for certain model cases. The considered general setup is common in the data-sparse treatment of high dimensional problems such as sparse approximation and low rank tensor calculus.
Scientific computations or measurements may result in huge volumes of data. Often these can be thought of representing a real-valued function on a high-dimensional domain, and can be conceptually arranged in the format of a tensor of high degree in some truncated or lossy compressed format. We look at some common post-processing tasks which are not obvious in the compressed format, as such huge data sets can not be stored in their entirety, and the value of an element is not readily accessible through simple look-up. The tasks we consider are finding the location of maximum or minimum, or minimum and maximum of a function of the data, or finding the indices of all elements in some interval --- i.e. level sets, the number of elements with a value in such a level set, the probability of an element being in a particular level set, and the mean and variance of the total collection. The algorithms to be described are fixed point iterations of particular functions of the tensor, which will then exhibit the desired result. For this, the data is considered as an element of a high degree tensor space, although in an abstract sense, the algorithms are independent of the representation of the data as a tensor. All that we require is that the data can be considered as an element of an associative, commutative algebra with an inner product. Such an algebra is isomorphic to a commutative sub-algebra of the usual matrix algebra, allowing the use of matrix algorithms to accomplish the mentioned tasks. We allow the actual computational representation to be a lossy compression, and we allow the algebra operations to be performed in an approximate fashion, so as to maintain a high compression level. One such example which we address explicitly is the representation of data as a tensor with compression in the form of a low-rank representation.
The discussion of topological tensor spaces has been started by Schatten [167] and Grothendieck [79, 80]. In Sect. 4.2 we discuss the question how the norms of V and W are related to the norm of $$V \otimes W.$$ From the viewpoint of functional analysis, tensor spaces of order 2 are of particular interest, since they are related to certain operator spaces (cf. §4.2.13). However, for our applications we are more interested in tensor spaces of order ≥ 3. These spaces are considered in Sect. 4.3. As preparation for the aforementioned sections and later ones, we need more or less well-known results from Banach space theory, which we provide in Sect. 4.1. Section 4.4 discusses the case of Hilbert spaces. This is important, since many applications are of this kind. Many of the numerical methods require scalar products. The reason is that, unfortunately, the solution of approximation problems with re- spect to general Banach norms is much more involved than those with respect to a scalar product.
The r-term representation v = Xr =1 Od j=1 v(j) ;i.e., a representation by sums of r elementary tensors, is already used in the algebraic definition (3.9) of tensors. In different fields, the r-term representation has different names: ‘canonical decomposition’ in psychometrics (cf. [48]), ‘parallel factors model’ (cf. [158]) in chemometrics.1 The word ‘representation’ is often replaced by ‘format’. The short form ‘CP’ is proposed by Comon [62] meaning ‘canonical polyadic decomposition’. Sometimes the polyadic decomposition (PD) is distinguished from the canonical polyadic decomposition (CPD) by the fact that in the latter case the used number of terms should be minimal (equal to the tensor rank; cf. Vervliet et al [296]). Here, the notation ‘r-term representation’ is used, where ‘r’ is considered as a variable from N0; which may be replaced by other variable names or numbers. Before we discuss the r-term representation in Section 7.3, we consider representations in general (Section 7.1) and the full representation (Section 7.2). The sensitivity of the r-term representation is analysed in Section 7.4. Section 7.5 discusses possible representations of the vectors v(j) 2Vj : We briefly mention the conversion from full format to r-term representation (cf. 7.6.1). We conclude with the discussion of (anti-)symmetric tensors (in Section 7.7) and other modifications of the r-term format (in Section 7.8). The discussion of arithmetical operations with tensors in r-term representation is postponed to Chapter 13. In this chapter we restrict our considerations to the exact representation in the r-term format. Approximations, which are of greater interest in practice, will be discussed in Chapter 9.
We consider elliptic partial differential equations in d variables and their discretisation in a product grid $$\mathbf{I} = \times^{d}_{j=1}I_{j}$$ . The solution of the discrete system is a grid function, which can directly be viewed as a tensor in $$\mathbf{V} = {\bigotimes}^{d}_{j=1}\mathbb{K}^{I_{j}}$$ . In Sect. 16.1 we compare the standard strategy of local refinement with the tensor approach involving regular grids. It turns out that the tensor approach can be more efficient. In Sect. 16.2 the solution of boundary value problems is discussed. A related problem is the eigenvalue problem discussed in Sect. 16.3. We concentrate ourselves to elliptic boundary value problems of second order. However, elliptic boundary value problems of higher order or parabolic problems lead to similar results.
In general, one tries to approximate a tensor v by another tensor u requiring less data. The reason is twofold: the memory size should decrease and, hopefully, operations involving u should require less computational work. In fact, $${\rm u} \in \mathcal{R}_{{\rm r}}$$ leads to decreasing cost for storage and operations as r decreases. However, the other side of the coin is an increasing approximation error. Correspondingly, in Sect. 9.1 two approximation strategies are presented, where either the representation rank r of u or the accuracy is prescribed. Before we study the approximation problem in general, two particular situations are discussed. Section 9.2 is devoted to r = 1, when $${\rm u} \in \mathcal{R}_{1}$$ is an elementary tensor. The matrix case d = 2 is recalled in Sect. 9.3. The properties observed in the latter two sections contrast with the true tensor case studied in Sect. 9.4. Numerical algorithms solving the approximation problem will be discussed in Sect. 9.5. Modified approximation problems are addressed in Sect. 9.6.
Tensors are in general large-scale data which require a special representation. These representations are also called a format. After mentioning the r -term and tensor subspace formats, we describe the hierarchical tensor format which is the most flexible one. Since operations with tensors often produce tensors of larger memory cost, truncation to reduced ranks is of utmost importance. The so-called higher-order singular-value decomposition (HOSVD) provides a save truncation with explicit error control. The paper explains in detail how the HOSVD procedure is performed within the hierarchical tensor format. Finally, we state special favourable properties of the HOSVD truncation.
The term ‘matrix-product state’ (MPS) is introduced in quantum physics (see, e.g., Verstraete-Cirac [294], [171, Eq. (2)]). The related tensor representation can be found already in Vidal [297] without a special naming of the representation. The method has been reinvented by Oseledets and Tyrtyshnikov ([234], [237], [242]) and called ‘TT decomposition’. We start in Section 12.1 with the finite-dimensional case. In Section 12.3 we show that the TT representation is a special form of the hierarchical format. Finally three conversions are considered: conversion from r-term format to TT format (cf. 12.4.1), from TT format into hierarchical format with a general tree Td (cf. 12.4.2), and vice versa, from general hierarchical format into TT format (cf. 12.4.3). A closely related variant of the TT format is the cyclic matrix product format. As we shall see in Section 12.5, the change from the tree structure to a proper graph structure may have negative consequences since, in general, such formats are nonclosed. The algorithms for obtaining HOSVD bases and for truncations are mentioned only briefly. The reason is the equivalence to the hierarchical format, so that the algorithms defined there can easily be transferred. The interested reader finds such algorithms in Oseledets [237].
The approximation of the function 1/x by exponential sums has several interesting applications. It is well known that best approximations with respect to the maximum norm exist. Moreover, the error estimates exhibit exponential decay as the number of terms increases. Here we focus on the computation of the best approximations. In principle, the problem can be solved by the Remez algorithm, however, because of the very sensitive behaviour of the problem the standard approach fails for a larger number of terms. The remedy described in the paper is the use of other independent variables of the exponential sum. We discuss the approximation error of the computed exponential sums up to 63 terms and hint to a webpage containing the corresponding coefficients.
The notion of minimal subspaces is closely connected with the representations of tensors, provided these representations can be characterised by (dimensions of) subspaces. A separate description of the theory of minimal subspaces can be found in Falcó-Hackbusch [57]. The tensor representations discussed in the later Chapters 8, 11, 12 will lead to subsets $${T_r}, {\mathcal{H}_r}, \mathbb{T}_{\rho}$$ of a tensor space. The results of this chapter will prove weak closedness of these sets. Another result concerns the question of a best approxima- tion: is the infimum also a minimum? In the positive case, it is guaranteed that the best approximation can be found in the same set. For tensors $${\rm{v}}\,\,\epsilon\,\,_{a}\bigotimes_{j=1}^{d}\,\,V_j$$ we shall define ‘minimal subspaces’ $$U_{j}^{\hbox{min}} (v) \subset V_j$$ in Sects. 6.1-6.4. In Sect. 6.5 we consider weakly convergent sequences $${\rm{v}}_n \rightharpoonup {\rm{v}}$$ and analyse the connection between $$U_{j}^{\hbox{min}} ({v}_{n})$$ and $$U_{j}^{\hbox{min}} (v)$$ . The main result will be presented in Theorem 6.24. While Sects. 6.1-6.5 discuss minimal subspaces of algebraic tensors $${\rm{v}}\,\,\epsilon\,\,_{a}{\bigotimes_{j=1}^{d}}\,\,V_j$$ , Sect. 6.6 investigates $$U_{j}^{\hbox{min}} (v)$$ for topological tensors $${\rm{v}}\,\,\epsilon\,\,_{\parallel.\parallel}{\bigotimes_{j=1}^{d}}\,\,V_j$$ . The final Sect. 6.7 is concerned with intersection spaces.
In 4.6 several tensor operations have been described. The numerical tensor calculus requires the practical realisation of these operations. In this chapter we describe the performance and arithmetical cost of the operations for the different formats. The discussed operations are the addition in Section 13.1, evaluation of tensor entries in Section 13.2, the scalar product and partial scalar product in Section 13.3, the change of bases in Section 13.4, general binary operations in Section 13.5, Hadamard product in Section 13.6, convolution of tensors in Section 13.7, matrixmatrix multiplication in Section 13.8, and matrix-vector multiplication in Section 13.9. Section 13.10 is devoted to special functions applied to tensors. In the last Section 13.11 we comment on the operations required for the treatment of Hartree– Fock and Kohn-Sham applications in quantum chemistry. In connection with the tensorisation discussed in Chapter 14, further operations and their cost will be discussed.
The exact representation of $${\rm v} \in {\rm V} = {\bigotimes}_{j=1}^{d} V_{j}$$ by a tensor subspace representation (8.6b) may be too expensive because of the high dimensions of the involved subspaces or even impossible since v is a topological tensor admitting no finite representation. In such cases we must be satisfied with an approximation u ≈ v which is easier to handle. We require that $$\mathbf{u} \in {\mathcal{T}}_{\mathbf{r}}$$ , i.e., there are bases $$\{{b}_{1}^{(j)},\ldots,{b}_{{r}_{j}}^{(j)}\} \subset {V }_{j}$$ such that 1 $$\mathbf{u} ={ \sum \nolimits }_{{i}_{1}=1}^{{r}_{1} }\cdots {\sum \nolimits }_{{i}_{d}=1}^{{r}_{d} }\mathbf{a}[{i}_{1}\cdots {i}_{d}]{\bigotimes}_{j=1}^{d}{b}_{{ i}_{j}}^{(j)}.$$ The basic task of this chapter is the following problem: 2 $$\text{ Given }\mathbf{v} \in \mathbf{V}\text{, find a suitable approximation }\mathbf{u} \in {\mathcal{T}}_{\mathbf{r}} \subset \mathbf{V,}$$ where $$\mathbf{r} = \left ({r}_{1},\ldots,{r}_{d}\right ) \in {\mathbb{N}}^{d}.$$ Finding $$\mathbf{u} \in {\mathcal{T}}_{\mathbf{r}}$$ means finding coefficients $$\mathbf{a}[{i}_{1}\cdots {i}_{d}]$$ as well as basis vectors b i (j) ∈ V j in (10.1). Problem (10.2) is formulated rather vaguely. If an accuracy ε > 0 is prescribed, r ∈ ℕ d as well as $$\mathbf{u} \in {\mathcal{T}}_{\mathbf{r}}$$ are to be determined. The strict minimisation of $$\left \Vert \mathbf{v} -\mathbf{u}\right \Vert$$ is often replaced by an appropriate approximation $$\mathbf{u}$$ requiring low computational cost. Instead of ε > 0, we may prescribe the rank vector r ∈ ℕ d in (10.2).Optimal approximations (so-called ‘best approximations’) will be studied in Sect. 10.2. While best approximations require an iterative computation, quasi-optimal approximations can be determined explicitly using the HOSVD basis introduced in Sect. 8.3. The latter approach is explained in Sect. 10.1.
The hierarchical tensor representation (notation: Hr) allows to keep the advantages of the subspace structure of the tensor subspace format Tr, but has only linear cost with respect to the order d concerning storage and operations. The hierarchy mentioned in the name is given by a ‘dimension partition tree’. The fact that the tree is binary, allows a simple application of the singular-value decomposition and enables an easy truncation procedure. After an introduction in Section 11.1, the algebraic structure of the hierarchical tensor representation is described in Section 11.2. While the algebraic representation uses subspaces, the concrete representation in Section 11.3 introduces frames or bases and the associated coefficient matrices in the hierarchy. Again, higher order singular-value decompositions (HOSVD) can be applied and the left singular vectors can be used as basis. In Section 11.4, the approximation in the Hr format is studied with respect to two aspects. First, the best approximation within Hr can be considered. Second, the HOSVD bases allow a quasi-optimal truncation. Section 11.5 discusses the joining of two representations. This important feature is needed if two tensors described by two different hierarchical tensor representations require a common representation.
In order to treat high-dimensional problems, one has to find data-sparse representations. Starting with a six-dimensional problem, we first introduce the low-rank approximation of matrices. One purpose is the reduction of memory requirements, another advantage is that now vector operations instead of matrix operations can be applied. In the considered problem, the vectors correspond to grid functions defined on a three-dimensional grid. This leads to the next separation: these grid functions are tensors in \(\mathbb {R}^{n}\otimes \mathbb {R}^{n}\otimes \mathbb {R}^{n}\) and can be represented by the hierarchical tensor format. Typical operations as the Hadamard product and the convolution are now reduced to operations between \(\mathbb {R}^{n}\) vectors. Standard algorithms for operations with vectors from \(\mathbb {R}^{n}\) are of order \(\mathcal {O}(n)\) or larger. The tensorisation method is a representation method introducing additional data-sparsity. In many cases, the data size can be reduced from \(\mathcal {O}(n)\) to \(\mathcal {O}(\log n)\). Even more important, operations as the convolution can be performed with a cost corresponding to these data sizes.
An important feature is the computation of a tensor from comparably few tensor entries. The input tensor v 2 V is assumed to be given in a full functional representation so that, on request, any entry can be determined. This partial information can be used to determine an approximation ~v 2 V. In the matrix case (d=2) an algorithm for this purpose is well-known under the name ‘cross approximation’ or ‘adaptive cross approximation’ (ACA). The generalisation to the multi-dimensional case is not straightforward. We present a multivariate cross approximation, which fits to the hierarchical format. If v 2 Hr holds with a known rank vector r, this tensor can be reproduced exactly, i.e., ~v = v. In the general case, the approximation is heuristic. Exact error estimates require either inspection of all tensor entries (which is practically impossible) or strong theoretical a priori knowledge. There are many different applications, where tensors are given in a full functional representation. Section 15.1 gives examples of multivariate functions constructed via integrals and describes multiparametric solutions of partial differential equations, which may originate from stochastic coefficients. Section 15.2 introduces the definitions of fibres and crosses. The matrix case is recalled in Section 15.3, while the true tensor case (d>3) is considered in Section 15.4.
Rainer E Burkard合作论文数Technical University Graz
Institute of Optimization and Discrete Mathematics (Mathematics B)
4
A. A. Reusken合作论文数RWTH Aachen Technical University2