Stock market returns are typically analyzed using standard regression models yet they reside on irregular domains, a natural scenario for graph signal processing. This motivates us to consider a market graph as an intuitive way to represent the relationships between financial assets. Traditional methods for estimating asset-return covariance operate under the assumption of statistical time-invariance, and are thus unable to appropriately infer the underlying structure of the market graph. To this end, this work introduces a class of graph spectral estimators which cater for the nonstationarity inherent to asset price movements, as a basis to represent the time-varying interactions between assets through a dynamic spectral market graph. Such an account of the time-varying nature of the asset-return covariance allows us to introduce the notion of dynamic spectral portfolio cuts, whereby the graph is partitioned into time-evolving clusters, thus allowing for robust and online asset allocation. The advantages of the proposed framework over traditional methods are demonstrated through numerical case studies using real-world price data.
A class of doubly stochastic graph shift operators (GSO) is proposed, which is shown to exhibit: (i) lower and upper L2-boundedness for locally stationary random graph signals, (ii) L2-isometry for i.i.d. random graph signals with the asymptotic increase in the incoming neighbourhood size of vertices, and (iii) preservation of the mean of any graph signal - all prerequisites for reliable graph neural networks. These properties are obtained through a statistical consistency analysis of the proposed graph shift operator, and by exploiting the dual role of the doubly stochastic GSO as a Markov (diffusion) matrix and as an unbiased expectation operator. For generality, we consider directed graphs which exhibit asymmetric connectivity matrices. The proposed approach is validated through an example on the estimation of a vector field.
Principal component analysis (PCA) is considered a quintessential data analysis technique when it comes to describing linear relationships between the features of a dataset. However, the well-known lack of robustness of PCA for non-Gaussian data and/or outliers often makes its practical use unreliable. To this end, we introduce a robust formulation of PCA based on the maximum correntropy criterion (MCC). By virtue of MCC, robust operation is achieved by maximising the expected likelihood of Gaussian distributed reconstruction errors. The analysis shows that the proposed solution reduces to a generalised power iteration, whereby: (i) robust estimates of the principal components are obtained even in the presence of outliers; (ii) the number of principal components need not be specified in advance; and (iii) the entire set of principal components can be obtained, unlike existing approaches. The advantages of the proposed maximum correntropy power iteration (MCPI) are demonstrated through an intuitive numerical example.
Global fixed-income returns exhibit highly structured correlations across maturities and economies (data modes), and their modeling therefore requires analysis tools that are capable of directly capturing the inherent multiway couplings present in such multimodal data. Yet, current analyses typically employ “flat-view” multivariate matrix models and their associated linear algebras; these are agnostic to the global data structure and can only describe local linear pairwise relationships between data entries. To address this issue, the authors first show that global fixed-income returns naturally reside on multi-modal lattice data structures, referred to as tensors. This serves as a basis to introduce a multilinear algebraic approach, inherent in tensors, to the modeling of the global term structure underlying multiple interest rate curves. Owing to the enhanced flexibility of multilinear algebra, statistical descriptors, such as correlations, exist between tensor columns and rows (fibers), as opposed to between individual entries in standard matrix analysis. This allows for the expression of the covariance of global returns as a joint multilinear decomposition of the maturity-domain and country-domain covariances. This not only drastically reduces the number of parameters required to fully capture the global return covariance structure, but also makes it possible to devise rigorous and tractable global portfolio management strategies; the authors tailor these specifically to each of the data domains and thereby fully exploit the lattice structure of global fixed-income returns. The ability of the proposed multilinear tensor approach to compactly describe the macroeconomic environment through economically meaningful factors is validated via empirical analysis that demonstrates the existence of maturity-domain and country-domain covariances underlying the interest rate curves of eight developed economies. TOPICS: Fixed income and structured finance, statistical methods, portfolio construction Key Findings ▪ This article shows that global fixed-income returns naturally reside on multimodal lattice data structures, referred to as tensors, which exhibit highly structured correlations across maturities and economies. ▪ The authors introduce a multilinear algebraic approach to the modeling of the global term structure underlying multiple interest rate curves, which allows for the expression of the covariance of global returns as a joint multilinear decomposition of the maturity-domain and country-domain covariances. ▪ This drastically reduces the number of parameters required to fully capture the global return covariance structure and makes it possible to devise rigorous and tractable global portfolio management strategies that are tailored to each of the data domains.
Classical portfolio optimization methods typically determine an optimal capital allocation through the implicit, yet critical, assumption of statistical time-invariance. Such models are inadequate for real-world markets as they employ standard time-averaging based estimators which suffer significant information loss if the market observables are non-stationary. To this end, we reformulate the portfolio optimization problem in the spectral domain to cater for the nonstationarity inherent to asset price movements and, in this way, allow for optimal capital allocations to be time-varying. Unlike existing spectral portfolio techniques, the proposed framework employs augmented complex statistics in order to exploit the interactions between the real and imaginary parts of the complex spectral variables, which in turn allows for the modelling of both harmonics and cyclostationarity in the time domain. The advantages of the proposed framework over traditional methods are demonstrated through numerical simulations using real-world price data.
A class of multivariate spectral representations for real-valued nonstationary random variables is introduced, which is characterised by a general complex Gaussian distribution. In this way, the temporal signal properties -- harmonicity, wide-sense stationarity and cyclostationarity -- are designated respectively by the mean, Hermitian variance and pseudo-variance of the associated time-frequency representation (TFR). For rigour, the estimators of the TFR distribution parameters are derived within a maximum likelihood framework and are shown to be statistically consistent, owing to the statistical identifiability of the proposed distribution parametrization. By virtue of the assumed probabilistic model, a generalised likelihood ratio test (GLRT) for nonstationarity detection is also proposed. Intuitive examples demonstrate the utility of the derived probabilistic framework for spectral analysis in low-SNR environments.
Many modern data analytics applications on graphs operate on domains where graph topology is not known a priori, and hence its determination becomes part of the problem definition, rather than serving as prior knowledge which aids the problem solution. Part III of this monograph starts by addressing ways to learn graph topology, from the case where the physics of the problem already suggest a possible topology, through to most general cases where the graph topology is learned from the data. A particular emphasis is on graph topology definition based on the correlation and precision matrices of the observed data, combined with additional prior knowledge and structural conditions, such as the smoothness or sparsity of graph connections. For learning sparse graphs (with small number of edges), the least absolute shrinkage and selection operator, known as LASSO is employed, along with its graph specific variant, graphical LASSO. For completeness, both variants of LASSO are derived in an intuitive way, and explained. An in-depth elaboration of the graph topology learning paradigm is provided through several examples on physically well defined graphs, such as electric circuits, linear heat transfer, social and computer networks, and spring-mass systems. As many graph neural networks (GNN) and convolutional graph networks (GCN) are emerging, we have also reviewed the main trends in GNNs and GCNs, from the perspective of graph signal filtering. Tensor representation of lattice-structured graphs is next considered, and it is shown that tensors (multidimensional data arrays) are a special class of graph signals, whereby the graph vertices reside on a high-dimensional regular lattice structure. This part of monograph concludes with two emerging applications in financial data processing and underground transportation networks modeling.
The area of Data Analytics on graphs promises a paradigm shift, as we approach information processing of new classes of data which are typically acquired on irregular but structured domains (such as social networks, various ad-hoc sensor networks). Yet, despite the long history of Graph Theory, current approaches tend to focus on aspects of optimisation of graphs themselves rather than on eliciting strategies relevant to the objective application of the graph paradigm, such as detection, estimation, statistical and probabilistic inference, clustering and separation from signals and data acquired on graphs. In order to bridge this gap, we first revisit graph topologies from a Data Analytics point of view, to establish a taxonomy of graph networks through a linear algebraic formalism of graph topology (vertices, connections, directivity). This serves as a basis for spectral analysis of graphs, whereby the eigenvalues and eigenvectors of graph Laplacian and adjacency matrices are shown to convey physical meaning related to both graph topology and higher-order graph properties, such as cuts, walks, paths, and neighborhoods. Through a number of carefully chosen examples, we demonstrate that the isomorphic nature of graphs enables both the basic properties of data observed on graphs and their descriptors (features) to be preserved throughout the data analytics process, even in the case of reordering of graph vertices, where classical approaches fail. Next, to illustrate the richness and flexibility of estimation strategies performed on graph signals, spectral analysis of graphs is introduced through eigenanalysis of mathematical descriptors of graphs and in a generic way. Finally, benefiting from enhanced degrees of freedom associated with graph representations, a framework for vertex clustering and graph segmentation is established based on graph spectral representation (eigenanalysis) which demonstrates the power of graphs in various data association tasks, from image clustering and segmentation trough to low-dimensional manifold representation. The supporting examples demonstrate the promise of Graph Data Analytics in modeling structural and functional/semantic inferences. At the same time, Part I serves as a basis for Part II and Part III which deal with theory, methods and applications of processing Data on Graphs and Graph Topology Learning from data.
Graph signal processing deals with signals which are observed on an irregular graph domain. While many approaches have been developed in classical graph theory to cluster vertices and segment large graphs in a signal independent way, signal localization based approaches to the analysis of data on graph represent a new research direction which is also a key to big data analytics on graphs. To this end, after an overview of the basic definitions in graphs and graph signals, we present and discuss a localized form of the graph Fourier transform. To establish an analogy with classical signal processing, spectral- and vertex-domain definitions of the localization window are given next. The spectral and vertex localization kernels are then related to the wavelet transform, followed by a study of filtering and inversion of the localized graph Fourier transform. For rigour, the analysis of energy representation and frames in the localized graph Fourier transform is extended to the energy forms of vertex-frequency distributions, which operate even without the need to apply localization windows. Another link with classical signal processing is established through the concept of local smoothness, which is subsequently related to the particular paradigm of signal smoothness on graphs. This all represents a comprehensive account of the relation of general vertex-frequency analysis with classical time-frequency analysis, and important but missing link for more advanced applications of graphs signal processing. The theory is supported by illustrative and practically relevant examples.
The area of Data Analytics on graphs deals with informationprocessing of data acquired on irregular but structured graphdomains. The focus of Part I of this monograph has beenon both the fundamental and higher-order graph properties,graph topologies, and spectral representations of graphs.Part I also establishes rigorous frameworks for vertex clusteringand graph segmentation, and illustrates the power ofgraphs in various data association tasks. Part II embarkson these concepts to address the algorithmic and practicalissues related to data/signal processing on graphs, withthe focus on the analysis and estimation of both deterministicand random data on graphs. The fundamental ideasrelated to graph signals are introduced through a simple andintuitive, yet general enough case study of multisensor temperaturefield estimation. The concept of systems on graphis defined using graph signal shift operators, which generalizethe corresponding principles from traditional learningsystems. At the core of the spectral domain representationof graph signals and systems is the Graph Fourier Transform(GFT), defined based on the eigendecomposition of both theadjacency matrix and the graph Laplacian. Spectral domainrepresentations are then used as the basis to introduce graphsignal filtering concepts and address their design, includingChebyshev series polynomial approximation. Ideas related tothe sampling of graph signals, and in particular the challengingtopic of data dimensionality reduction through graphsubsampling, are presented and further linked with compressivesensing. The principles of time-varying signals on graphsand basic definitions related to random graph signals arenext reviewed. Localized graph signal analysis in the jointvertex-spectral domain is referred to as the vertex-frequencyanalysis, since it can be considered as an extension of classicaltime-frequency analysis to the graph serving as signaldomain. Important aspects of the local graph Fourier transform(LGFT) are covered, together with its various formsincluding the graph spectral and vertex domain windowsand the inversion conditions and relations. A link betweenthe LGFT with a varying spectral window and the spectralgraph wavelet transform (SGWT) is also established.Realizations of the LGFT and SGWT using polynomial(Chebyshev) approximations of the spectral functions arefurther considered and supported by examples. Finally, energyversions of the vertex-frequency representations areintroduced, along with their relations with classical timefrequencyanalysis, including a vertex-frequency distributionthat can satisfy the marginal properties. The material issupported by illustrative examples.
The area of Data Analytics on graphs promises a paradigm shift as we approach information processing of classes of data, which are typically acquired on irregular but structured domains (social networks, various ad-hoc sensor networks). Yet, despite its long history, current approaches mostly focus on the optimization of graphs themselves, rather than on directly inferring learning strategies, such as detection, estimation, statistical and probabilistic inference, clustering and separation from signals and data acquired on graphs. To fill this void, we first revisit graph topologies from a Data Analytics point of view, and establish a taxonomy of graph networks through a linear algebraic formalism of graph topology (vertices, connections, directivity). This serves as a basis for spectral analysis of graphs, whereby the eigenvalues and eigenvectors of graph Laplacian and adjacency matrices are shown to convey physical meaning related to both graph topology and higher-order graph properties, such as cuts, walks, paths, and neighborhoods. Next, to illustrate estimation strategies performed on graph signals, spectral analysis of graphs is introduced through eigenanalysis of mathematical descriptors of graphs and in a generic way. Finally, a framework for vertex clustering and graph segmentation is established based on graph spectral representation (eigenanalysis) which illustrates the power of graphs in various data association tasks. The supporting examples demonstrate the promise of Graph Data Analytics in modeling structural and functional/semantic inferences. At the same time, Part I serves as a basis for Part II and Part III which deal with theory, methods and applications of processing Data on Graphs and Graph Topology Learning from data.
The focus of Part I of this monograph has been on both the fundamental properties, graph topologies, and spectral representations of graphs. Part II embarks on these concepts to address the algorithmic and practical issues centered round data/signal processing on graphs, that is, the focus is on the analysis and estimation of both deterministic and random data on graphs. The fundamental ideas related to graph signals are introduced through a simple and intuitive, yet illustrative and general enough case study of multisensor temperature field estimation. The concept of systems on graph is defined using graph signal shift operators, which generalize the corresponding principles from traditional learning systems. At the core of the spectral domain representation of graph signals and systems is the Graph Discrete Fourier Transform (GDFT). The spectral domain representations are then used as the basis to introduce graph signal filtering concepts and address their design, including Chebyshev polynomial approximation series. Ideas related to the sampling of graph signals are presented and further linked with compressive sensing. Localized graph signal analysis in the joint vertex-spectral domain is referred to as the vertex-frequency analysis, since it can be considered as an extension of classical time-frequency analysis to the graph domain of a signal. Important topics related to the local graph Fourier transform (LGFT) are covered, together with its various forms including the graph spectral and vertex domain windows and the inversion conditions and relations. A link between the LGFT with spectral varying window and the spectral graph wavelet transform (SGWT) is also established. Realizations of the LGFT and SGWT using polynomial (Chebyshev) approximations of the spectral functions are further considered. Finally, energy versions of the vertex-frequency representations are introduced.
Anthony G. Constantinides合作论文数Communications and Signal Processing Group of the Department of Electrical and Electronic Engineering3