In psychometric sciences, such as social or behavioral sciences, and, similarly, in medical sciences, it is increasingly common to deal with longitudinal data organized as high-dimensional multidimensional arrays, also known as tensors. Within this framework, the time-continuous property of longitudinal data often implies a smooth functional structure on one of the tensor modes. To help researchers investigate such data, we introduce a new tensor decomposition approach based on the PARAFAC decomposition. Our approach allows researchers to represent a high-dimensional functional tensor as a low-dimensional set of functions and feature matrices. Furthermore, to capture the underlying randomness of the statistical setting more efficiently, we introduce a probabilistic latent model in the decomposition. A covariance-based block-relaxation algorithm is derived to obtain estimates of model parameters. Thanks to the covariance formulation of the solving procedure and thanks to the probabilistic modeling, the method can be used in sparse and irregular sampling schemes, making it applicable in numerous settings. Our approach is applied in the psychometric setting to help characterize multiple neurocognitive scores observed over time in the Alzheimer's Disease Neuroimaging Initiative study. Finally, intensive simulations show a notable advantage of our method in reconstructing tensors.
Identifying latent variables and their causal relationships with the observed variables is an important yet challenging problem. Existing works may face empirical challenges such as testing-order dependency, error propagation, and selection of an appropriate significance level. The complex nature of the problem also poses difficulties in scaling up to large data. In this work, we tackle these challenges in the problem known as measurement model learning. We propose a computationally efficient two-step procedure involving low-rank subspace identification and clustering. The first step is to retrieve from the observed covariance matrix a low-rank matrix that reflects the measurement model structure. The second step conducts clustering of the observed variables given the found low-rank matrix using what we call selective flipping. Our theoretical analysis shows the optimality of the proposed method in the large sample limit. Experimental results demonstrate that the proposed method achieves improvements in both learning accuracy and computational efficiency.
Regularized Generalized Canonical Correlation Analysis (RGCCA) is a general statistical framework for multiblock data analysis. RGCCA enables deciphering relationships between several sets of variables and subsumes many well-known multivariate analysis methods as special cases. However, RGCCA only deals with vector-valued blocks, disregarding their possible higher-order structures. This paper presents Tensor GCCA (TGCCA), a new method for analyzing higher-order tensors with canonical vectors admitting an orthogonal rank-R CP decomposition. Moreover, two algorithms for TGCCA, based on whether a separable covariance structure is imposed or not, are presented along with convergence guarantees. The efficiency and usefulness of TGCCA are evaluated on simulated and real data and compared favorably to state-of-the-art approaches.
In this paper, we introduce Functional Generalized Canonical Correlation Analysis (FGCCA), a new framework for exploring associations between multiple random processes observed jointly. The framework is based on the multiblock Regularized Generalized Canonical Correlation Analysis (RGCCA) framework. It is robust to sparsely and irregularly observed data, making it applicable in many settings. We establish the monotonic property of the solving procedure and introduce a Bayesian approach for estimating canonical components. We propose an extension of the framework that allows the integration of a univariate or multivariate response into the analysis, paving the way for predictive applications. We evaluate the method's efficiency in simulation studies and present a use case on a longitudinal dataset.
Regularized generalized canonical correlation analysis (RGCCA) is a general statistical framework for multiblock data analysis. However, multiblock data often have missing structure, i.e., data in one or more blocks may be completely unobserved for a sample. In this work, several solutions were investigated to properly handle missing data structures within the framework of RGCCA then compared on simulations.
In this paper, we propose a novel variational approach for supervised classification based on transform learning. Our approach consists of formulating an optimization problem on both the transform matrix and the centroids of the classes in a low-dimensional transformed space. The loss function is based on the distance to the centroids, which can be chosen in a flexible manner. To avoid trivial solutions or highly correlated clusters, our model incorporates a penalty term on the centroids, which encourages them to be separated. The resulting non-convex and non-smooth minimization problem is then solved by a primal-dual alternating minimization strategy. We assess the performance of our method on a bunch of supervised classification problems and compare it to state-of-the-art methods.
A B S T R A C T Modeling multidimensional data using tensor models, particularly through the Canonical Polyadic (CP) model, can be found in large numbers of timely and important signal-based applications. However, the computational complexity in the case of high-order and large-scale tensors remains a challenge that prevents the implementation of the CP model in practice. While some algorithms in the literature deal with large-scale problems, others target high-order tensors. Nevertheless, these algorithms encounter major issues when both problems are present. In this paper, we propose a parallelizable strategy based on the tensor network theory, to deal simultaneously with both high-order and large-scale problems. We show the usefulness of the proposed strategy in reducing the computation time on a realistic electroencephalography data set.(c) 2022 Elsevier B.V. All rights reserved.
Regularized generalized canonical correlation analysis (RGCCA) is a general multiblock data analysis framework that encompasses several important multivariate analysis methods such as principal component analysis, partial least squares regression, and several versions of generalized canonical correlation analysis. In this article, we extend RGCCA to the case where at least one block has a tensor structure. This method is called multiway generalized canonical correlation analysis (MGCCA). Convergence properties of the MGCCA algorithm are studied, and computation of higher-level components are discussed. The usefulness of MGCCA is shown on simulation and on the analysis of a cognitive study in human infants using electroencephalography (EEG).
This paper presents a Baseline Removal method in the context of spectrometry gamma. The method implements an estimator for the full continuum based on the observation of local minima. This estimator is constructed from the statistical properties of the signal and is therefore easily explainable. The method involves a limited number of fixed parameters, which allows the automation of the process. Moreover, the method is adaptable to any peaks width, which makes it suitable for both HPGe spectrometers and scintillators. Application to real gamma spectrometry measurements are presented, as well as a discussion about the choice of the parameters, for which an adjustment is proposed.
Early life stages are vulnerable to environmental hazards and present important windows of opportunity for lifelong disease prevention. This makes early life a relevant starting point for exposome studies. The Advancing Tools for Human Early Lifecourse Exposome Research and Translation (ATHLETE) project aims to develop a toolbox of exposome tools and a Europe-wide exposome cohort that will be used to systematically quantify the effects of a wide range of community- and individual-level environmental risk factors on mental, cardiometabolic, and respiratory health outcomes and associated biological pathways, longitudinally from early pregnancy through to adolescence. Exposome tool and data development include as follows: (1) a findable, accessible, interoperable, reusable (FAIR) data infrastructure for early life exposome cohort data, including 16 prospective birth cohorts in 11 European countries; (2) targeted and nontargeted approaches to measure a wide range of environmental exposures (urban, chemical, physical, behavioral, social); (3) advanced statistical and toxicological strategies to analyze complex multidimensional exposome data; (4) estimation of associations between the exposome and early organ development, health trajectories, and biological (metagenomic, metabolomic, epigenetic, aging, and stress) pathways; (5) intervention strategies to improve early life urban and chemical exposomes, co-produced with local communities; and (6) child health impacts and associated costs related to the exposome. Data, tools, and results will be assembled in an openly accessible toolbox, which will provide great opportunities for researchers, policymakers, and other stakeholders, beyond the duration of the project. ATHLETE’s results will help to better understand and prevent health damage from environmental exposures and their mixtures from the earliest parts of the life course onward.
Multidimensional signal processing is receiving a lot of interest recently due to the wide spread appearance of multidimensional signals in different applications of data science. Many of these fields rely on prior knowledge of particular properties, such as sparsity for instance, in order to enhance the performance and the efficiency of the estimation algorithms. However, these multidimensional signals are, often, structured into high-order tensors, where the computational complexity and storage requirements become an issue for growing tensor orders. In this paper, we present a sparse-based Joint dImensionality Reduction And Factors rEtrieval (JIRAFE). More specifically, we assume that an arbitrary factor admits a decomposition into a redundant dictionary coded as a sparse matrix, called the sparse coding matrix. The goal is to estimate the sparse coding matrix in the Tensor-Train model framework.
This study focused on the assessment of radio-frequency electromagnetic fields (RF-EMF) exposure in a realistic apartment due to the presence of a WiFi source deployed in uncertain position. In order to describe the 2D spatial distribution of electric field induced in the whole apartment for whatever position of the WiFi source, an innovative approach that combines Principal Component Analysis (PCA) and Gaussian process regression (Kriging method) was applied. The 2D surrogate model was used to investigate the exposure in three different usage scenarios of the WiFi sources, i.e. surfing to a new web site, using a Skype video call and watching a You Tube video at 1080p, evaluating the electric field E induced at each location of the apartment for 10,000 different positions of the source. Across all the examined conditions, we found E values distributions with median values in the range 2.2-96.1 mV/m and 90 th percentiles in the range 4.9-209.3 mV/m. The 2D surrogate model allowed obtaining a complete statistical description of the exposure for any positions of the WiFi source in the apartment, with a computational effort equal to about 10% of the one needed by using only the WiCa Heuristic Indoor Propagation Prediction (WHIPP) network planner.
This paper presents a method to estimate the continuum of a gamma rays spectrum through the observation of local minima. The method is simple, automatable and has a large scope of application. Indeed, it is not limited by the peaks width, and consequently it is usable with GeHP as well as with scintillators spectra. In the extent where the method exploits signal properties, its operation is easily explainable. It involves a limited set of meaningful parameters for which an adjustment is proposed. The potential of this method is demonstrated through simulations but also through real gamma spectrometry measurements.
In different application fields, heterogeneous data sets are structured into either matrices or higher-order tensors. In some cases, these structures present the property of having common underlying factors, which is used to improve the efficiency of factor-matrices estimation in the process of the so-called coupled matrix-tensor factorization (CMTF). Many methods target the CMTF problem relying on alternating algorithms or gradient approaches. However, computational complexity remains a challenge when the data sets are tensors of high-order, which is linked to the well-known "curse of dimensionality". In this paper, we present a methodological approach, using the Joint dImensionality Reduction And Factors rEtrieval (JIRAFE) algorithm for joint factorization of high-order tensor and matrix. This approach reduces the high-order CMTF problem into a set of 3-order CMTF and canonical polyadic decomposition (CPD) problems. The proposed algorithm is evaluated on simulation and compared with a gradient-based method.
This study focused on the evaluation of the electric field 2D spatial distribution of the E-field in a one-floor apartment when a WiFi source is placed in uncertain position. An innovative approach that combines Principal Component Analysis and Kriging model in order to build space-dependent surrogate models was applied and validated. Preliminary results showed the feasibility of the approach.
PURPOSE:To demonstrate that fast-kz spokes can be used in parallel transmission to homogenize flip angle ramp profiles (known as TONE) in slab selections, and thereby improve Time-Of-Flight angiography of the whole human brain at 7T. METHODS:B1+ and B0 maps were measured on seven human brains with a z-segmented coil connected to an 8-channel pTx system. Tailored two-spoke pulses were designed under strict hardware and SAR constraints for uniform slab profile before transforming their subpulse waveforms for linearly-increasing flip-angle ramps. Increasing angulations along the feet-head direction were prescribed in 2-slab and 3-slab TOF acquisitions. Excitation patterns were simulated and compared with RF-shimmed (single spoke) ramp pulses. Excitation performances were assessed in ~10-min TOF acquisitions by visually inspecting Maximal Intensity Projections angiograms. RESULTS:The flip-angle ramp fidelity achieved by double spokes inside slabs of interest was improved by 30-40% compared to RF-shimmed ramps. This allowed better homogenizing signal along arteries, and depicting small vessels in distal areas of the brain, in comparison with RF-shimmed ramp pulses or double-spoke uniform excitation. CONCLUSION:Ramp double spokes used in conjunction with parallel transmission yield better blood saturation compensation and more finely resolved TOF angiograms than mere double spokes or ramp single spokes at 7T.
In this study, an innovative approach that combines Principal Component Analysis (PCA) and Gaussian process regression (Kriging method), never used before in the assessment of human exposure to electromagnetic fields (EMF), was applied to build space‐dependent surrogate models of the 3D spatial distribution of the electric field induced in central nervous system (CNS) of children of different ages exposed to uniform magnetic field at 50 Hz of 200 μT of amplitude with uncertain orientation. The 3D surrogate models showed very low normalized percentage mean square error (MSE) values, always lower than 0.16%, confirming the feasibility and accuracy of the approach in estimating the 3D spatial distribution of E with a low number of components. Results showed that the electric field values induced in CNS tissues of children were within the ICNIRP basic restrictions for general public, with 99th percentiles of the E values obtained for each orientation showing median values in the range 1.9–2.1 mV/m. Similar 3D spatial distributions of the electric fields were found to be induced in CNS tissues of children of different ages. Bioelectromagnetics. 9999:1–10, 2018. © 2019 Bioelectromagnetics Society.