Blind source separation (BSS), particularly independent component analysis (ICA), has been widely used in various fields of science such as biomedical signal processing to recover latent source signals from the observed mixture. While ICA is typically applied to individual datasets, many real-world applications share underlying sources across datasets. Independent vector analysis (IVA) extends ICA to jointly analyze multiple datasets by exploiting statistical dependencies across them. While various IVA methods have been presented in signal processing literature, the statistical properties of methods remains largely unexplored. This article introduces the IVA model, numerous density models used in IVA, and various classical IVA methods to statistics community highlighting the need for further theoretical developments.
The growing use of high-throughput sequencing (HTS) has enabled the large-scale production of compositional count data, driving progress in microbiome research. However, such count data are often high-dimensional, over-dispersed, and heavily zero-inflated, and they conflict with the continuity assumptions underlying log-ratio-based compositional data analysis (CoDA), creating substantial methodological challenges. This review organizes the scattered literature into a structured overview of established zero-handling strategies, categorizing them into zero-tolerant transformations, imputation approaches for rounded zeros, and statistical models for essential zeros. We further highlight the theoretical and practical consequences of applying continuous log-ratio frameworks to discrete compositional count data. Theoretically, violations of continuity may induce numerical instability and distort the geometric assumptions underlying CoDA. Practically, we examine how existing continuous imputation strategies behave when applied to lattice-valued count data within the classical log-ratio framework, comparing their performance in terms of matrix reconstruction and downstream analyses. Our results show that reconstruction accuracy alone does not always reflect downstream analytical performance. Overall, this review consolidates current zero-handling strategies and clarifies their appropriate scope of application. It also identifies methodological gaps that motivate future developments for high-dimensional compositional count data.
In this discussion, we comment on the paper by Leyder et al., which introduces a new framework for robust independent component analysis based on robust distance correlation. Our discussion focuses on the role of whitening, the choice of scatter matrices, and the implications of the independence property in robust ICA. We further argue that the minimum distance index is preferable to the Amari index as a performance measure, as it admits a direct connection to the limiting distribution of ICA estimators. Overall, the paper by Leyder et al. constitutes an important contribution to robust blind source separation and opens several promising avenues for future research.
ABSTRACT Satellite observations of atmospheric pollutants are often available only at coarse spatial resolution, which limits their use in local‐scale environmental analysis. Spatial downscaling methods aim to transform such data into high‐resolution fields. In this work, two widely used deep learning architectures—the super‐resolution deep residual network (SRDRN) and the encoder–decoder‐based UNet—for spatial downscaling, are extended with a lightweight temporal module that encodes observation time using either sinusoidal or radial basis function representations and integrates temporal features with spatial information. The proposed time‐aware extensions are evaluated in a case study on ozone downscaling over Italy. Results show that, with only a minor increase in computational cost, incorporating temporal information significantly improves downscaling accuracy and accelerates model convergence.
We evaluated four existing Monte Carlo-based methods for approximating prediction error variances and reliabilities in a multiple-trait genomic prediction model. All four methods yielded consistent approximations of the exact values. Genomic prediction models, such as genomic best linear unbiased prediction (GBLUP), use genomic data to improve the accuracy of estimated genetic values. As the number of genotypes and traits increases, the exact calculation of prediction error variances (PEVs) and reliabilities becomes computationally infeasible due to the need to invert the coefficient matrix of the mixed model equations, whose dimension increases directly with the number of individuals and traits. The objective of this study was to evaluate the applicability of the Monte Carlo (MC) sampling method to approximate PEVs and reliabilities in a multiple-trait GBLUP framework relevant to hybrid breeding. The MC method avoids direct matrix inversion by repeatedly sampling genetic values from their assumed distributions to approximate PEVs. We applied the MC method using four previously published formulas to approximate PEVs and reliabilities. All formulas produced consistent estimates of PEVs and reliabilities, with convergence rates depending on the formula, the level of reliability, and the MC sample size.
ABSTRACT The formation of ground‐level ozone follows complex nonlinear photochemical processes that depend on multiple environmental factors and have strong spatio‐temporal structures. Environmental data used to study these dynamics usually originate from multiple sources, including in situ monitoring stations and satellite observations. While in situ data provide more accurate measurements, they often suffer from missing values and limited coverage, making it beneficial to incorporate additional satellite‐based covariates. To address these challenges for spatio‐temporal interpolation and forecasting aims, novel identifiable variational autoencoder (iVAE)‐based methods are introduced that explicitly integrate satellite‐derived predictors and handle missing values within a nonlinear blind source separation framework. The proposed extensions preserve the identifiability guarantees of the iVAE framework under the missing at random assumption. Performance is benchmarked against established statistical and deep learning approaches by using daily average ozone concentrations in Northern Italy: in interpolation, the proposed method achieves an appreciable reduction of the estimation errors and in forecasting it shows 10 times less training time with respect to the best competing method. The results establish nonlinear blind source separation as a promising approach for spatio‐temporal prediction of environmental data.
Model-based ordination of ecological community data has gained recently significant popularity among practitioners, largely due to increased availability and utilization of computational resources. Specifically, generalized linear latent variable models (GLLVMs)–a factor-analytic and rank-reduced form of mixed effect models–have proven to be both accurate and computationally efficient. GLLVMs have been implemented for a wide range of response types common to ecological community data; presence-absence, biomass, overdispersed and/or zero-inflated counts serving as examples. In this paper, we demonstrate how GLLVMs can be applied in the analysis of high-dimensional compositional count data. These methods are useful for example in the analysis of microbiome data, which are typically collected using modern lab-based sampling tools and are inherently compositional due to the finite capacity of sequencing instruments. We use simulation studies to compare the ordination methods based on GLLVMs with algorithmic compositional data analysis methods that rely on log-transformations. Also recently developed fast model-based ordination methods that utilize Gaussian copula models are included in our comparisons. The methods are illustrated with a microbiome data example.
Genomic prediction models, such as genomic best linear unbiased prediction (GBLUP), use genomic data to improve the accuracy of estimated genetic values. As the number of genotypes and traits increases, the exact calculation of prediction error variances (PEVs) and reliabilities becomes computationally infeasible due to the need to invert the coefficient matrix of the mixed model equations, whose dimension increases directly with the number of individuals and traits. The objective of this study was to evaluate the applicability of the Monte Carlo (MC) sampling method to approximate PEVs and reliabilities in a multiple-trait GBLUP framework relevant to hybrid breeding. The MC method avoids direct matrix inversion by repeatedly sampling genetic values from their assumed distributions to approximate PEVs. We applied the MC method using four previously published formulas to approximate PEVs and reliabilities. All formulas produced consistent estimates of PEVs and reliabilities, with convergence rates depending on the formula, the level of reliability, and the MC sample size.
Background Over the past decade, joint species distribution models (JSDMs) and model-based ordination have emerged as powerful tools for the analysis of community ecology data. Generalized linear latent variable models (GLLVMs) offer a flexible framework for multivariate analysis of a wide range of data types, based on including a small number of latent variables to perform dimension reduction while accounting for residual correlation between species. Fast estimation methods The R package gllvm implements a wide range of GLLVMs, with estimation performed via fast approximate likelihood-based techniques; including the recently proposed extended variational approximation, which is applicable to almost any combination of response type and link function. Since its original development and accompanying software paper, the gllvm package has undergone a significant overhaul, consolidating its place as a general framework for joint modeling of community ecology datasets. Expanded functionalities Some of the key new features of gllvm include model-based constrained and concurrent ordination methods, capacity to account for nested/hierarchical sampling designs, and (phylogenetic) random effects. On top of this, other notable improvements include a great expansion of the response types that it can handle, enhanced capabilities of GLLVM inference, selection and prediction, and an easier-to-use interface for model fitting.
The modeling and prediction of multivariate spatio-temporal data involve numerous challenges. Dimension reduction methods can significantly simplify this process, provided that they account for the complex dependencies between variables and across time and space. Nonlinear blind source separation has emerged as a promising approach, particularly following recent advances in identifiability results. Building on these developments, we introduce the identifiable autoregressive variational autoencoder, which ensures the identifiability of latent components consisting of nonstationary autoregressive processes. The blind source separation efficacy of the proposed method is showcased through a simulation study, where it is compared against state-of-the-art methods, and the spatio-temporal prediction performance is evaluated against several competitors on air pollution and weather datasets.
Monitoring performance-related characteristics of athletes can reveal changes that facilitate training adaptations. Here, we examine the relationships between submaximal running, maximal jump performance (CMJ), concentrations of blood lactate, sleep duration (SD) and latency (SL), and perceived stress (PSS) in junior cross-country skiers during pre-season training. These parameters were monitored in 15 male and 14 females (17 +/- 1 years) for the 12-weeks prior to the competition season, and the data was analysed using linear and mixed-effect models. An increase in SD exerted a decrease in both PSS (B = -2.79, p <= 0.01) and blood lactate concentrations during submaximal running (B = -0.623, p <= 0.05). In addition, there was a negative relationship between SL and CMJ (B = -0.09, p = 0.08). Compared to males, females exhibited higher PSS scores and little or no change in performance-related tests. A significant interaction between time and sex was present in CMJ with males displaying an effect of time on CMJ performance. For all athletes, lower PSS appeared to be associated with longer overnight sleep. Since the females experienced higher levels of stress, monitoring of their PSS might be beneficial. These findings have implications for the preparation of young athletes' competition season.
Modelling multivariate spatio-temporal data with complex dependency structures is a challenging task but can be simplified by assuming that the original variables are generated from independent latent components. If these components are found, they can be modelled univariately. Blind source separation aims to recover the latent components by estimating the unknown linear or nonlinear unmixing transformation based on the observed data only. In this paper, we extend recently introduced identifiable variational autoencoder to the nonlinear nonstationary spatio-temporal blind source separation setting and demonstrate its performance using comprehensive simulation studies. Additionally, we introduce two alternative methods for the latent dimension estimation, which is a crucial task in order to obtain the correct latent representation. Finally, we illustrate the proposed methods using a meteorological application, where we estimate the latent dimension and the latent components, interpret the components, and show how nonstationarity can be accounted and prediction accuracy can be improved by using the proposed nonlinear blind source separation method as a preprocessing method.
1. Joint species distribution models (JSDMs) have gained considerable traction among ecologists over the past decade, due to their capacity to answer a wide range of questions at both the species- and the community-level. The family of generalized linear latent variable models in particular has proven popular for building JSDMs, being able to handle many response types including presence-absence data, biomass, overdispersed and/or zero-inflated counts. 2. We extend latent variable models to handle percent cover data, with vegetation, sessile invertebrate, and macroalgal cover data representing the prime examples of such data arising in community ecology. 3. Sparsity is a commonly encountered challenge with percent cover data. Responses are typically recorded as percentages covered per plot, though some species may be completely absent or present, i.e., have 0% or 100% cover respectively, rendering the use of beta distribution inadequate. 4. We propose two JSDMs suitable for percent cover data, namely a hurdle beta model and an ordered beta model. We compare the two proposed approaches to a beta distribution for shifted responses, transformed presence-absence data, and an ordinal model for percent cover classes. Results demonstrate the hurdle beta JSDM was generally the most accurate at retrieving the latent variables and predicting ecological percent cover data.
In stationary subspace analysis (SSA) one assumes that the observable p-variate time series is a linear mixture of a k-variate nonstationary time series and a (p−k)-variate stationary time series. The aim is then to estimate the unmixing matrix which transforms the observed multivariate time series onto stationary and nonstationary components. In the classical approach multivariate data are projected onto stationary and nonstationary subspaces by minimizing a Kullback–Leibler divergence between Gaussian distributions, and the method only detects nonstationarities in the first two moments. In this paper we consider SSA in a more general multivariate time series setting and propose SSA methods which are able to detect nonstationarities in mean, variance and autocorrelation, or in all of them. Simulation studies illustrate the performances of proposed methods, and it is shown that especially the method that detects all three types of nonstationarities performs well in various time series settings. The paper is concluded with an illustrative example.
In spatial blind source separation the observed multivariate random fields are assumed to be mixtures of latent spatially dependent random fields. The objective is to recover latent random fields by estimating the unmixing transformation. Currently, the algorithms for spatial blind source separation can only estimate linear unmixing transformations. Nonlinear blind source separation methods for spatial data are scarce. In this paper we extend an identifiable variational autoencoder that can estimate nonlinear unmixing transformations to spatially dependent data and demonstrate its performance for both stationary and nonstationary spatial data using simulations. In addition, we introduce scaled mean absolute Shapley additive explanations for interpreting the latent components through nonlinear mixing transformation. The spatial identifiable variational autoencoder is applied to a geochemical dataset to find the latent random fields, which are then interpreted by using the scaled mean absolute Shapley additive explanations. Finally, we illustrate how the proposed method can be used as a pre-processing method when making multivariate predictions.
Ecosystem restoration will increase following the ambitious international targets, which calls for a rigorous evaluation of restoration effectiveness. Here, we present results from a long-term before-after control-impact experiment on the restoration of forestry-drained boreal peatland ecosystems. Our data comprise 151 sites, representing six ecosystem types. Species-level vegetation sampling has been conducted before, two, five, and ten years after restoration. With joint species distribution modelling, we show that, on average, not restoring leads to further degradation, but restoration stops and reverses this trend. The variation in restoration outcomes largely arises from ecosystem types: restoration of nutrient-poor ecosystems has a higher probability of failure. Yet, the ten-year study period is insufficient to capture the restoration effects in slow-recovering ecosystems. Altogether, restoration can effectively halt the biodiversity loss of degraded ecosystems, although ecosystem attributes affect the outcome. This variability in outcomes underlies the need for evidence-based prioritization of restoration efforts across ecosystems. Restoration halts and reverses degradation of boreal peatlands in nutrient-rich ecosystems, though the impact may be weak in nutrient-poor ones, according to a long-term experiment in Finland comprising 151 sites and 6 ecosystem types
Animals host complex bacterial communities in their gastrointestinal tracts, with which they share a mutualistic interaction. The numerous effects these interactions grant to the host include regulation of the immune system, defense against pathogen invasion, digestion of otherwise undigestible foodstuffs, and impacts on host behaviour. Exposure to stressors, such as environmental pollution, parasites, and/or predators, can alter the composition of the gut microbiome, potentially affecting host-microbiome interactions that can be manifest in the host as, for example, metabolic dysfunction or inflammation. However, whether a change in gut microbiota in wild animals associates with a change in host condition is seldom examined. Thus, we quantified whether wild bank voles inhabiting a polluted environment, areas where there are environmental radionuclides, exhibited a change in gut microbiota (using 16S amplicon sequencing) and concomitant change in host health using a combined approach of transcriptomics, histological staining analyses of colon tissue, and quantification of short-chain fatty acids in faeces and blood. Concomitant with a change in gut microbiota in animals inhabiting contaminated areas, we found evidence of poor gut health in the host, such as hypotrophy of goblet cells and likely weakened mucus layer and related changes in Clca1 and Agr2 gene expression, but no visible inflammation in colon tissue. Through this case study we show that inhabiting a polluted environment can have wide reaching effects on the gut health of affected animals, and that gut health and other host health parameters should be examined together with gut microbiota in ecotoxicological studies.
Generalized linear latent variable models (GLLVMs) have become mainstream models in this analysis of correlated, m‐dimensional data. GLLVMs can be seen as a reduced‐rank version of generalized linear mixed models (GLMMs) as the latent variables which are of dimension p≪m$$ p\ll m $$ induce a reduced‐rank covariance structure for the model. Models are flexible and can be used for various purposes, including exploratory analysis, that is, ordination analysis, estimating patterns of residual correlation, multivariate inference about measured predictors, and prediction. Recent advances in computational tools allow the development of efficient, scalable algorithms for fitting GLLMVs for any response distribution. In this article, we discuss the basics of GLLVMs and review some options for model fitting. We focus on methods that are based on likelihood inference. The implementations available in R are compared via simulation studies and an example illustrates how GLLVMs can be applied as an exploratory tool in the analysis of data from community ecology.
ABSTRACT Ecosystem restoration will increase following the ambitious international targets, which calls for a rigorous evaluation of restoration effectiveness. Studies addressing restoration effectiveness across ecosystems have thus far shown varying and unpredictable patterns. A rigorous assessment of the factors influencing restoration effectiveness is best done with large-scale and long-term experimental data. Here, we present results from a well replicated long-term before-after control-impact experiment on restoration of forestry-drained boreal peatland ecosystems. Our data comprise 151 sites, representing six ecosystem types. Vegetation sampling has been conducted to the species level before restoration and two, five and ten years after restoration. We show that, on average, restoration stops and reverses the trend of further degradation. The variation in restoration outcomes largely arises from ecosystem types: restoration of nutrient-poor ecosystems has higher probability of failure. Our experiment provides clear evidence that restoration can be effective in halting the biodiversity loss of degraded ecosystems, although ecosystem attributes can affect the restoration outcome. These findings underlie the need for evidence-based prioritization of restoration efforts across ecosystems.
Consider a spatial blind source separation model in which the observed multivariate spatial data are assumed to be a linear mixture of latent stationary spatially uncorrelated random fields. The objective is to recover an unknown mixing procedure as well as the latent random fields. Recently, spatial blind source separation methods that are based on the simultaneous diagonalization of two or more scatter matrices were proposed. In cases involving uncontaminated data, such methods can solve the blind source separation problem, however, in the presence of outlying observations, these methods perform poorly. We propose a robust blind source separation method that employs robust global and local covariance matrices based on generalized spatial signs in simultaneous diagonalization. Simulation studies are employed to illustrate the robustness and efficiency of the proposed methods in various scenarios.