Summary This article focuses on covariance estimation for multi-view data. Popular approaches rely on factor-analytic decompositions that have shared and view-specific latent factors. Posterior computation is conducted via expensive and brittle Markov chain Monte Carlo (MCMC) sampling or variational approximations that underestimate uncertainty and lack theoretical guarantees. Our proposed methodology employs spectral decompositions to estimate and align latent factors that are active in at least one view. Conditionally on these factors, we choose jointly conjugate prior distributions for factor loadings and residual variances. The resulting posterior is a simple product of normal-inverse gamma distributions for each variable, bypassing MCMC and facilitating posterior computation. We prove favorable increasing-dimension asymptotic properties, including posterior contraction and central limit theorems for point estimators. We show excellent performance in simulations, including accurate uncertainty quantification, and apply the methodology to integrate four high-dimensional views from a multi-omics dataset of cancer cell samples.
It is often of interest to infer lower-dimensional structure underlying complex data. As a flexible class of nonlinear structures, it is common to focus on Riemannian manifolds. Most existing manifold-learning algorithms replace the original data with lower-dimensional coordinates without providing an estimate of the manifold or using it to denoise the original data. This article proposes a new methodology to address these issues, allowing interpolation of the estimated manifold between the fitted data points. The proposed approach is motivated by the novel theoretical properties of local covariance matrices constructed from samples near a manifold. Our results enable the transformation of a global manifold-reconstruction problem into a local regression problem, allowing the application of Gaussian processes for probabilistic manifold reconstruction. In addition to the theory justifying our methodology, we provide simulated and real data examples to illustrate its performance.
This article is motivated by challenges in conducting Bayesian inferences on unknown discrete distributions, with a particular focus on count data. To avoid the computational disadvantages of traditional mixture models, we develop a novel Bayesian predictive approach. In particular, our Metropolis-adjusted Dirichlet (MAD) sequence model characterizes the predictive measure as a mixture of a base measure and Metropolis-Hastings kernels centered on previous data points. The resulting MAD sequence is asymptotically exchangeable and the posterior on the data generator takes the form of a martingale posterior. This structure leads to straightforward algorithms for inference on count distributions, with easy extensions to multivariate, regression, and binary data cases. We obtain a useful asymptotic Gaussian approximation and illustrate the methodology on a variety of applications.
Recent advances in Passive Acoustic Monitoring (PAM) offer an opportunity to obtain ecological spatial point-process data at unprecedented scale. However, realizing this opportunity necessitates the development of accurate and scalable localization methods. In real-world outdoor soundscapes, however, the assumptions underlying classical localization methods such as hyperbolic and score-based localization are routinely violated by multipath dominance, near-field effects, and complex propagation. Under these conditions, classical localization methods become brittle, with extreme errors possible even in small detection arrays. Rather than statistically replacing the underlying physics, we propose a method to refine it and increase robustness outside of ideal operating conditions: a learned model operating on physics-informed acoustic features corrects a fast hyperbolic solver where it produces implausible solutions, substantially reducing catastrophic worst-case errors while matching its median accuracy on field data. We further provide calibrated, geometry-aware uncertainty estimates suitable for propagation into downstream spatial models. Evaluating on distributed microphone arrays in real and simulated outdoor environments, we demonstrate that the proposed method yields robust, uncertainty-aware localization, providing a step toward scalable automated wildlife monitoring in complex acoustic environments.
It is increasingly common to collect data of multiple different types on the same set of samples. Our focus is on studying relationships between such multiview features and responses. A motivating application arises in the context of precision medicine where multiomics data are collected to correlate with clinical outcomes. It is of interest to infer dependence within and across views while combining multimodal information to improve the prediction of outcomes. The signal-to-noise ratio can vary substantially across views, motivating more nuanced statistical tools beyond standard late and early fusion. This challenge comes with the need to preserve interpretability, select features, and obtain accurate uncertainty quantification. To address these challenges, we introduce two complementary factor regression models. A baseline joint factor regression (jfr) captures combined variation across views via a single factor set, and a more nuanced Joint Additive FActor Regression (jafar) that decomposes variation into shared and view-specific components. For JFR, we use independent cumulative shrinkage process (I-CUSP) priors, while for JAFAR, we develop a dependent version (D-CUSP) designed to ensure identifiability of the components. We develop Gibbs samplers that exploit the model structure and accommodate flexible feature and outcome distributions. Prediction of time-to-labor onset from immunome, metabolome, and proteome data illustrates performance gains against state-of-the-art competitors.
In including random effects to account for dependent observations, the odds ratio interpretation of logistic regression coefficients is changed from population-averaged to subject-specific. This is unappealing in many applications, motivating a rich literature on methods that maintain the marginal logistic regression structure without random effects, such as generalized estimating equations. However, for spatial data, random effect approaches are appealing in providing a full probabilistic characterization of the data that can be used for prediction. We propose a new class of spatial logistic regression models that maintain both population-averaged and subject-specific interpretations through a novel class of bridge processes for spatial random effects. These processes are shown to have appealing computational and theoretical properties, including a scale mixture of normal representation. The new methodology is illustrated with simulations and an analysis of childhood malaria prevalence data in Gambia.
models often rely on the forward-backward sampler. This makes them computationally slow as the length of the time series increases, motivating the development of sub-sampling-based approaches. These approximate the full posterior by using small random subsequences of the data at each MCMC iteration within stochastic gradient MCMC. In the presence of imbalanced data resulting from rare latent states, subsequences often exclude rare latent state data, leading to inaccurate inference and prediction/detection of rare events. We propose a targeted subsampling (TASS) approach that over-samples observations corresponding to rare latent states when calculating the stochastic gradient of parameters associated with them. TASS uses an initial clustering of the data to construct subsequence weights that reduce the variance in gradient estimation. This leads to improved sampling efficiency, in particular in settings where the rare latent states correspond to extreme observations. We demonstrate substantial gains in predictive and inferential accuracy on real and synthetic examples.
BACKGROUND: Human exposure to complex, changing, and variably correlated mixtures of environmental chemicals has presented analytical challenges to epidemiologists and human health researchers. There has been a wide variety of recent advances in statistical methods for analyzing mixtures data, with most methods having open-source software for implementation. However, there is no one-size-fits-all method for analyzing mixture data given the considerable heterogeneity in scientific focus and study design. For example, some methods focus on predicting the overall health effect of a mixture and others seek to disentangle main effects and pairwise interactions. Some methods are only appropriate for cross-sectional designs, while other methods can accommodate longitudinally measured exposures or outcomes. OBJECTIVES: This article focuses on simplifying the task of identifying which methods are most appropriate to a particular study design, data type, and scientific focus. METHODS: We present an organized workflow for statistical analysis considerations in environmental mixtures data and two example applications implementing the workflow. This systematic strategy builds on epidemiological and statistical principles, considering specific nuances for the mixtures' context. We also present an accompanying methods repository to increase awareness of and inform application of existing methods and new methods as they are developed. DISCUSSION: We note several methods may be equally appropriate for a specific context. This article does not present a comparison or contrast of methods or recommend one method over another. Rather, the presented workflow can be used to identify a set of methods that are appropriate for a given application. Accordingly, this effort will inform application, educate researchers (e.g., new researchers or trainees), and identify research gaps in statistical methods for environmental mixtures that warrant further development.
Abstract Understanding global biodiversity patterns and their drivers is a prerequisite for countering the biodiversity crisis. In this paper, we introduce a novel generalized linear model, Hubbell regression, to estimate a key biodiversity descriptor, the fundamental biodiversity number. This can be converted into a set of biodiversity descriptors, including Shannon and Simpson indices, and more. Hence, quantifying the impact of environmental conditions on the fundamental biodiversity number allows us to predict the general properties of local biodiversity in any setting. In addition to having a strong mathematical foundation, Hubbell regression consistently outperformed current state-of-the-art models in predicting global biodiversity. We apply the method to arthropods, which account for the majority of terrestrial biodiversity. By parameterizing the models using samples of 1.78 million arthropods from 2415 samples collected at 135 sites spanning all continents, we pinpoint the drivers of arthropod biodiversity and its features at the global scale. We find that actual evapotranspiration is the single largest predictor of arthropod diversity and explains nearly 30% of the variation in richness. Moreover, we infer that high human activity has led to a 21.3 % and 29.2% decrease in potential insect richness in tropical and dry zones, respectively, but increased insect richness in polar regions. These insights bring a new foundation for biodiversity research and action.
Many datasets include a small set of variables, such as biomarkers or clinical outcomes, whose relationships to the broader system are of primary scientific interest. Estimating the full network of inter-variable relationships in such settings often obscures local structures around these targets, limiting interpretability. To address this fundamental problem, we introduce local graph estimation, a statistical framework for inferring substructures around target variables. We show that traditional graph estimation methods often fail to recover local structure, and present pathwise feature selection (PFS) as an effective alternative. PFS estimates local subgraphs by iteratively applying feature selection and propagating uncertainty along network paths, providing rigorous finite-sample false discovery control even in settings with mixed variable types and nonlinear dependencies. In four distinct applications spanning environmental and public health, multiomics, brain connectomics, and single-nucleus RNA sequencing, PFS recovers interpretable networks consistent with domain knowledge, highlighting its ability to uncover established mechanisms and generate novel hypotheses.
Beta regression is used routinely for continuous proportional data, but it often encounters practical issues such as a lack of robustness of regression parameter estimates to misspecification of the beta distribution. We develop an improved class of generalized linear models starting with the continuous binomial (cobin) distribution and further extending to dispersion mixtures of cobin distributions (micobin). The proposed cobin regression and micobin regression models have attractive robustness, computation, and flexibility properties. A key innovation is the Kolmogorov-Gamma data augmentation scheme, which facilitates Gibbs sampling for Bayesian computation, including in hierarchical cases involving nested, longitudinal, or spatial data. We demonstrate robustness, ability to handle responses exactly at the boundary (0 or 1), and computational efficiency relative to beta regression in simulation experiments and through analysis of the benthic macroinvertebrate multimetric index of US lakes using lake watershed covariates.
Citizen science provides large amounts of biodiversity data. Key challenges in unlocking its full potential include engaging citizens with limited species identification skills and accelerating the transition from data collection to research and monitoring outputs. Here we use a large dataset from Finland to show how even citizens who cannot identify birds themselves can contribute to real-time predictions of avian distributions. This is achieved through a digital twin that combines smartphone-based citizen science with long-term knowledge in a continuously updating model. The app submits raw audio to a backend that classifies birds with machine learning, reducing variation in data quality and enabling validation and reclassification by continuously improving classifiers. We counteracted spatiotemporal sampling biases by interval recordings and permanent point count networks. Over 2 years, the app generated 15 million bird detections. Independent test data show that the digital-twin-informed models are more accurate at predicting bird spatiotemporal distributions. Because our approach is highly scalable and has the potential to generate biomonitoring data even in understudied areas, it could accelerate the flow of reliable biodiversity information and increase inclusivity in citizen science projects. Citizen science data are increasingly used in biodiversity monitoring. This study applies a digital twin approach to biodiversity monitoring using a large citizen science dataset on birds from Finland, demonstrating its potential for ecological forecasting.
Gibbs posteriors are proportional to a prior distribution multiplied by an exponentiated loss function, with a key tuning parameter weighting information in the loss relative to the prior and providing a control of posterior uncertainty. Gibbs posteriors provide a principled framework for likelihood-free Bayesian inference, but in many situations, including a single tuning parameter inevitably leads to poor uncertainty quantification. In particular, regardless of the value of the parameter, credible regions have far from the nominal frequentist coverage even in large samples. We propose a sequential extension to Gibbs posteriors to address this problem. We prove the proposed sequential posterior exhibits concentration and a Bernstein-von Mises theorem, which holds under easy to verify conditions in Euclidean space and on manifolds. As a byproduct, we obtain the first Bernstein-von Mises theorem for traditional likelihood-based Bayesian posteriors on manifolds. All methods are illustrated with an application to principal component analysis.
Abstract Ecological networks offer powerful insights into community function but, without first characterizing these networks accurately, our ability to detect and interpret changes under environmental stress is limited. We develop an extension of the Covariate‐Informed Link Prediction (COIL) framework to reduce bias in ecological link prediction when interaction data are derived from studies focused on a small number of species. Because the absence of an observed interaction is only informative when the species involved actually co‐occur, uncertainty in species occurrence can strongly influence how non‐interactions are interpreted in meta‐network datasets. Our extended model (COIL+) employs a latent factor structure that borrows information across species, incorporates species traits and phylogeny, and integrates observations from multiple studies to account for uncertainty in species occurrence under extreme taxonomic bias (i.e. studies focussing on only a small subset of species). We additionally introduce a trait‐matching procedure that quantifies how the influence of species traits on interaction probability varies across partner species. We illustrate the use of the model with a literature‐based dataset of 268 sources reporting Afrotropical frugivory and compare performance with and without correction for occurrence uncertainty. COIL+ substantially improves link prediction by increasing out‐of‐sample discrimination and reduces sampling bias, revealing 5637 likely but unobserved frugivory interactions (a median of nine additional interactions per frugivore). Newly predicted interactions are concentrated among poorly sampled frugivores, such as the water chevrotain (Hyemoschus aquaticus, a small forest‐dwelling ungulate) and the rufous‐bellied helmetshrike (Prionops rufiventris, a passerine bird of East African tropical forests). Additionally, the method better distinguishes true interactions from non‐interactions compared to existing approaches under strong taxonomic bias and narrow study focus. This framework generalizes to diverse meta‐network contexts and provides a useful tool for link prediction in the face of biased interaction data.
Beta diversity quantifies variation in species composition across ecological communities and is fundamental for understanding biodiversity patterns across space and environmental gradients. Statistical inference on beta diversity is challenging: species occurrence data are high dimensional, many species remain unobserved despite extensive sampling, and surveys are subject to imperfect detection. Existing approaches are typically based on empirical dissimilarity indices with limited uncertainty quantification or on models that rely on unrealistic exchangeability and perfect detection assumptions. We introduce a new class of Bayesian feature allocation models for partially exchangeable species occurrence data with imperfect detection. The framework combines latent feature allocation models with occupancy-based detection mechanisms, allowing heterogeneous species compositions across sites, explicitly accounting for false negatives, and accommodating the discovery of previously unobserved species. We develop coherent probabilistic inference for species sharing and beta diversity, deriving explicit posterior and predictive distributions for between-community heterogeneity, including the number of shared species across sites and the number expected under future sampling. These analytical results yield interpretable posterior summaries of compositional heterogeneity and facilitate scalable inference in high-dimensional biodiversity studies. Simulation studies and an application to global fungal biodiversity data demonstrate improved inference on species sharing and between-community diversity.
Factor models are popular approaches for analyzing high-dimensional data to extract low-rank signals and estimate covariances. They decompose the covariance matrix as the sum of low-rank and diagonal components. A key issue is how to choose the latent dimension k, which is particularly challenging when the factor model only holds approximately and in low signal-to-noise scenarios. Bayesian overfitted factor models specify an upper bound on k and rely on structured shrinkage priors to effectively remove extra components. Such approaches are popular and effective, but computationally expensive. We propose a much faster approach that provides valid uncertainty quantification, based on spectral estimation of latent factors and adaptive empirical Bayes calibration of key hyperparameters. The resulting posterior distribution factorizes across outcomes and is analytically tractable, bypassing Markov chain Monte Carlo. We show that adapts to the signal-to-noise ratio of each outcome and latent dimension, while shrinking superfluous latent components to zero. We establish favorable asymptotic properties and demonstrate strong empirical performance in numerical experiments and a genomics application, where EigenBayes outperforms state-of-the-art alternatives.
Dirichlet process mixtures are particularly sensitive to the value of the so-called precision parameter, which controls the behavior of the underlying latent partition. Randomization of the precision through a prior distribution is a common solution, which leads to more robust inferential procedures. However, existing prior choices do not allow for transparent elicitation, due to the lack of analytical results. We introduce and investigate a novel prior for the Dirichlet process precision, the Stirling-gamma distribution. We study the distributional properties of the induced random partition, with an emphasis on the number of clusters. Our theoretical investigation clarifies the reasons of the improved robustness properties of the proposed prior. Moreover, we show that, under specific choices of its hyperparameters, the Stirling-gamma distribution is conjugate to the random partition of a Dirichlet process. We illustrate with an ecological application the usefulness of our approach for the detection of communities of ant workers.
We study how the posterior contraction rate under a Gaussian process (GP) prior depends on the intrinsic dimension of the predictors and the smoothness of the regression function. An open question is whether a generic GP prior that does not incorporate knowledge of the intrinsic lower-dimensional structure of the predictors can attain an adaptive rate for a broad class of such structures. We show that this is indeed the case, establishing conditions under which the posterior contraction rates become adaptive to the intrinsic dimension in terms of the covering number of the data domain (the Minkowski dimension) and prove the nonparametric posterior contraction rate, up to a logarithmic factor. When the domain is a compact manifold, we prove the RKHS approximation to intrinsically defined H & ouml;lder functions on the manifold of any order of smoothness by a novel analysis, leading to the optimal adaptive posterior contraction rate. We propose an empirical Bayes prior on the kernel bandwidth using kernel affinity and k-nearest neighbor statistics, bypassing explicit estimation of the intrinsic dimension. The efficiency of the proposed Bayesian regression approach is demonstrated in various numerical experiments.
Feature selection is a critical task in machine learning and statistics. However, existing feature selection methods either (i) rely on parametric methods such as linear or generalized linear models, (ii) lack theoretical false discovery control, or (iii) identify few true positives. Here, we introduce a general feature selection method with finite-sample false discovery control based on applying integrated path stability selection (IPSS) to arbitrary feature importance scores. The method is nonparametric whenever the importance scores are nonparametric, and it estimates q-values, which are better suited to high-dimensional data than p-values. We focus on two special cases using importance scores from gradient boosting (IPSSGB) and random forests (IPSSRF). Extensive nonlinear simulations with RNA sequencing data show that both methods accurately control the false discovery rate and detect more true positives than existing methods. Both methods are also efficient, running in under 20 seconds when there are 500 samples and 5000 features. We apply IPSSGB and IPSSRF to detect microRNAs and genes related to cancer, finding that they yield better predictions with fewer features than existing approaches.
Our interest is in multiplex network data with multiple network samples observed across the same set of nodes. Examples originate from a variety of fields, including brain connectivity, international trade networks, and social networks, among others. Our goal is to infer a hierarchical structure of the nodes at a population level, while performing multi-resolution clustering of the individual replicates. To accomplish this, we propose a Bayesian hierarchical model, provide theoretical support in terms of identifiability and posterior consistency, and design efficient methods for posterior computation. We provide novel technical tools for proving model identifiability, which are of independent interest. Our proposed methodology is demonstrated through numerical simulation and an application to brain connectome data.