Abscission, the shedding of organismal parts, depends on physiological events whose optimal timing is crucial for species survival. Environmental variations impact species development, particularly abscission processes across multiple developmental stages, identifying which environmental factors modulate abscission and when is essential in the context of climate change. Treating environmental variables as time series-groups of temporally correlated variables raises statistical challenges for selecting relevant groups (environmental variables) and their correlated components (time periods). We address these objectives by introducing the Bayesian fused and fusion priors through a general parameterization. We highlight a trade-off between priors used on differences versus coefficients, demonstrating that horseshoe-type priors on both differences and coefficients, with appropriate parameterizations, achieve effective selection, estimation and algorithmic stability regardless of group number or size. Our study focuses on fruit abscission in oil palms, which affects bunch harvest timing. Abscission disruption can impact oil yield and quality, consequently affecting economic returns. This application, based on experimental data from Benin, illustrates how our proposed priors successfully select both environmental variables and developmental stages involved in bunch harvest timing.
The link function is the key component of regression models for binary response variables. Despite the diverse potential fits obtained from different link functions, only the logit and the probit links have been widely popularized. Maximum likelihood estimations in models generated from these links are known to be non-robust in the presence of outliers. We show that this problem is exacerbated when the two response levels are strongly separated in the explanatory space. To address this shortcoming, we propose and encourage the use of the maximum likelihood estimation with the Student link function. We highlight its robustness to outliers and also to noisy variables, particularly when the data exhibit a strong separation setting, still keeping all the maximum likelihood estimation's properties.
In statistical modeling, there is a wide variety of generalized linear models for categorical response variables (nominal or ordinal responses); yet, there is no software embracing all these models together in a unique and generic framework. We propose and present GLMcat, an R package to estimate generalized linear models implemented under the unified specification (r, F, Z) where r represents the ratio of probabilities (reference, cumulative, adjacent, or sequential), F the cumulative distribution function for the linkage, and Z the design matrix. All classical models (and their variations) for categorical data can be written as an (r, F, Z) triplet, thus, they can be fitted with GLMcat. The functions in the package are intuitive and user-friendly. For each of the three components, there are multiple alternatives from which the user should thoroughly select those that best address the objectives of the analysis. The main strengths of the GLMcat package are the possibility of choosing from a large number of link functions (defined by the composition of F and r) and the simplicity for setting constraints in the linear prediction, either on the intercepts or on the slopes. This paper proposes a methodological and practical guide for the appropriate selection of a model considering the concordance between the nature of the data and the properties of the model.
The influence of canopy structure on tropical tree growth has been scantly studied because of the difficulties making field measurements in these dense multi-layered ecosystems. The recent advent of unmanned aerial vehicles (UAVs), has made it easier to collect canopy data, so offering a way to gain a better understanding of forest productivity and thereby improve forest management. In this study, we assessed tree growth prediction using UAV-derived crown measurements as an alternative for field data. Four experimental 9 ha plots were sampled in two forest sites, Yoko in the Democratic Republic of the Congo and Loundoungou in the Republic of Congo. Field inventories were made between 2015 and 2020. For each tree, we computed the diameter increment (DBHI) using censuses and diameter-based competition indices (diameter-based CIs) using the first census. High-resolution orthoimages and digital surface models were acquired with UAVs in 2016 and 2018 in the two sites. They gave estimates of crown characteristics (size, relative elevation, shape) and crown-based competition indices (crown-based CIs). Co-recorded UAV and field measurements were obtained for 1558 trees. The diameter increment of these trees was then modelled using supervised component generalized linear regression, and 20 % of trees were kept for cross-validation. Combined field and UAV data predicted tree DBHI twice better than either taken separately. Diameter at breast height (DBH) and crown area (CA) were found to be complementary predictors. Crown-based CIs significantly improved predictions of models already containing DBH and CA. Adding diameter-based CIs to models containing DBH, CA, and crown-based CIs only marginally improved growth predictions, showing that tree competition can be well-described with UAV data. The model calibrated at one site predicted the growth at the other site well, suggesting that a general model could be devised for multiple sites. Growth variance was better explained in the site (Yoko) where the crown density was higher and the crown smaller. Further data are now needed from multiple sites with ranging stand structures and compositions to build a general model.
In a context of component-based multivariate modeling we propose to model the residual dependence of the responses. Each response of a response vector is assumed to depend, through a Generalized Linear Model, on a set of explanatory variables. The vast majority of explanatory variables are partitioned into conceptually homogeneous variable groups, viewed as explanatory themes. Variables in themes are supposed many and some of them are highly correlated or even collinear. Thus, generalized linear regression demands dimension reduction and regularization with respect to each theme. Besides them, we consider a small set of "additional" covariates not conceptually linked to the themes, and demanding no regularization. Supervised Component Generalized Linear Regression proposed to both regularize and reduce the dimension of the explanatory space by searching each theme for an appropriate number of orthogonal components, which both contribute to predict the responses and capture relevant structural information in themes. In this paper, we introduce random latent variables (a.k.a. factors) so as to model the covariance matrix of the linear predictors of the responses conditional on the components. To estimate the model, we present an algorithm combining supervised component-based model estimation with factor model estimation. This methodology is tested on simulated data and then applied to an agricultural ecology dataset.
In this article, we propose to cluster responses in order to identify groups predicted by specific explanatory components. A response matrix is assumed to depend on a set of explanatory variables and a set of additional covariates. Explanatory variables are supposed many and redundant, which implies some dimension reduction and regularization. By contrast, additional covariates contain few selected variables which are forced into the regression model, as they demand no regularization. The response matrix is assumed partitioned into several unknown groups of responses. We suppose that the responses in each group are predictable from an appropriate number of specific orthogonal supervised components of explanatory variables. The classification is based on a mixture model of the responses. To estimate the model, we propose a criterion extending that of Supervised Component-based Generalized Linear Regression, a Partial Least Squares-type method, and develop an algorithm combining component-based model and Expectation Maximization estimation. This new methodology is tested on simulated data and then applied to a floristic ecology dataset.
Comprendre l’influence des facteurs environnementaux sur la distribution des communautés d’espèces est crucial face aux changements globaux. Pour surmonter les limitations des méthodes de régression usuelles (réponse multivariée et/ou covariables redondantes), nous proposons la régression linéaire généralisée sur composantes supervisées. Les composantes construites prédisent au mieux la distribution de l’ensemble des espèces, tout en synthétisant l’information contenue dans les covariables.
Here, using an dataset of 6 million trees in more than 180,000 field plots, we jointly model the distribution in abundance of the most dominant central African tree taxa and produce the first continuous maps of the floristic and functional composition of central African forests. Our results show that the uncertainty in taxon-specific distributions averages out at the community level, revealing highly deterministic assemblages. We uncover contrasting floristic and functional compositions across climate, soil types and anthropogenic gradients, with functional convergence among floristically dissimilar forest types. Combining these spatial predictions with global change scenarios suggests a high vulnerability of the northern and southern forest margins, the Atlantic forests and of most forests from the Democratic Republic of Congo where both climate and anthropogenic threats are expected to increase sharply by 2085. These results constitute key quantitative benchmarks for scientists and policy makers to shape transnational conservation and management strategies aiming at providing a sustainable future for central African forests. These models were generated using change-factor downscaling approaches to model spatial variation at local scales while correcting for differences between observed and simulated baseline climates (see Platts et al. 87 for more details). We here concentrated on one representative concentration pathway of the IPCC-AR5 (RCP 4.5) for the late 21st century (2071-2100, hereafter named 2085) and reconstructed the three SCGLR selected CCs from the climatic predictions as follows: let X r cp 4.5 be the predicted future climatic conditions. Let m = X and S = sd ( X ) be the mean and standard deviation matrices of the current climatic conditions. The predictive climatic components under future scenarios are then equal to f rc p 4.5 = ( X rc p 4.5 −m ) S ^ u , where ^ u represents SCGLR CCs. We then calculated the euclidean distance between the three current and the three predicted CCs for each of the 18 models and then estimated the exposure to climate change as the mean distance over the 18 models.
Africa is forecasted to experience large and rapid climate change 1 and population growth 2 during the twenty-first century, which threatens the world’s second largest rainforest. Protecting and sustainably managing these African forests requires an increased understanding of their compositional heterogeneity, the environmental drivers of forest composition and their vulnerability to ongoing changes. Here, using a very large dataset of 6 million trees in more than 180,000 field plots, we jointly model the distribution in abundance of the most dominant tree taxa in central Africa, and produce continuous maps of the floristic and functional composition of central African forests. Our results show that the uncertainty in taxon-specific distributions averages out at the community level, and reveal highly deterministic assemblages. We uncover contrasting floristic and functional compositions across climates, soil types and anthropogenic gradients, with functional convergence among types of forest that are floristically dissimilar. Combining these spatial predictions with scenarios of climatic and anthropogenic global change suggests a high vulnerability of the northern and southern forest margins, the Atlantic forests and most forests in the Democratic Republic of the Congo, where both climate and anthropogenic threats are expected to increase sharply by 2085. These results constitute key quantitative benchmarks for scientists and policymakers to shape transnational conservation and management strategies that aim to provide a sustainable future for central African forests.
In statistical modeling, there is a wide variety of regression models for categorical responses. Yet, no software encapsulates all of these models in a standardized format. We introduce and illustrate the utility of glmcat, the R package we developed to estimate generalized linear models implemented under the unified specification ( r, F, Z ), where r represents the ratio of probabilities (reference, cumulative, adjacent, or sequential), F the cumulative cdf function for the linkage
How does the genetic architecture of quantitative traits evolve over time? Answering this question is crucial for many applied fields such as human genetics and plant or animal breeding. In the last decades, high-throughput genome techniques have been used to better understand links between genetic information and quantitative traits. Recently, high-throughput phenotyping methods are also being used to provide huge information at a phenotypic scale. In particular, these methods allow traits to be measured over time, and this, for a large number of individuals. Combining both information might provide evidence on how genetic architecture evolves over time. However, such data raise new statistical challenges related to, among others, high dimensionality, time dependencies, time varying effects. In this work, we propose a Bayesian varying coefficient model allowing, in a single step, the identification of genetic markers involved in the variability of phenotypic traits and the estimation of their dynamic effects. We evaluate the use of spike-and-slab priors for the variable selection with either P-spline interpolation or non-functional techniques to model the dynamic effects. Numerical results are shown on simulations and on a functional mapping study performed on an Arabidopsis thaliana (L. Heynh) data which motivated these developments.
We address component-based regularization of a multivariate generalized linear model (GLM). A vector of random responses [Formula: see text] is assumed to depend, through a GLM, on a set [Formula: see text] of explanatory variables, as well as on a set [Formula: see text] of additional covariates. [Formula: see text] is partitioned into [Formula: see text] conceptually homogenous variable groups [Formula: see text], viewed as explanatory themes. Variables in each [Formula: see text] are assumed many and redundant. Thus, generalized linear regression demands dimension reduction and regularization with respect to each [Formula: see text]. By contrast, variables in [Formula: see text] are assumed few and selected so as to demand no regularization. Regularization is performed searching each [Formula: see text] for an appropriate number of orthogonal components that both contribute to model [Formula: see text] and capture relevant structural information in [Formula: see text]. To estimate a single-theme model, we first propose an enhanced version of Supervised Component Generalized Linear Regression (SCGLR), based on a flexible measure of structural relevance of components, and able to deal with mixed-type explanatory variables. Then, to estimate the multiple-theme model, we develop an algorithm encapsulating this enhanced SCGLR: THEME-SCGLR. The method is tested on simulated data and then applied to rainforest data in order to model the abundance of tree species.