Robust estimation of the covariance matrix and detection of outliers remain major challenges in statistical data analysis, particularly when the proportion of contaminated observations increases with the size of the dataset. Outliers can severely bias parameter estimates and induce a masking effect, whereby some outliers conceal the presence of other outliers, further complicating their detection. Although many approaches have been proposed for covariance estimation and outlier detection, to our knowledge, none of these methods have been implemented in an online setting. In this paper, we focus on online covariance matrix estimation and outlier detection. Specifically, we propose a new method for simultaneously and online estimating the geometric median and variance, which allows us to calculate the Mahalanobis distance for each incoming data point before deciding whether it should be considered an outlier. To mitigate the masking effect, robust estimation techniques for the mean and variance are required. Our approach uses the geometric median for robust estimation of the location and the median covariance matrix for robust estimation of the dispersion parameters. The new online methods proposed for parameter estimation and outlier detection allow real-time identification of outliers as data are observed sequentially. The performance of our methods is demonstrated on simulated datasets.
Abundance data are used in ecology for species monitoring and conservation. These count data often display several specific characteristics like numerous missing data, high variance, and a high proportion of zeros, particularly when monitoring rare species. We present a model that aims to impute missing data and estimate the effect of covariates on species presence and abundance. It is based on the log-normal Poisson model, which offers more flexibility in the variance of counts than a Poisson model. A latent variable is added for the overrepresentation of zeros in the data. The imputation of missing data is made possible by assuming that the latent variance matrix has low rank and the inclusion of covariates. We demonstrate the identifiability in the presence of missing data. Since maximum likelihood inference is intractable, we use a variational expectation-maximization algorithm to infer the parameters. We provide an estimate of the asymptotic variance of the estimators and derive prediction intervals for the imputations, an estimate of the temporal trend, and a procedure for detecting a potential change in this trend. We evaluate our imputations and associated prediction intervals using artificially degraded monitoring data set. We conclude with an illustration on a monitoring waterbirds data set.
We consider the problem of estimating a high-dimensional covariance matrix from a small number of observations when covariates on pairs of variables are available and the variables can have spatial structure. This is motivated by the problem arising in demography of estimating the covariance matrix of the total fertility rate (TFR) of 195 different countries when only 11 observations are available. We construct an estimator for high-dimensional covariance matrices by exploiting information about pairwise covariates, such as whether pairs of variables belong to the same cluster, or spatial structure of the variables, and interactions between the covariates. We reformulate the problem in terms of a mixed effects model. This requires the estimation of only a small number of parameters, which are easy to interpret and which can be selected using standard procedures. The estimator is consistent under general conditions, and asymptotically normal. It works if the mean and variance structure of the data is already specified or if some of the data are missing. We assess its performance under our model assumptions, as well as under model misspecification, using simulations. We find that it outperforms several popular alternatives. We apply it to the TFR dataset and draw some conclusions.
The aim of change-point detection is to identify behavioral shifts within time series data. This article focuses on scenarios where the data are derived from a heterogeneous Poisson process or a marked Poisson process. We present a methodology for detecting multiple offline change-points using a minimum contrast estimator. Specifically, we address how to manage the continuous nature of the process given the available discrete observations. Additionally, we select the appropriate number of changes via a cross-validation procedure, which is particularly effective given the characteristics of the Poisson process. Lastly, we show how to use this methodology for self-exciting processes with changes in the intensity. Through experiments, with both simulated and real datasets, we showcase the advantages of the proposed method, which has been implemented in the R package.
The robust estimation of the parameters of multivariate Gaussian linear regression models is considered by using robust versions of the usual (Mahalanobis) least-square criterion, with or without Ridge regularization. Two methods of estimation are introduced: (i) online stochastic gradient descent algorithms and their averaged variants, and (ii) offline fixed-point algorithms. These methods are applied to both the standard and Mahalanobis least-squares criteria, as well as to their regularized counterparts. Under weak assumptions, the resulting estimators are shown to be asymptotically normal. Since the noise covariance matrix is generally unknown, a robust estimate of this matrix is incorporated into the Mahalanobis-based stochastic gradient descent algorithms. Numerical experiments on synthetic data demonstrate a substantial gain in robustness compared with classical least-squares estimators, while also highlighting the computational efficiency of the online procedures. All proposed algorithms are implemented in the R package RobRegression, available on CRAN.
The Poisson log-normal model is a latent variable model that provides a generic framework for the analysis of multivariate count data. Inferring its parameters can be a daunting task since the conditional distribution of the latent variables given the observed ones is intractable. For this model, variational approaches are the golden standard solution as they prove to be computationally efficient but lack theoretical guarantees on the estimates. Sampling-based solutions are quite the opposite. We first define a Monte Carlo EM algorithm that can achieve maximum likelihood estimators, but that is computationally efficient only for low-dimensional latent spaces. We then propose a novel inference procedure combining the EM framework with composite likelihood and importance sampling estimates. The algorithm preserves the desirable asymptotic properties of maximum likelihood estimators while circumventing the high-dimensional integration bottleneck, thus maintaining computational feasibility for moderately large datasets. This approach enables grounded parameter estimation, confidence intervals, and hypothesis testing. Application to the Barents Sea fish dataset demonstrates the algorithm capacity to identify significant environmental effects and residual interspecies correlations.
We consider a broad class of random bipartite networks, the distribution of which is invariant under permutation within each type of nodes. We are interested in $U$-statistics defined on the adjacency matrix of such a network, for which we define a new type of Hoeffding decomposition. This decomposition enables us to characterize non-degenerate $U$-statistics -- which are then asymptotically normal -- and provides us with a natural and easy-to-implement estimator of their asymptotic variance. \\ We illustrate the use of this general approach on some typical random graph models and use it to estimate or test some quantities characterizing the topology of the associated network. We also assess the accuracy and the power of the proposed estimates or tests, via a simulation study.
This Package focuses on multivariate robust Guassian linear regression.We provide a function Robust_Mahalanobis_regression which enables to obtain robust estimates of the parameters of Multivariate Gaussian Linear Models with the help of the Mahalanobis distance, using a Stochastic Gradient algorithm or a Fix point.This is based on the function Robust_Variance which allows to obtain robust estimation of the variance, and so, also for low rank matrices (see Godichon-Baggioni and RObin (2024) )Robust methods for estimating the parameters of multivariate Gaussian linear models. .
RNA sample integrity variability introduces biases and obscures natural RNA degradation, posing a significant challenge in transcriptomics. To address this, we developed the Direct Transcriptome Integrity (DTI) measure, a universal and robust RNA integrity metric based on nanopore sequencing. By accurately modeling RNA fragmentation, DTI provides a reliable assessment of sample quality. Integrated into the INDEGRA package (freely available at https://github.com/Arnaroo/INDEGRA), we provide tools to correct false discoveries and enable precise differential expression and RNA degradation analyses, even for challenging sample types. INDEGRA software can be used to accurately measure RNA DTI stability metric, isolate biological component of RNA degradation from technical biases, compare biological RNA stability transcriptome-wide and suppress false degradation-induced differential gene expression hits to allow broad comparisons across samples of different quality DTI offers a straightforward and accurate method for assessing RNA degradation, characterizing both overall sample integrity and transcript-specific degradation rates using direct RNA sequencing (DRS) data. Calculated through INDEGRA, DTI reveals inter- and intra-transcript variability in degradation, while INDEGRA separates RNA degradation from mapping inaccuracies, and connects degradation profiles to RNA fragmentation rates. By leveraging INDEGRA, researchers can minimize false differential transcript abundance findings caused by variations in overall sample integrity, while preserving genuine transcript-specific differences in stability and degradation. INDEGRA supports integration with widely used differential transcript abundance tools like DESeq2, limma-voom, and edgeR, enabling seamless analysis pipelines. INDEGRA enhances the accuracy and reliability of RNA quantification in high-throughput data and simplifies comparisons across diverse transcriptomic datasets, including those derived from different tissues, species, or experimental protocols. ### Competing Interest Statement The authors have declared no competing interest.
Chapter 3 Evolutionary Models of Continuous Traits Paul BASTIDE, Paul BASTIDE IMAG, CNRS, Université de Montpellier, FranceSearch for more papers by this authorMahendra MARIADASSOU, Mahendra MARIADASSOU MaIAGE, INRAE, Université Paris-Saclay, Jouy-en-Josas, FranceSearch for more papers by this authorStéphane ROBIN, Stéphane ROBIN LPSM, Sorbonne Université, Paris, FranceSearch for more papers by this author Paul BASTIDE, Paul BASTIDE IMAG, CNRS, Université de Montpellier, FranceSearch for more papers by this authorMahendra MARIADASSOU, Mahendra MARIADASSOU MaIAGE, INRAE, Université Paris-Saclay, Jouy-en-Josas, FranceSearch for more papers by this authorStéphane ROBIN, Stéphane ROBIN LPSM, Sorbonne Université, Paris, FranceSearch for more papers by this author Gilles Didier, Gilles DidierSearch for more papers by this authorStéphane Guindon, Stéphane GuindonSearch for more papers by this author Book Author(s):Gilles Didier, Gilles DidierSearch for more papers by this authorStéphane Guindon, Stéphane GuindonSearch for more papers by this author First published: 12 April 2024 https://doi.org/10.1002/9781394284252.ch3 AboutPDFPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShareShare a linkShare onEmailFacebookTwitterLinkedInRedditWechat Summary The evolutionary models all have in common the use of a low cardinal state space or alphabet: four nucleotides for DNA and RNA and 20 amino acids for proteins. As with DNA sequences, it is expected that closely related species will possess more similar trait values than two more distantly related species. This link between sequence similarity and relatedness is the basis of many phylogenetic tree reconstruction methods. The purpose of continuous trait evolution models is to describe this similarity in order to understand or correct it. This chapter shows how the multivariate model allows the correlated evolution of multiple traits on a phylogenetic tree to be modeled. It discusses some non-Gaussian models that capture more sophisticated dynamics at the cost of an often much more complex estimation. The Brownian motion can then be used for modeling, in a stochastic environment, the evolution of adaptive traits over phylogenies. References Akaike , H. ( 1974 ). A new look at the statistical model identification . IEEE Transactions on Automatic Control , 19 ( 6 ), 716 – 723 . 10.1109/TAC.1974.1100705 CASWeb of Science®Google Scholar Alizon , S. , von Wyl , V. , Stadler , T. , Kouyos , D.R. , Yerly , S. , Hirschel , B. , Böni , J. , Shah , C. , Klimkait , T. , Furrer , H. et al. ( 2010 ). Phylogenetic approach reveals that virus genotype largely determines HIV set-point viral load . PLOS Pathogens , 6 ( 9 ). 10.1371/journal.ppat.1001123 PubMedWeb of Science®Google Scholar Aristide , L. and Morlon , H. ( 2019 ). Understanding the effect of competition during evolutionary radiations: An integrated model of phenotypic and species diversification . Ecology Letters , 22 ( 12 ), 2006 – 2017 . 10.1111/ele.13385 PubMedWeb of Science®Google Scholar Aristide , L. , dos Reis , S.F. , Machado , A.C. , Lima , I. , Lopes , R.T. , Perez , S.I. ( 2016 ). Brain shape convergence in the adaptive radiation of New World monkeys . Proceedings of the National Academy of Sciences , 113 ( 8 ), 2158 – 2163 . 10.1073/pnas.1514473113 CASPubMedWeb of Science®Google Scholar Aristide , L. , Bastide , P. , dos Reis , S.F. , Pires dos Santos , T.M. , Lopes , R.T. , Perez , S.I. ( 2018 ). Multiple factors behind early diversification of skull morphology in the continental radiation of New World monkeys . Evolution , 72 ( 12 ), 2697 – 2711 . 10.1111/evo.13609 PubMedWeb of Science®Google Scholar Bartoszek , K. , Pienaar , J. , Mostad , P. , Andersson , S. , Hansen , T.F. ( 2012 ). A phylogenetic comparative method for studying multivariate adaptation . Journal of Theoretical Biology , 314 , 204 – 215 . 10.1016/j.jtbi.2012.08.005 PubMedWeb of Science®Google Scholar Bartoszek , K. , Glémin , S. , Kaj , I. , Lascoux , M. ( 2017 ). Using the Ornstein–Uhlenbeck process to model the evolution of interacting populations . Journal of Theoretical Biology , 429 , 35 – 45 . 10.1016/j.jtbi.2017.06.011 PubMedWeb of Science®Google Scholar Bastide , P. , Mariadassou , M. , Robin , S. ( 2017 ). Detection of adaptive shifts on phylogenies by using shifted stochastic processes on a tree . Journal of the Royal Statistical Society: Series B (Statistical Methodology) , 79 ( 4 ), 1067 – 1093 . 10.1111/rssb.12206 Web of Science®Google Scholar Bastide , P. , Ané , C. , Robin , S. , Mariadassou , M. ( 2018a ). Inference of adaptive shifts for multivariate correlated traits . Systematic Biology , 67 ( 4 ), 662 – 680 . 10.1093/sysbio/syy005 PubMedWeb of Science®Google Scholar Bastide , P. , Solís-Lemus , C. , Kriebel , R. , Sparks , K.W. , Ané , C. ( 2018b ). Phylogenetic comparative methods on phylogenetic networks with reticulations . Systematic Biology , 67 ( 5 ), 800 – 820 . 10.1093/sysbio/syy033 PubMedWeb of Science®Google Scholar Beaulieu , J.M. , Jhwueng , D.-C. , Boettiger , C. , O'Meara , B.C. ( 2012 ). Modeling stabilizing selection: Expanding the Ornstein–Uhlenbeck model of adaptive evolution . Evolution , 66 ( 8 ), 2369 – 2383 . 10.1111/j.1558-5646.2012.01619.x PubMedWeb of Science®Google Scholar Blanquart , F. , Wymant , C. , Cornelissen , M. , Gall , A. , Bakker , M. , Bezemer , D. , Hall , M. , Hillebregt , M. , Ong , S.H. , Albert , J. et al. ( 2017 ). Viral genetic variation accounts for a third of variability in HIV-1 set-point viral load in Europe . PLOS Biology , 15 ( 6 ), 1 – 26 . 10.1371/journal.pbio.2001855 Web of Science®Google Scholar Blomberg , S.P. , Garland , T. , Ives , A.R. ( 2003 ). Testing for phylogenetic signal in comparative data: Behavioral traits are more labile . Evolution , 57 ( 4 ), 717 – 745 . 10.1111/j.0014-3820.2003.tb00285.x PubMedWeb of Science®Google Scholar Boucher , F.C. , Démery , V. , Conti , E. , Harmon , L.J. , Uyeda , J. ( 2018 ). A general model for estimating macroevolutionary landscapes . Systematic Biology , 67 ( 2 ), 304 – 319 . 10.1093/sysbio/syx075 PubMedWeb of Science®Google Scholar Butler , M.A. and King , A.A. ( 2004 ). Phylogenetic comparative analysis: A modeling approach for adaptive evolution . The American Naturalist , 164 ( 6 ), 683 – 695 . 10.1086/426002 PubMedWeb of Science®Google Scholar Ceccarelli , F. , Koch , N.M. , Soto , E.M. , Barone , M.L. , Arnedo , M.A. , Ramírez , M.J. ( 2018 ). The grass was greener: Repeated evolution of specialized morphologies and habitat shifts in ghost spiders following grassland expansion in South America . Systematic Biology , 68 ( 1 ), 63 – 77 . Web of Science®Google Scholar Clavel , J. , Escarguel , G. , Merceron , G. ( 2015 ). mvmorph: An r package for fitting multivariate evolutionary models to morphometric data . Methods in Ecology and Evolution , 6 ( 11 ), 1311 – 1319 . 10.1111/2041-210X.12420 Web of Science®Google Scholar Clavel , J. , Aristide , L. , Morlon , H. ( 2019 ). A penalized likelihood framework for high-dimensional phylogenetic comparative methods and an application to new-world monkeys brain evolution . Systematic Biology , 68 ( 1 ), 93 – 116 . 10.1093/sysbio/syy045 PubMedWeb of Science®Google Scholar Cooper , N. , Thomas , G.H. , Venditti , C. , Meade , A. , Freckleton , R.P. ( 2016 ). A cautionary note on the use of Ornstein–Uhlenbeck models in macroevolutionary studies . Biological Journal of the Linnean Society , 118 ( 1 ), 64 – 77 . 10.1111/bij.12701 PubMedWeb of Science®Google Scholar Crow , J.F. , Kimura , M. ( 1970 ). An Introduction to Population Genetics Theory . Harper & Row , New York . 10.1006/tpbi.1995.1025 Google Scholar Cybis , G.B. , Sinsheimer , J.S. , Bedford , T. , Mather , A.E. , Lemey , P. , Suchard , M.A. ( 2015 ). Assessing phenotypic correlation through the multivariate phylogenetic latent liability model . The Annals of Applied Statistics , 9 ( 2 ), 969 – 991 . 10.1214/15-AOAS821 PubMedWeb of Science®Google Scholar Drury , J. , Clavel , J. , Manceau , M. , Morlon , H. ( 2016 ). Estimating the effect of competition on trait evolution using maximum likelihood inference . Systematic Biology , 65 ( 4 ), 700 – 710 . 10.1093/sysbio/syw020 PubMedWeb of Science®Google Scholar Drury , J.P. , Grether , G.F. , Garland , T. , Morlon , H. ( 2018 ). An assessment of phylogenetic tools for analyzing the interplay between interspecific interactions and phenotypic evolution . Systematic Biology , 67 ( 3 ), 413 – 427 . 10.1093/sysbio/syx079 CASPubMedWeb of Science®Google Scholar Duchen , P. , Leuenberger , C. , Szilágyi , S.M. , Harmon , L.J. , Eastman , J.M. , Schweizer , M. , Wegmann , D. ( 2017 ). Inference of evolutionary jumps in large phylogenies using Lévy processes . Systematic Biology , 66 ( 6 ), 1 – 14 . 10.1093/sysbio/syx028 PubMedWeb of Science®Google Scholar Eastman , J.M. , Alfaro , M.E. , Joyce , P. , Hipp , A.L. , Harmon , L.J. ( 2011 ). A novel comparative method for identifying shifts in the rate of character evolution on trees . Evolution , 65 ( 12 ), 3578 – 3589 . 10.1111/j.1558-5646.2011.01401.x PubMedWeb of Science®Google Scholar Eastman , J.M. , Wegmann , D. , Leuenberger , C. , Harmon , L.J. ( 2013 ). Simpsonian "evolution by jumps' in an adaptive radiation of anolis lizards . arXiv , 1305.4216. Google Scholar Felsenstein , J. ( 1973 ). Maximum likelihood and minimum-steps methods for estimating evolutionary trees from data on discrete characters . Systematic Biology , 22 ( 3 ), 240 – 249 . 10.1093/sysbio/22.3.240 Google Scholar Felsenstein , J. ( 1985 ). Phylogenies and the comparative method . The American Naturalist , 125 ( 1 ), 1 – 15 . 10.1086/284325 Web of Science®Google Scholar Felsenstein , J. ( 2004 ). Inferring Phylogenies . Sinauer Associates , Sunderland . Google Scholar Felsenstein , J. ( 2005 ). Using the quantitative genetic threshold model for inferences between and within species . Philosophical Transactions of the Royal Society B: Biological Sciences , 360 ( 1459 ), 1427 – 1434 . 10.1098/rstb.2005.1669 PubMedWeb of Science®Google Scholar Felsenstein , J. ( 2008 ). Comparative methods with sampling error and within-species variation: Contrasts revisited and revised . The American Naturalist , 171 ( 6 ), 713 – 725 . 10.1086/587525 PubMedWeb of Science®Google Scholar Felsenstein , J. ( 2012 ). A comparative method for both discrete and continuous characters using the threshold model . The American Naturalist , 179 ( 2 ), 145 – 156 . 10.1086/663681 PubMedWeb of Science®Google Scholar Fitzjohn , R.G. ( 2010 ). Quantitative traits and diversification . Systematic Biology , 59 ( 6 ), 619 – 633 . 10.1093/sysbio/syq053 PubMedWeb of Science®Google Scholar Fitzjohn , R.G. ( 2012 ). Diversitree: Comparative phylogenetic analyses of diversification in R . Methods in Ecology and Evolution , 3 ( 6 ), 1084 – 1092 . 10.1111/j.2041-210X.2012.00234.x Web of Science®Google Scholar Fitzjohn , R.G. , Maddison , W.P. , Otto , S.P. ( 2009 ). Estimating trait-dependent speciation and extinction rates from incompletely resolved phylogenies . Systematic Biology , 58 ( 6 ), 595 – 611 . 10.1093/sysbio/syp067 PubMedWeb of Science®Google Scholar Freckleton , R.P. ( 2012 ). Fast likelihood calculations for comparative analyses . Methods in Ecology and Evolution , 3 ( 5 ), 940 – 947 . 10.1111/j.2041-210X.2012.00220.x Web of Science®Google Scholar Giraud , C. ( 2014 ). Introduction to High-Dimensional Statistics . CRC Press , Hoboken . 10.1201/b17895 Google Scholar Goldberg , E.E. and Foo , J. ( 2019 ). Memory in trait macroevolution . The American Naturalist , The University of Chicago Press , 195 ( 2 ), 300 – 314 . 10.1086/705992 PubMedWeb of Science®Google Scholar Goldberg , E.E. , Lancaster , L.T. , Ree , R.H. ( 2011 ). Phylogenetic inference of reciprocal effects between geographic range evolution and diversification . Systematic Biology , 60 ( 4 ), 451 – 465 . 10.1093/sysbio/syr046 PubMedWeb of Science®Google Scholar Goolsby , E.W. ( 2016 ). Likelihood-based parameter estimation for high-dimensional phylogenetic comparative models: Overcoming the limitations of "distance-based" methods . Systematic Biology , 65 ( 5 ), 852 – 870 . 10.1093/sysbio/syw051 PubMedWeb of Science®Google Scholar Goolsby , E.W. , Bruggeman , J. , Ané , C. ( 2017 ). Rphylopars: Fast multivariate phylogenetic comparative methods for missing data and within-species variation . Methods in Ecology and Evolution , 8 ( 1 ), 22 – 27 . 10.1111/2041-210X.12612 Web of Science®Google Scholar Grafen , A. ( 1989 ). The phylogenetic regression . Philosophical Transactions of the Royal Society B: Biological Sciences , 326 ( 1233 ), 119 – 157 . 10.1098/rstb.1989.0106 CASPubMedWeb of Science®Google Scholar Grafen , A. ( 1992 ). The uniqueness of the phylogenetic regression . Journal of Theoretical Biology , 156 ( 4 ), 405 – 423 . 10.1016/S0022-5193(05)80635-6 Web of Science®Google Scholar Green , P.J. ( 1995 ). Reversible jump Markov chain Monte Carlo computation and Bayesian model determination . Biometrika , 82 ( 4 ), 711 – 732 . 10.1093/biomet/82.4.711 Web of Science®Google Scholar Hadfield , J.D. and Nakagawa , S. ( 2010 ). General quantitative genetic methods for comparative biology: Phylogenies, taxonomies and multi-trait models for continuous and categorical characters . Journal of Evolutionary Biology , 23 ( 3 ), 494 – 508 . 10.1111/j.1420-9101.2009.01915.x CASPubMedWeb of Science®Google Scholar Hansen , T.F. ( 1997 ). Stabilizing selection and the comparative analysis of adaptation . Evolution , 51 ( 5 ), 1341 . 10.1111/j.1558-5646.1997.tb01457.x PubMedWeb of Science®Google Scholar Hansen , T.F. and Houle , D. ( 2004 ). Evolvability, stabilizing selection, and the problem of stasis . In Phenotypic Integration: Studying the Ecology and Evolution of Complex Phenotypes , M. Pigliucci and K. Preston (eds). Oxford University Press , New York . 10.1093/oso/9780195160437.003.0006 Web of Science®Google Scholar Hansen , T.F. and Martins , E.P. ( 1996 ). Translating between microevolutionary process and macroevolutionary patterns: The correlation structure of interspecific data . Evolution , 50 ( 4 ), 1404 . 10.1111/j.1558-5646.1996.tb03914.x PubMedWeb of Science®Google Scholar Hansen , T.F. and Orzack , S.H. ( 2005 ). Assessing current adaptation and phylogenetic inertia as explanations of trait evolution: The need for controlled comparisons . Evolution , 59 ( 10 ), 2063 – 2072 . 10.1111/j.0014-3820.2005.tb00917.x PubMedWeb of Science®Google Scholar Hansen , T.F. , Pienaar , J. , Orzack , S.H. ( 2008 ). A comparative method for studying adaptation to a randomly evolving environment . Evolution , 62 ( 8 ), 1965 – 1977 . 10.1111/j.1558-5646.2008.00412.x PubMedWeb of Science®Google Scholar Harmon , L.J. ( 2019 ). Phylogenetic comparative methods . [Online]. Available at: https://lukejharmon.github.io/pcm/ . Google Scholar Harmon , L.J. , Losos , J.B. , Jonathan , D.T. , Gillespie , R.G. , Gittleman , J.L. , Bryan , J.W. , Kozak , K.H. , McPeek , M.A. , Moreno-Roark , F. , Near , T.J. et al. ( 2010 ). Early bursts of body size and shape evolution are rare in comparative data . Evolution , 64 ( 8 ), 2385 – 2396 . 10.1111/j.1558-5646.2010.01025.x PubMedWeb of Science®Google Scholar Henderson , C.R. ( 1976 ). A simple method for computing the inverse of a numerator relationship matrix used in prediction of breeding values . Biometrics , 32 ( 1 ), 69 – 83 . 10.2307/2529339 Web of Science®Google Scholar Hiscott , G. , Fox , C. , Parry , M. , Bryant , D. ( 2016 ). Efficient recycled algorithms for quantitative trait models on phylogenies . Genome Biology and Evolution , 8 ( 5 ), 1338 – 1350 . 10.1093/gbe/evw064 PubMedWeb of Science®Google Scholar Ho , L.S.T. and Ané , C. ( 2014a ). A linear-time algorithm for Gaussian and non-Gaussian trait evolution models . Systematic Biology , 63 ( 3 ), 397 – 408 . 10.1093/sysbio/syu005 PubMedWeb of Science®Google Scholar Ho , L.S.T. and Ané , C. ( 2014b ). Intrinsic inference difficulties for trait evolution with Ornstein–Uhlenbeck models . Methods in Ecology and Evolution , 5 ( 11 ), 1133 – 1146 . 10.1111/2041-210X.12285 Web of Science®Google Scholar Housworth , E.A. , Martins , E.P. , Lynch , M. ( 2004 ). The phylogenetic mixed model . The American Naturalist , 163 ( 1 ), 84 – 96 . 10.1086/380570 PubMedWeb of Science®Google Scholar Hunt , G. and Rabosky , D.L. ( 2014 ). Phenotypic evolution in fossil species: Pattern and process . Annual Review of Earth and Planetary Sciences , 42 ( 1 ), 421 – 441 . 10.1146/annurev-earth-040809-152524 CASGoogle Scholar Hunt , G. , Bell , M.A. , Travis , M.P. ( 2008 ). Evolution toward a new adaptive optimum: Phenotypic evolution in a fossil stickleback lineage . Evolution , 62 ( 3 ), 700 – 710 . 10.1111/j.1558-5646.2007.00310.x PubMedWeb of Science®Google Scholar Ives , A.R. , Midford , P.E. , Garland , T. , Oakley , T. ( 2007 ). Within-species variation and measurement error in phylogenetic comparative methods . Systematic Biology , 56 ( 2 ), 252 – 270 . 10.1080/10635150701313830 PubMedWeb of Science®Google Scholar Jaffe , A.L. , Slater , G.J. , Alfaro , M.E. ( 2011 ). The evolution of island gigantism and body size variation in tortoises and turtles . Biology Letters , 7 ( 4 ), 558 – 561 . 10.1098/rsbl.2010.1084 PubMedWeb of Science®Google Scholar Jetz , W. , Thomas , G. , Joy , J. , Hartmann , K. , Mooers , A. ( 2012 ). The global diversity of birds in space and time . Nature , 491 ( 7424 ), 444 – 448 . 10.1038/nature11631 CASPubMedWeb of Science®Google Scholar Khabbazian , M. , Kriebel , R. , Rohe , K. , Ané , C. ( 2016 ). Fast and accurate detection of evolutionary shifts in Ornstein–Uhlenbeck models . Methods in Ecology and Evolution , 7 ( 7 ), 811 – 824 . 10.1111/2041-210X.12534 Web of Science®Google Scholar Kim , H. and Perl , J. ( 1983 ). A computational model for combined causal and diagnostic reasoning in inference systems . In Proceedings of the Eighth International Joint Conference on Artificial Intelligence . Morgan-Kaufmann , San Mateo . Google Scholar Labra , A. , Pienaar , J. , Hansen , T.F. ( 2009 ). Evolution of thermal physiology in Liolaemus lizards: Adaptation, phylogenetic inertia, and niche tracking . The American Naturalist , 174 ( 2 ), 204 – 220 . 10.1086/600088 PubMedWeb of Science®Google Scholar Lande , R. ( 1976 ). Natural selection and random genetic drift in phenotypic evolution . Evolution , 30 ( 2 ), 314 . 10.1111/j.1558-5646.1976.tb00911.x PubMedWeb of Science®Google Scholar Landis , M.J. , Schraiber , J.G. , Liang , M. ( 2013 ). Phylogenetic analysis using Lévy processes: Finding jumps in the evolution of continuous traits . Systematic Biology , 62 ( 2 ), 193 – 204 . 10.1093/sysbio/sys086 CASPubMedWeb of Science®Google Scholar Lartillot , N. ( 2014 ). A phylogenetic Kalman filter for ancestral trait reconstruction using molecular data . Bioinformatics , 30 ( 4 ), 488 – 496 . 10.1093/bioinformatics/btt707 CASPubMedWeb of Science®Google Scholar Law , C.J. , Slater , G.J. , Mehta , R.S. ( 2018 ). Lineage diversity and size disparity in musteloidea: Testing patterns of adaptive radiation using molecular and fossil-based methods . Systematic Biology , 67 ( 1 ), 127 – 144 . 10.1093/sysbio/syx047 CASPubMedWeb of Science®Google Scholar Lemey , P. , Rambaut , A. , Welch , J.J. , Suchard , M.A. ( 2010 ). Phylogeography takes a relaxed random walk in continuous space and time . Molecular Biology and Evolution , 27 ( 8 ), 1877 – 1885 . 10.1093/molbev/msq067 CASPubMedWeb of Science®Google Scholar Leventhal , G.E. and Bonhoeffer , S. ( 2016 ). Potential pitfalls in estimating viral load heritability . Trends in Microbiology , 24 ( 9 ), 687 – 698 . 10.1016/j.tim.2016.04.008 CASPubMedWeb of Science®Google Scholar Lynch , M. ( 1991 ). Methods for the analysis of comparative data in evolutionary biology . Evolution , 45 ( 5 ), 1065 – 1080 . 10.1111/j.1558-5646.1991.tb04375.x PubMedWeb of Science®Google Scholar Maddison , W.P. , Midford , P.E. , Otto , S.P. ( 2007 ). Estimating a binary character's effect on speciation and extinction . Systematic Biology , 56 ( 5 ), 701 – 710 . 10.1080/10635150701607033 PubMedWeb of Science®Google Scholar Mahler , D.L. , Ingram , T. , Revell , L.J. , Losos , J.B. ( 2013 ). Exceptional convergence on the macroevolutionary landscape in island lizard radiations . Science , 341 ( 6143 ), 292 – 295 . 10.1126/science.1232392 CASPubMedWeb of Science®Google Scholar Manceau , M. , Lambert , A. , Morlon , H. ( 2016 ). A unifying comparative phylogenetic framework including traits coevolving across interacting lineages . Systematic Biology , 66 ( 4 ), syw115 . 10.1093/sysbio/syw115 Web of Science®Google Scholar Mardia , K.V. , Kent , J.T. , Bibby , J.M. ( 1979 ). Multivariate Analysis, Probability and Mathematical Statistics . Academic Press , New York . Google Scholar Méléard , S. ( 2016 ). Modèles aléatoires en écologie et évolution. Mathématiques et applications . Springer , Berlin/Heidelberg . Google Scholar Meredith , R.W. , Janečka , J.E. , Gatesy , J. , Ryder , O.A. , Fisher , C.A. , Teeling , E.C. , Goodbla , A. , Eizirik , E. , Simão , T.L.L. , Stadler , T. et al. ( 2011 ). Impacts of the cretaceous terrestrial revolution and KPg extinction on mammal diversification . Science , 334 ( 6055 ), 521 – 524 . 10.1126/science.1211028 CASPubMedWeb of Science®Google Scholar Meucci , A. ( 2009 ). Review of statistical arbitrage, cointegration, and multivariate Ornstein–Uhlenbeck . SSRN Electronic Journal , 20 . Google Scholar Mitov , V. , Bartoszek , K. , Stadler , T. ( 2019 ). Automatic generation of evolutionary hypotheses using mixed Gaussian phylogenetic models . Proceedings of the National Academy of Sciences , 201813823 . Web of Science®Google Scholar Mitov , V. and Stadler , T. ( 2018 ). A practical guide to estimating the heritability of pathogen traits . Molecular Biology and Evolution , 35 ( 3 ), 756 – 772 . 10.1093/molbev/msx328 CASPubMedWeb of Science®Google Scholar Nuismer , S.L. and Harmon , L.J. ( 2015 ). Predicting rates of interspecific interaction from phylogenetic trees . Ecology Letters , 18 ( 1 ), 17 – 27 . 10.1111/ele.12384 PubMedWeb of Science®Google Scholar O'Meara , B.C. , Ané , C. , Sanderson , M.J. , Wainwright , P.C. ( 2006 ). Testing for different rates of continuous trait evolution using likelihood . Evolution , 60 ( 5 ), 922 – 933 . 10.1111/j.0014-3820.2006.tb01171.x PubMedWeb of Science®Google Scholar Pagel , M. ( 1999 ). Inferring the historical patterns of biological evolution . Nature , 401 ( 6756 ), 877 – 884 . 10.1038/44766 CASPubMedWeb of Science®Google Scholar Paradis , E. , Claude , J. , Strimmer , K. ( 2004 ). APE: Analyses of phylogenetics and evolution in R language . Bioinformatics , 20 ( 2 ), 289 – 290 . 10.1093/bioinformatics/btg412 CASPubMedWeb of Science®Google Scholar Pennell , M.W. , Eastman , J.M. , Slater , G.J. , Brown , J.W. , Uyeda , J.C. , FitzJohn , R.G. , Alfaro , M.E. , Harmon , L.J. ( 2014 ). geiger v2.0: An expanded suite of methods for fitting macroevolutionary models to phylogenetic trees . Bioinformatics , 30 ( 15 ), 2216 – 2218 . 10.1093/bioinformatics/btu181 CASPubMedWeb of Science®Google Scholar Pybus , O.G. , Suchard , M.A. , Lemey , P. , Bernardin , F.J. , Rambaut , A. , Crawford , F.W. , Gray , R.R. , Arinaminpathy , N. , Stramer , S.L. , Busch , M.P. et al. ( 2012 ). Unifying the spatial epidemiology and molecular evolution of emerging epidemics . Proceedings of the National Academy of Sciences , 109 ( 37 ), 15066 – 15071 . 10.1073/pnas.1206598109 CASPubMedWeb of Science®Google Scholar Rabosky , D.L. ( 2014 ). Automatic detection of key innovations, rate shifts, and diversity-dependence on phylogenetic trees . PLOS ONE , 9 ( 2 ). 10.1371/journal.pone.0089543 PubMedWeb of Science®Google Scholar Rabosky , D.L. and Huang , H. ( 2016 ). A robust semi-parametric test for detecting trait-dependent diversification . Systematic Biology , 65 ( 2 ), 181 – 193 . 10.1093/sysbio/syv066 PubMedWeb of Science®Google Scholar Raz , R. ( 2003 ). On the complexity of matrix product . SIAM Journal on Computing , 32 ( 5 ), 1356 – 1369 . 10.1137/S0097539702402147 Web of Science®Google Scholar Revell , L.J. ( 2009 ). Size-correction and principal components for interspecific comparative studies . Evolution , 63 ( 12 ), 3258 – 3268 . 10.1111/j.1558-5646.2009.00804.x PubMedWeb of Science®Google Scholar Revell , L.J. ( 2010 ). Phylogenetic signal and linear regression on species data . Methods in Ecology and Evolution , 1 ( 4 ), 319 – 329 . 10.1111/j.2041-210X.2010.00044.x Web of Science®Google Scholar Revell , L.J. ( 2012 ). Phytools: An R package for phylogenetic comparative biology (and other things) . Methods in Ecology and Evolution , 3 ( 2 ), 217 – 223 . 10.1111/j.2041-210X.2011.00169.x Web of Science®Google Scholar Revell , L.J. and Harmon , L.J. ( 2022 ). Phylogenetic Comparative Methods . Princeton University Press . Google Scholar Revell , L.J. , Harmon , L.J. , Collar , D.C. ( 2008 ). Phylogenetic signal, evolutionary process, and rate . Systematic Biology , 57 ( 4 ), 591 – 601 . 10.1080/10635150802302427 PubMedWeb of Science®Google Scholar Rose , J.P. , Kriebel , R. , Sytsma , K.J. ( 2016 ). Shape analysis of moss (Bryophyta) sporophytes: Insights into land plant evolution . American Journal of Botany , 103 ( 4 ), 652 – 662 . 10.3732/ajb.1500394 CASPubMedWeb of Science®Google Scholar Silvestro , D. , Kostikova , A. , Litsios , G. , Pearman , P.B. , Salamin , N. ( 2015 ). Measurement errors should always be incorporated in phylogenetic comparative analysis . Methods in Ecology and Evolution , 6 ( 3 ), 340 – 346 . 10.1111/2041-210X.12337 Web of Science®Google Scholar Slater , G.J. and Pennell , M.W. ( 2014 ). Robust regression and posterior predictive simulation increase power to detect early bursts of trait evolution . Systematic Biology , 63 ( 3 ), 293 – 308 . 10.1093/sysbio/syt066 PubMedWeb of Science®Google Scholar Stuart , Y.E. , Campbell , T.S. , Hohenlohe , P.A. , Reynolds , R.G. , Revell , L.J. , Losos , J.B. ( 2014 ). Rapid evolution of a native species following invasion by a congener . Science , 346 ( 6208 ), 463 – 466 . 10.1126/science.1257008 CASPubMedWeb of Science®Google Scholar Thompson , E.A. ( 2000 ). Statistical Inference from Genetic Data on Pedigrees: V 6 (Nsf-Cbms Confererence Series in Probability & Statistics Volume 6) . Institute of Mathematical Statistics , Beachwood . 10.1214/cbms/1462106037 Google Scholar Tibshirani , R. ( 1996 ). Regression selection and shrinkage via the lasso . Journal of the Royal Statistical Society. Series B (Methodological) , 58 ( 1 ), 267 – 288 . 10.1111/j.2517-6161.1996.tb02080.x Web of Science®Google Scholar Tolkoff , M.R. , Alfaro , M.E. , Baele , G. , Lemey , P. , Suchard , M.A. ( 2018 ). Phylogenetic factor analysis . Systematic Biology , 67 ( 3 ), 384 – 399 . 10.1093/sysbio/syx066 PubMedWeb of Science®Google Scholar Uyeda , J.C. and Harmon , L.J. ( 2014 ). A novel Bayesian method for inferring and interpreting the dynamics of adaptive landscapes from phylogenetic comparative data . Systematic Biology , 63 ( 6 ), 902 – 918 . 10.1093/sysbio/syu057 PubMedWeb of Science®Google Scholar Uyeda , J.C. , Caetano , D.S. , Pennell , M.W. ( 2015 ). Comparative analysis of principal components can be misleading . Systematic Biology , 64 ( 4 ), 677 – 689 . 10.1093/sysbio/syv019 CASPubMedWeb of Science®Google Scholar Vrancken , B. , Lemey , P. , Rambaut , A. , Bedford , T. , Longdon , B. , Günthard , H.F. , Suchard , M.A. ( 2015 ). Simultaneously estimating evolutionary history and repeated traits phylogenetic signal: Applications to viral and host phenotypic evolution . Methods in Ecology and Evolution , 6 ( 1 ), 67 – 82 . 10.1111/2041-210X.12293 PubMedWeb of Science®Google Scholar Models and Methods for Biological Evolution: Mathematical Models and Algorithms to Study Evolution ReferencesRelatedInformation
Next-generation biomonitoring proposes to combine machine-learning algorithms with environmental DNA data to automate the monitoring of the Earth's major ecosystems. In the present study, we searched for molecular biomarkers of tree water status to develop next-generation biomonitoring of forest ecosystems. Because phyllosphere microbial communities respond to both tree physiology and climate change, we investigated whether environmental DNA data from tree phyllosphere could be used as molecular biomarkers of tree water status in forest ecosystems. Using an amplicon sequencing approach, we analysed phyllosphere microbial communities of four tree species (Quercus ilex, Quercus robur, Pinus pinaster and Betula pendula) in a forest experiment composed of irrigated and non-irrigated plots. We used these microbial community data to train a machine-learning algorithm (Random Forest) to classify irrigated and non-irrigated trees. The Random Forest algorithm detected tree water status from phyllosphere microbial community composition with more than 90% accuracy for oak species, and more than 75% for pine and birch. Phyllosphere fungal communities were more informative than phyllosphere bacterial communities in all tree species. Seven fungal amplicon sequence variants were identified as candidates for the development of molecular biomarkers of water status in oak trees. Altogether, our results show that microbial community data from tree phyllosphere provides information on tree water status in forest ecosystems and could be included in next-generation biomonitoring programmes that would use in situ, real-time sequencing of environmental DNA to help monitor the health of European temperate forest ecosystems.
Grouping observations into homogeneous groups is a recurrent task in statistical data analysis. We consider Gaussian Mixture Models, which are the most famous parametric model-based clustering method. We propose a new robust approach for model-based clustering, which consists in a modification of the EM algorithm (more specifically, the M-step) by replacing the estimates of the mean and the variance by robust versions based on the median and the median covariation matrix. All the proposed methods are available in the R package RGMM accessible on CRAN.
The entropy is a measure of uncertainty that plays a central role in information theory. When the distribution of the data is unknown, an estimate of the entropy needs to be obtained from the data sample itself. A semi-parametric estimate is proposed based on a mixture model approximation of the distribution of interest. A Gaussian mixture model is used to illustrate the accuracy and versatility of the proposal, although the estimate can rely on any type of mixture. Performance of the proposed approach is assessed through a series of simulation studies. Two real-life data examples are also provided to illustrate its use.
Le modèle Poisson log-normal multivarié propose une modélisation conjointe des abondances des espèces d’une communauté distinguant les effets environnementaux (abiotiques) des interactions entre espèces (biotiques). Ses différentes variantes permettent la visualisation par réduction de dimension ou l’inférence du réseau d’interactions directes entre les espèces. Ces approches sont utilisées pour analyser l’écosystème marin de la forêt de kelp de l’île d’Anacapa.
After electricity liberalization, the "energy-only market" design lacks effective incentives to invest in new capacity. Across the world, capacity remuneration mechanisms have been taking hold as an alternative to the energy-only market to ensure adequate power generation capacity. The way these designs are characterized reflects the insight from general economic theory on market power and strategic behavior. Indeed, Europe is heading toward a patchwork of different, uncoordinated national mechanisms. In practice, design choices are driven by national policies, needs and constraints. In addition, they are progressively converging toward a market-based mechanism with a forward period and market power mitigation. In this paper, we investigate their efficiency properties in terms of new investments, the reduction in unserved energy frequency and the energy prices of a generic capacity remuneration mechanism that is impervious to the so-called forward capacity market. The results from laboratory experiments confirm that this mechanism gives better incentives to invest in new capacity than the energy-only market alone and provides an empirical understanding of the investment decision process. The forward capacity market contributes to lower energy market prices in peak demand and extra-high peak demand periods and average energy costs if we consider the social cost of unserved energy.
Bipartite networks are a natural representation of the interactions between entities from two different types. The organization (or topology) of such networks gives insight to understand the systems they describe as a whole. Here, we rely on motifs which provide a meso-scale description of the topology. Moreover, we consider the bipartite expected degree distribution (B-EDD) model which accounts for both the density of the network and possible imbalances between the degrees of the nodes. Under the B-EDD model, we prove the asymptotic normality of the count of any given motif, considering sparsity conditions. We also provide close-form expressions for the mean and the variance of this count. This allows to avoid computationally prohibitive resampling procedures. Based on these results, we define a goodness-of-fit test for the B-EDD model and propose a family of tests for network comparisons. We assess the asymptotic normality of the test statistics and the power of the proposed tests on synthetic experiments and illustrate their use on ecological data sets.
On s'intéresse ici à la probabilité de passer d'un état A à un état B mais dans le cas où le caractère d'intérêt est un trait quantitatif comme la taille ou le poids. Les modèles utilisés diffèrent du cas discrets et dérivent principalement du mouvement brownien. Ce chapitre présente les principaux modèles d'évolution de traits quantitatifs ainsi que les méthodes permettant de les appliquer dans le contexte évolutif.
A lot of what we know about past speciation and extinction dynamics is based on statistically fitting birth-death processes to phylogenies of extant species. Despite their wide use, the reliability of these tools is regularly questioned. It was recently demonstrated that vast 'congruent' sets of alternative diversification histories cannot be distinguished (i.e., are not identifiable) using extant phylogenies alone, reanimating the debate about the limits of phylogenetic diversification analysis. Here, we summarize what we know about the identifiability of the birth-death process and how identifiability issues can be addressed. We conclude that extant phylogenies, when combined with appropriate prior hypotheses and regularization techniques, can still tell us a lot about past diversification dynamics.
Joint Species Distribution Models (JSDM) provide a general multivariate framework to study the joint abundances of all species from a community. JSDM account for both structuring factors (environmental characteristics or gradients, such as habitat type or nutrient availability) and potential interactions between the species (competition, mutualism, parasitism, etc.), which is instrumental in disentangling meaningful ecological interactions from mere statistical associations. Modeling the dependency between the species is challenging because of the count-valued nature of abundance data and most JSDM rely on Gaussian latent layer to encode the dependencies between species in a covariance matrix. The multivariate Poisson-lognormal (PLN) model is one such model, which can be viewed as a multivariate mixed Poisson regression model. Inferring such models raises both statistical and computational issues, many of which were solved in recent contributions using variational techniques and convex optimization tools. The PLN model turns out to be a versatile framework, within which a variety of analyses can be performed, including multivariate sample comparison, clustering of sites or samples, dimension reduction (ordination) for visualization purposes, or inferring interaction networks. This paper presents the general PLN framework and illustrates its use on a series a typical experimental datasets. All the models and methods are implemented in the R package PLNmodels, available from cran.r-project.org.
Sophie Schbath合作论文数Institut National de la Recherche Agrononique
Unité Mathématique8