
ABSTRACT The boxplot as a summary display of a one‐dimensional dataset is typically credited to John Tukey (1915–2000). His ideas for these displays became widespread when Exploratory Data Analysis (EDA) (Tukey 1977) was published; there he defined first the “box‐and‐whisker plot” and then the “schematic plot,” later called “boxplot”. Some displays with a similar flavor appeared 2–3 decades earlier. As a testament to its popularity, enhancements to the boxplot and extensions to weighted and multidimensional data have been proposed in the 50 years since EDA . In this article, I describe the boxplot's predecessors, define the standard boxplot, and review several other versions and recent enhancements that have appeared in the past 50 years. References to R implementations are cited where available. The goal is to clarify the boxplot's definition and origins, provide a review of boxplot‐inspired displays and extensions, including some for bivariate data (but excluding circular and functional data), and illustrate its widespread use for displaying data across many fields. This article is categorized under: Statistical and Graphical Methods of Data Analysis > Statistical Graphics and Visualization Statistical Learning and Exploratory Methods of the Data Sciences > Exploratory Data Analysis Statistical and Graphical Methods of Data Analysis > Robust Methods
ABSTRACT Network‐valued random vectors (NVRVs) provide a statistical framework for settings in which each observational unit is a network rather than a scalar, vector, image, or functional observation. Such data occur in social networks, omics and gene‐regulatory systems, functional brain connectivity, policy and intervention networks, and other domains where relational structure is itself the object of inference. NVRVs incorporate dependence through structured interactions among nodes and edges, thereby presenting significant challenges for statistical modeling, regression, and inference. This article provides an advanced review of inference techniques for NVRVs, with technical expositions focusing on the problem of measuring and testing network change in dynamic and heterogeneous settings. We review approaches for modeling networks as both responses and covariates, emphasizing key strategies such as edge‐wise models, summary‐based regression, latent variable methods, and Bayesian hierarchical formulations. We discuss how heterogeneity and temporal dynamics complicate inference, particularly in the context of network change detection, where current approaches frequently target isolated network features instead of the full network distribution. We consider the perspectives from both frequentist and Bayesian paradigms to identify fundamental gaps in current methodology including limitations in global network testing and the lack of theoretical guarantees in heterogeneous network data. Using the MRN‐114 dataset from the Mind Research Network database as empirical test cases, we perform integrative experiments demonstrating use of the techniques, discussing the advantages and pitfalls of each. Overall, our reviews compare techniques covering the representation of network‐valued observations, the measurement of dynamic change, and the inferential consequence of heterogeneity.
ABSTRACT Functional graphical models (FGMs) extend classical graphical models from multivariate vectors to multivariate random functions, enabling inference on conditional dependence and, in directed settings, directional or causal relationships across functional domains. This review provides a unified overview of recent developments in both undirected and directed FGMs. We summarize key theoretical foundations, estimation methods, and computational challenges arising from infinite dimensionality, operator non‐invertibility, dimension reduction, and non‐Gaussian behavior. Applications in brain connectivity, protein signaling, transportation systems, and longitudinal microbiome studies illustrate the broad potential of FGMs. We also highlight open research directions, including dynamic networks, adaptive truncation, and methodological extensions beyond Gaussian assumptions.
ABSTRACT Signal‐plus‐noise type matrices provide a natural extension of classical sample covariance models and arise widely in high‐dimensional statistics, signal processing, wireless communications, and machine learning. This review surveys recent advances in the spectral theory of such matrices from both global and local perspectives. We first introduce the basic signal‐plus‐noise and generalized signal‐plus‐noise models, emphasizing the role of deterministic signal components and heterogeneous noise structures. We then review the limiting spectral distribution and its characterization through Stieltjes transform equations, together with related analytic properties such as density regularity and support determination. Building on these first‐order results, we summarize spectral separation and exact separation phenomena, which describe how gaps in the limiting spectrum determine the location and number of sample eigenvalues. We further discuss edge behavior of extreme eigenvalues, including Tracy–Widom law, as well as the asymptotic theory of spiked eigenvalues and eigenvectors under general noise environments. In addition, the review covers recent progress on linear spectral statistics and their central limit theorems, highlighting the impact of deterministic signals on second‐order fluctuations. Overall, this survey provides a unified overview of the main spectral phenomena of signal‐plus‐noise type matrices and illustrates their theoretical significance and practical value in high‐dimensional inference, wireless communications, and modern machine learning problems such as noisy manifold learning and kernel‐based sensor fusion.
ABSTRACT Sufficient dimension reduction (SDR) refers to supervised methods of dimension reduction that apply in the context of regression, interpreted broadly. SDR started in the early 1990's with methodology to reduce linearly the predictor dimension without loss of information about the conditional distribution of the response given the predictors. The field grew quickly and today it is vast. The early ideas and methods have been formalized, extended, specialized and adapted to many problems in statistics. A comprehensive synopsis of everything covered by SDR would be truly substantial. Instead, the focus of this overview is on the ideas and philosophy of SDR. While the field is vast, there are a few foundational ideas that define the area generally and have been adapted to different problems. In this overview, we elucidate these ideas by explaining the historical ambience, exploring the genesis of the ideas and discussing how they are adapted for various problems. There is also literature on ‘dimensionality reduction’ that is beyond the scope of this overview. Examples include uniform manifold approximation and projection, neural PCA, kernel PCA, and locally linear embedding. Dimension reduction of data that are not meaningfully stochastic is also outside the scope of this article. The worlds of dimensionality reduction and sufficient dimension reduction were largely developed independently, even when in retrospect the developments involve overlap.
ABSTRACT Subgroup identification is a significant research area in statistics and machine learning, aiming to partition a heterogeneous population into more homogeneous subgroups to enable precise inference and personalized decision‐making. Among the various tools available, change point analysis has emerged as a powerful approach for detecting structural changes in data sequences, and it plays an increasingly important role in subgroup identification. In this paper, we provide a systematic review of recent advances in subgroup identification methods based on efficient multiple change point detection methods: (a) We first review the two‐step multiple change point detection method (TSMCD) and its application from linear regression to survival analysis. (b) We then discuss the construction of the threshold variable, including the recent developments in the change plane regression model and the change surface regression model. This review aims to provide researchers with a comprehensive perspective to promote the further application and development of change point analysis in subgroup identification and precision medicine.
ABSTRACT Reliable estimation of the covariance matrix is fundamental to multivariate analysis, second in importance only to the mean. Its accuracy directly impacts applications in economics, finance, chemistry, health science, bioinformatics, climate research, signal processing, and social network analysis. Modern data environments often involve wide data, where the number of variables is comparable to or exceeds the number of observations, making classical estimators unstable or singular. This review synthesizes recent advances in high‐dimensional covariance estimation, focusing on thresholding procedures, linear and nonlinear shrinkage methods, graphical model‐based approaches, and estimation techniques using random matrix theory. A unifying taxonomy is proposed that organizes these diverse techniques under a single conceptual framework, highlighting their interconnections and guiding the selection of appropriate estimators in wide‐data settings.
ABSTRACT Text mining has become central to bibliometrics, providing quantitative insight into the semantic structure of scientific communication. This review surveys current methodological approaches to text‐based science mapping, including geometric embeddings, probabilistic models, network techniques, and neural embedding methods. The discussion examines how these approaches operate across different representations of text and evaluates their interpretability, stability, and statistical assumptions. Key issues include data quality, model validation, reproducibility, and the growing influence of large language models. Persistent challenges—language bias, topic instability, limited full‐text access, and model opacity—raise open questions about dynamic, multimodal, and ethically grounded science mapping.
ABSTRACT The Hurst exponent () plays a key role in understanding long‐range dependence and self‐similarity in time series data. Wavelet‐based methods have gained popularity for estimating because they efficiently capture patterns across multiple scales. This review explores how these methods have developed from their theoretical roots to real‐world applications in fields like biology, engineering, and telecommunications. The review aims to highlight key techniques, compare their strengths and limitations, and point out challenges that still need to be addressed, such as handling noise and non‐stationary data. The field is well‐established, but there is growing interest in combining traditional wavelet‐based models with modern machine learning to push the boundaries even further. This review offers a clear starting point and roadmap for future exploration. We critically evaluate methodological robustness, computational scalability, and practical adoption, identifying key challenges that define the current frontier of self‐similarity estimation.
Marginal likelihood plays a central role in Bayesian model comparison and hypothesis testing, but its computation is often challenging in practice. This article reviews recent Monte Carlo methods that rely on the availability of Markov chain Monte Carlo (MCMC) samples from the posterior and prior distributions along with the corresponding unnormalized kernels that can be evaluated numerically. Within this scope, we summarize the strengths, limitations, differences, and connections of different methods. Two in‐depth applications are presented to illustrate their relative performance. This article is categorized under: Statistical and Graphical Methods of Data Analysis > Monte Carlo Methods Statistical Models > Model Selection Statistical Models > Bayesian Models
In analytic survey inference, the attributes of units in a target or frame survey population are idealized as a sample from a superpopulation statistical model, and model parameters are estimated from survey data drawn from a probability sample of the frame population. The data structure in such survey inference consists of relevant attribute data together with survey weights associated with all survey respondents. These weights relate to the probability of inclusion of each unit within the respondent set, and they enable consistent estimation in large populations and samples of all frame-population averages of functions of the unit attributes. However, even when this is assumed correct, model parameters such as those for within-cluster dependence between survey attributes from distinct respondents may not be identifiable from survey data with weights. That is, even assuming the superpopulation model, with a parametric dependence structure for attributes within clusters, if sampled data are observed with precisely correct weights equal to the reciprocals of single-inclusion probabilities or of conditional probabilities of inclusion given unit data, multiple distinct values of the parameters of the superpopulation model may yield the same likelihood for the data for some sample designs compatible with the weights. This article first describes the background and existing methods for the design-based estimation of cluster-level model parameters from survey data on a clustered superpopulation using single-inclusion weights. Nonidentifiability results are presented rigorously as mathematical examples, proving that large-sample consistent estimation of within-cluster dependence parameters from survey data with single-inclusion weights is not always possible. This article is categorized under: Statistical and Graphical Methods of Data Analysis > Sampling Algorithms and Computational Methods > Maximum Likelihood Methods
The functional delta method for deriving asymptotic distributions is presented. Assuming our interest lies in T θ 0 $$ T\left({\boldsymbol{\theta}}_0\right) $$ where θ 0 $$ {\boldsymbol{\theta}}_0 $$ is an unknown, infinite-dimensional parameter and T $$ T $$ is a known functional. Like the delta method the functional delta method allows to immediately obtain an approximation of the distribution of the plug-in estimator T θ ̂ n $$ T\left({\hat{\boldsymbol{\theta}}}_n\right) $$ through the asymptotic distribution of r n T θ ̂ n − T θ 0 $$ {r}_n\left(T\left({\hat{\boldsymbol{\theta}}}_n\right)-T\left({\boldsymbol{\theta}}_0\right)\right) $$ subject to (a) the asymptotic distribution of r n θ ̂ n − θ 0 $$ {r}_n\left({\hat{\boldsymbol{\theta}}}_n-{\boldsymbol{\theta}}_0\right) $$ being known and (b) the existence of an appropriate functional derivative of T $$ T $$ at θ 0 $$ {\boldsymbol{\theta}}_0 $$ . This article is categorized under:
Optimal transport (OT) methods and their variants have become increasingly prominent tools in computer science and machine learning, owing to their appealing geometric properties and powerful potency. Despite broad applications, OT methods suffer from prohibitively high computational cost, limiting the scalability even for moderately sized datasets. To address this challenge, regularized OT formulations and the corresponding Sinkhorn algorithm have emerged as standard alternatives to improve efficiency. However, these methods still face the high per-iteration cost and slow convergence rate drawbacks. Sparsification techniques have emerged as an effective and practically valuable class of methods for mitigating these computational bottlenecks by leveraging inherent or induced sparsity in the matrices involved in OT optimization. Broadly, sparsification methods can be grouped into two main categories: (1) kernel-based sparsification building on the primal regularized OT formulation, and (2) Hessian-based sparsification, derived from the dual formulation. In this survey, we provide an extensive and comprehensive review of sparsification techniques developed for OT problems, highlighting their underlying motivations, algorithmic distinctions, and theoretical guarantees. This article is categorized under:
Quantile regression has emerged as a powerful tool for modeling heterogeneous effects and tail behavior across different parts of the response distribution. This review highlights recent advances that address the challenges of high-dimensional and complex data, including penalized estimation, debiasing, distributed learning, transfer learning, and machine learning–based approaches for quantile regression. Practical procedures, theoretical insights, and available software are summarized to support the application of modern quantile regression in diverse data environments. New advances make quantile regression smarter, faster, and ready for complex, high-dimensional data. This article is categorized under:
In modern industrial settings, advanced acquisition systems allow for the collection of data in the form of profiles, that is, as functional relationships linking responses to explanatory variables. In this context, statistical process monitoring (SPM) aims to assess the stability of profiles over time in order to detect unexpected behavior. This review focuses on SPM methods that model profiles as functional data, that is, smooth functions defined over a continuous domain, and apply functional data analysis (FDA) tools to address limitations of traditional monitoring techniques. A reference framework for monitoring multivariate functional data is first presented. This review then offers a focused survey of several recent FDA-based profile monitoring methods that extend this framework to address common challenges encountered in real-world applications. These include approaches that integrate additional functional covariates to enhance detection power, a robust method designed to accommodate outlying observations, a real-time monitoring technique for partially observed profiles, and two adaptive strategies that target the characteristics of the out-of-control distribution. These methods are all implemented in the R package funcharts , available on CRAN. Finally, a review of additional existing FDA-based profile monitoring methods is also presented, along with suggestions for future research. This article is categorized under:
ABSTRACT Contrastive dimension reduction (CDR) methods aim to extract signal unique to or enriched in a treatment (foreground) group relative to a control (background) group. This setting arises in many scientific domains, such as genomics, imaging, and time series analysis, where traditional dimension reduction techniques such as principal component analysis (PCA) may fail to isolate the signal of interest. In this review, we provide a systematic overview of existing CDR methods. We propose a pipeline for analyzing case–control studies together with a taxonomy of CDR methods based on their assumptions, objectives, and mathematical formulations, unifying disparate approaches under a shared conceptual framework. We highlight key applications and challenges in existing CDR methods and identify open questions and future directions. By providing a clear framework for CDR and its applications, we aim to facilitate broader adoption and motivate further developments in this emerging field. This article is categorized under: Statistical Learning and Exploratory Methods of the Data Sciences > Manifold Learning Statistical and Graphical Methods of Data Analysis > Dimension Reduction Statistical and Graphical Methods of Data Analysis > Analysis of High Dimensional Data
This review article focuses on hidden truncation models, a versatile framework for modeling a wide range of random phenomena. These models are characterized by the condition that the primary study variable(s) are observable only when a concomitant variable (or a set of concomitant variables in the multivariate case) meets specific criteria. Hidden truncation, where the truncation mechanism is not directly observable and must be inferred from data, adds a layer of complexity that has inspired extensive research. While hidden truncation for normal distributions is straightforward and mathematically manageable, extending hidden truncation models to non‐normal distributions often results in complex forms involving computationally complex normalizing constants. This article surveys the foundational development of hidden truncation models, their mathematical and structural properties, and their relationship with the skewing paradigm. Through illustrative examples, we highlight the applicability of hidden truncation models across diverse real‐world scenarios. We also examine key aspects of statistical inference, unresolved challenges, and promising directions for future research, offering a comprehensive resource for researchers and practitioners interested in this evolving area of study.
Nuclear norm, also known as trace norm, has been widely used in statistical machine learning. Nuclear norm regularization has emerged as an important tool for addressing various statistical problems involving the estimation of low‐rank matrices, particularly in tasks such as matrix completion and reduced rank regression. This review delves into the foundational models, practical implementations, and recent advancements in nuclear norm regularization. We discuss key implementation techniques, including semidefinite programming and singular value thresholding, which enable efficient solutions to low‐rank matrix estimation problems. Additionally, we examine the application of nuclear norm regularization in matrix covariate and matrix response regression, as well as its extension to tensor regression problems. Our study highlights the versatility and efficacy of nuclear norm regularization in providing both theoretical guarantees and scalable computational methods. Future research directions include improving computational efficiency, refining conditions for theoretical guarantees and extending applications to higher‐order tensors.
A surrogate endpoint is an endpoint that is used as a substitute for a direct clinical outcome measure of how a patient feels, functions, or survives. The use of surrogate endpoints offers potentially significant ethical, logistical, and/or economic advantages to clinical studies, thereby deserving serious consideration in clinical development planning. The means of establishing a biomarker as a validated or reasonably likely surrogate endpoint requires scientific evidence of the predictive nature of a potential surrogate endpoint and statistical validation of the predictivity of the surrogate endpoint on clinical outcome measures. In this review, we present statistical methodologies and regulatory considerations for establishing surrogate endpoints and provide a few applications across multiple disease areas.
Several notions of concentration function have been proposed in the statistical literature. Here we will refer to the definition by Cifarelli and Regazzini (1987), used for the comparison of probability measures in very general probability spaces. We will consider the concentration function in more restricted settings, typically those of interest to practitioners, that is, discrete or absolutely continuous (with respect to the Lebesgue measure) probability measures. We will show how a sophisticated mathematical tool can be used to provide graphs and indices that are easily interpreted by practitioners. Some examples will support this statement. This article is categorized under: Statistical Learning and Exploratory Methods of the Data Sciences > Exploratory Data Analysis Statistical and Graphical Methods of Data Analysis > Statistical Graphics and Visualization