
Scatter search is a population-based metaheuristic designed to solve complex optimization problems through structured solution combination and adaptive memory. Unlike traditional evolutionary algorithms, scatter search emphasizes deterministic strategies to balance intensification and diversification. We present a comprehensive review of scatter search and its connection to path relinking, covering their historical development, core methodology, and applications. Key components of scatter search include diversification generation, improvement, reference set updating, subset generation, and solution combination. Advanced strategies such as dynamic reference set updating, tiered memory structures, constructive and destructive neighborhoods, and vocabulary building enhance its performance and scalability. Scatter search has been successfully applied in scheduling, routing, bioinformatics, and software engineering. Hybridizations with other metaheuristics and integration with machine learning further expand its applicability. The review concludes with a tutorial on a scatter search Python implementation for 0-1 knapsack problems that includes a Jupyter Notebook with code, execution traces, visualizations, and didactic analyses.
Computing a neural network is, in essence, a large unconstrained optimization problem, though this fact is often obscured by machine learning jargon. We present a gentle introduction to deep learning for operations researchers, describing in a didactic manner the underlying optimization problem and providing examples developed from scratch. These examples are solved using both standard modeling languages and modern machine learning frameworks. We conclude by discussing applications of neural networks in operations research, illustrated through a concrete example involving the knapsack problem.
Reconciling multidimensional count data across multiple sources is a common challenge in social and economic research. Iterative proportional fitting is widely used for this purpose, but aligning indices under weighted sum-convex constraints calls for a more fexible approach. We introduce the weighted iterative proportional fitting algorithm, which incorporates sum-weighted constraints to adjust indicators-such as death-risk indices by wealth, habitat, and climate-while preserving marginal consistency. Weighted iterative proportional fitting has been implemented in an R package of the same name, enabling scholars, statisticians, and policymakers' advisors, among others, to apply it easily to multidimensional data.
State-of-the-art methods for estimating extreme quantiles (value at risk, VaR) and their tail expectations (conditional tail expectation, CTE) under covariate control primarily rely on quantile regression but lack explicit constraints for non-crossing conditions. To address this, we introduce the non-crossing dual neural network, a deep learning model that simultaneously estimates VaRs and CTEs across multiple quantile levels, incorporates covariate dependence, and enables the reconstruction of individual conditional distributions and risk profles while ensuring the natural order of quantile levels. Using a 2015 telematics dataset, the proposed methodology outperforms benchmark models while enforcing previously unaddressed conditions. The model can be used to identify risk within an insurance portfolio and to analyze extreme right-tail behaviour at the individual level.
Data scientists address real-world problems using multivariate and heterogeneous datasets, characterized by multiple variables of different natures. Selecting a suitable distance function between units is crucial, as many statistical techniques and machine learning algorithms depend on this concept. Traditional distances, such as Euclidean or Manhattan, are unsuitable for mixed-type data, and although Gower distance was designed to handle this kind of data, it may lead to suboptimal results in the presence of outlying units or underlying correlation structure. In this work robust distances for mixed-type data are defned and explored, namely robust generalized Gower and robust related metric scaling. A new Python package is developed, which enables to compute these robust proposals as well as classical ones.
The area under the ROC curve (AUC) plays an important role in the study of the predictive capacity of regression models. It is well known that an infated AUC may result when the same data are used for training and testing the model. In this paper optimism correction of the AUC in the presence of missing data is investigated. Complete case analysis, inverse probability weighting and multiple imputation are employed to address the issue of missing data. For each of these approaches, split-sample, K-fold cross-validation and leave-one-out cross-validation are employed to correct for the optimism of the AUC. The methods are compared through intensive Monte Carlo simulations in the particular setting of binary regression. Results suggest that all estimators are consistent with the exception of complete case analysis, which may be biased when missing is not completely at random. In general, a combined application of multiple imputation and leave-one-out cross-validation is recommended.
We propose a stochastic partial differential equation to model geo-referenced data in the plane, with spatially correlated noise and a temporal log-normal evolution. Discretization in space permits us to develop the model in a finite-dimensional framework, reducing it to a set of stochastic differential equations coupled by correlated Wiener processes. The correlations are considered time-varying and stochastic, with a transformed log-normal distribution. The final model is framed within a hierarchical structure, and parameter inference is conducted jointly using Bayesian methods. The statistical methodology is illustrated by analyzing crime activity in the city of Valencia, Spain.
When modeling time-to-event data that are subject to right censoring, it is commonly assumed that the survival time T and the censoring time C are independent. However, this assumption frequently fails in practice, leading to biased estimators and testing procedures having invalid type 1 error rates. To overcome this issue, several models relaxing the independent censoring assumption have been proposed in the literature. Among these, copula-based approaches have become popular due to their ability to separately model the marginal distributions of T and C and their dependence structure. This review paper gives a comprehensive overview of recent advances in copula-based methods for dependent censoring, along with a discussion of the most important historical papers on this topic. As it is well known that the distribution of (T,C) (and hence of T) is not identifed in a fully nonparametric way, we examine different strategies to achieve model identifability. These strategies consist of imposing assumptions on either the copula or the marginal distributions of T and C. Both of these approaches will be discussed, with and without covariates. We also consider the case where a dependent censoring time is accompanied by an additional latent independent censoring time. Lastly, we briefy explain alternative approaches that are not based on copulas.
We propose the geometric mean spatial conditional model for fitting spatial public health data, assuming that the disease incidence in one region depends on that of neighbouring regions, and incorporating an autoregressive spatial term based on their geometric mean. We explore alternative spatial weights matrices, including those based on contiguity, distance, covariate differences and individuals' mobility. A simulation study assesses the model's performance with mobility-based spatial correlation. We illustrate our proposals by analysing the COVID-19 spread in Flanders, Belgium, and comparing the proposed model with other commonly used spatial models. Our approach demonstrates advantages in interpretability, computational efficiency, and fexibility over the commonly used and previously existing methods.
The comparison of investments in financial derivatives is an appealing topic in the optimization of resources. A relevant derivative is the call ratio backspread. Motivated by the need to compare investments in such derivatives, a new family of stochastic orders is introduced. That permits to reach decisions on the allocations of funds in those derivatives under general conditions and without assuming specific probability distributions of the asset prices. Characterizations of the orders are developed. Special emphasis is placed on the existence of infima and suprema in such dominance criteria, which leads to lattice structures on some special spaces and to the reduction of some optimization problems with stochastic dominance constraints. The method is illustrated with an application using real data from financial markets.
This paper presents a Bayesian inferential framework for estimating joint, conditional, and marginal probabilities in directed acyclic graphs (DAGs) applied to the study of the progression of hospitalized patients with severe influenza. Using data from the PIDIRAC retrospective cohort study in Catalonia, we model patient pathways from admission through different stages of care until discharge, death, or transfer to a long-term care facility. Direct transition probabilities are estimated through a Bayesian approach combining conjugate Dirichlet-multinomial inferential processes, while posterior distributions associated to absorbing state or inverse probabilities are assessed via simulation techniques. Bayesian methodology quantifies uncertainty through posterior distributions, providing insights into disease progression and improving hospital resource planning during seasonal influenza peaks. These results support more effective patient management and decision making in healthcare systems. Keywords: Confirmed influenza hospitalization; Directed acyclic graphs (DAGs); Dirichlet-multinomial Bayesian inferential process; Healthcare decision-making; Transition probabilities.
The maxima and the minima of a randomly stopped sample of a random variable, $X$, together with two newly defined random variables that make $X$ into the maxima or minima of a randomly stopped sample of them, can be used to define statistical model transformation mechanisms. These transformations can be used to define models for extreme value data that are not grounded on large sample theory. The relationship between the stopping model and characteristics of the corresponding model transformations obtained is investigated. In particular, one looks into which stopping models make these model transformations into model extensions, and which stopping models lead to statistically stable extensions in the sense that using the model extension a second time leaves the extended model unchanged. The stopping models under which the extensions based on randomly stopped maxima and their inverses coincide with the extensions based on randomly stopped minima and their inverses are also characterized. The advantages of using models obtained through these model extension mechanisms instead of resorting to extreme value models grounded on asymptotic arguments is illustrated by way of examples.
Evaluating the predictive performance of a statistical model is commonly done using cross-validation. Although the leave-one-out method is frequently employed, its application is justified primarily for independent and identically distributed observations. However, this method tends to mimic interpolation rather than prediction when dealing with dependent observations. This paper proposes a modified cross-validation for dependent observations. This is achieved by excluding an automatically determined set of observations from the training set to mimic a more reasonable prediction scenario. Also, within the framework of latent Gaussian models, we illustrate a method to adjust the joint posterior for this modified cross-validation to avoid model refitting. This new approach is accessible in the R-INLA package (www.r-inla.org).
Joint modelling has gained attention in longitudinal studies incorporating biomarkers and survival data. In the context of chronic diseases, patient evolution is often tracked through multiple assessments, with patient-reported outcomes playing a crucial role. The Beta-Binomial distribution is suggested as a suitable model for these longitudinal variables. However, its integration into joint modelling remains unexplored. This study introduces an estimation procedure for analyzing longitudinal patient-reported outcomes and survival data together. We compare different estimation approaches through simu- lation experiments, including the proposed model. Furthermore, the methodologies are applied to real data from a follow-up study on chronic obstructive pulmonary disease pa tients.
In many biomedical studies, recurrent or consecutive events may occur during the follow-up of the individuals. This situation can be found, for example, in transplant studies, where there are two consecutive events which give rise to two times of interest subject to a common random right-censoring time, the first one being the elapsed time from acceptance into the transplantation program to transplant, and the second one the time from transplant to death. In this work, we incorporate the information of a continuous covariate into the bivariate distribution of the two gap times of interest and propose a non-parametric method to cope with it. We prove the asymptotic properties of the proposed method and carry out a simulation study to see the performance of this approach. Additionally, we illustrate its use with Stanford heart transplant data and colon cancer data.
Prediction of a traffc accident cost is one of the major problems in motor insurance. To identify the factors that infuence costs is one of the main challenges of actuarial modelling. Telematics data about individual driving patterns could help calculating the expected claim severity in motor insurance. We propose using single-index models to assess the marginal effects of covariates on the claim severity conditional distribution. Thus, drivers with a claim cost distribution that has a long tail can be identifed. These are risky drivers, who should pay a higher insurance premium and for whom preventa-tive actions can be designed. A new kernel approach to estimate the covariance matrix of coeffcients' estimator is outlined. Its statistical properties are described and an ap-plication to an innovative data set containing information on driving styles is presented. The method provides good results when the response variable is skewed.
Gaussian random felds with Mat ern covariance functions are popular models in spatial statistics and machine learning. In this work, we develop a spatio-temporal extension of the Gaussian Mat ern felds formulated as solutions to a stochastic partial differential equation. The spatially stationary subset of the models have marginal spatial Mat ern covariances, and the model also extends to Whittle-Matern felds on curved manifolds, and to more general non-stationary felds. In addition to the parameters of the spatial dependence (variance, smoothness, and practical correlation range) it additionally has parameters controlling the practical correlation range in time, the smoothness in time, and the type of non-separability of the spatio-temporal covariance. Through the separability parameter, the model also allows for separable covariance functions. We provide a sparse representation based on a fnite element approximation, that is well suited for statistical inference and which is implemented in the R-INLA software. The fexibility of the model is illustrated in an application to spatio-temporal modeling of global temperature data.
Household composition reveals vital aspects of the socioeconomic situation and major changes in developed countries for decision -making and mapping the distribution of single -person households is highly relevant and useful. Driven by the Spanish Household Budget Survey data, we propose a new statistical methodology for small area estimation of proportions and total counts of single -person households. Estimation domains are defned as crosses of province, sex and age group of the main breadwinner of the household. Predictors are based on area -level zero-infated Poisson mixed models. Model parameters are estimated by maximum likelihood and mean squared errors by parametric bootstrap. Several simulation experiments are carried out to empirically investigate the properties of these estimators and predictors. Finally, the paper concludes with an application to real data from 2016.
In this paper we review some methods proposed in the literature for combining a nonprobability and a probability sample with the purpose of obtaining an estimator with a smaller bias and standard error than the estimators that can be obtained using only the probability sample. We propose a new methodology based on the kernel weighting method. We discuss the properties of the new estimator when there is only selection bias and when there are both coverage and selection biases. We perform an extensive simulation study to better understand the behaviour of the proposed estimator.
Multistate models (MSM) are well developed for continuous and discrete times under a first order Markov assumption. Motivated by a cohort of COVID-19 patients, an MSM was designed based on 14 transitions among 7 states of a patient. Since a preliminary analysis showed that the first order Markov condition was not met for some transitions, we have developed a second order Markov model where the future evolution not only depends on the current but also on the preceding state. Under a discrete time analysis, assuming homogeneity and that past information is restricted to 2 consecutive times, we expanded the transition probability matrix and proposed an extension of the Chapman- Kolmogorov equations.