Latin hypercube designs (LHDs) play an important role in computer experiments, offering flexible and efficient space-filling properties under a variety of optimality criteria, including maximin distance, maximum projection, and orthogonality. Constructing optimal LHDs with flexible sizes is challenging due to the limited availability of theoretical results, namely, algebraic constructions with provable optimality guarantees, and the rapidly expanding search space encountered by search algorithms. For design sizes outside the scope of known algebraic constructions, search-based algorithms are widely used, but they often involve substantial computational effort and lack assurances of global optimality. This paper provides a comprehensive review and comparison of current popular algebraic constructions for optimal LHDs and widely adopted search algorithms for generating high-quality LHDs. By reviewing their theoretical properties and empirical performance across a range of criteria, we offer a unified perspective on their relative strengths and limitations. The comparisons presented herein aim to assist practitioners in selecting appropriate design strategies for different objectives and constraints. The insights from this work also highlight open challenges and may serve as a benchmark for future development in optimal LHD construction.
Discrete choice experiments (DCEs) are popular in business, marketing, health sciences, and many other fields. Panel mixed logit models are a special case of generalized linear mixed models that are natural choices for analyzing data arising from such experiments. In this article, we propose techniques for identifying optimal designs for panel mixed logit models. Here the information matrix does not have a closed form expression and is computationally intensive to evaluate numerically, which to date has made finding designs under these models using search algorithms difficult. To overcome this difficulty, we propose using alternative forms of the information matrix based on penalized quasi-likelihood (PQL), marginal quasi-likelihood (MQL) and the method of simulated moments (MSM). Our simulation results suggest that PQL is the best option when a design with a very high efficiency is required, but that the significantly faster MQL may be acceptable in many cases. We use the proposed methods to search for D-optimal designs for two nontrivial DCE examples reported in literature with a large number of attributes and choice sets, which would be prohibitively expensive using existing techniques. All approaches are implemented in an R package.
Finding D-optimal designs for generalized linear models (GLMs) is challenging due to the dependence of the Fisher information matrix on unknown parameters and the lack of closed-form solutions, particularly when input factors include both discrete and continuous variables. Although classical algorithms and recent metaheuristic approaches have offered partial solutions, there remains a need for robust and computationally efficient methods. In this paper, we propose a penalized Particle Swarm Optimization (PSO) approach, named $p$-PSO. Here we introduce a new, general-purpose penalty formulation for constrained optimization and demonstrate its effectiveness in optimal design problems. The formulation is algorithm-agnostic and applicable to a broad class of black-box optimization methods. Results show that the method is highly efficient, with its primary contribution being a penalty formulation that enables the direct use of an off-the-shelf PSO algorithm and extends naturally to more general constrained optimization tasks.
Optimal Latin hypercube designs (LHDs), including maximin distance LHDs, maximum projection LHDs and orthogonal LHDs, are widely used in computer experiments. It is challenging to construct such designs with flexible sizes, especially for large ones, for two main reasons. One reason is that theoretical results, such as algebraic constructions ensuring the maximin distance property or orthogonality, are only available for certain design sizes. For design sizes where theoretical results are unavailable, search algorithms can generate designs. However, their numerical performance is not guaranteed to be optimal. Another reason is that when design sizes increase, the number of permutations grows exponentially. Constructing optimal LHDs is a discrete optimization process, and enumeration is nearly impossible for large or moderate design sizes. Various search algorithms and algebraic constructions have been proposed to identify optimal LHDs, each having its own pros and cons. We develop the R package LHD which implements various search algorithms and algebraic constructions. We embedded different optimality criteria into each of the search algorithms, and they are capable of constructing different types of optimal LHDs even though they were originally invented to construct maximin distance LHDs only. Another input argument that controls maximum CPU time is added to each of the search algorithms to let users flexibly allocate their computational resources. We demonstrate functionalities of the package by using various examples, and we provide guidance for experimenters on finding suitable optimal designs. The LHD package is easy to use for practitioners and possibly serves as a benchmark for future developments in LHD.
A new type of experiment, called the order-of-addition factorial experiment, has recently received considerable attention in medical science and bioengineering. These experiments aim to simultaneously optimize the order of addition and dose levels of drug components. In the experimental design literature, the concept of dual-orthogonal arrays (DOAs), a class of optimal order-of-addition two-level factorial designs under the compound model, has recently been introduced for such experiments. However, constructing flexible DOAs remains a challenging task. In this article, we propose a novel theory-guided search method to efficiently identify DOAs of any feasible size. We also provide an algebraic construction method that immediately leads to certain DOAs. Moreover, to address the potential issue that DOA ignores interaction effects, we propose to construct a new type of optimal designs under the expanded compound model, named the strong DOA (SDOA). We provide two algebraic construction methods for the SDOA. We establish theoretical results on the optimality of both DOAs and SDOAs. Numerical studies illustrate the superiority of our proposed designs.
A new spline-based method was developed to transform over 100,000 multi-gigapixel plant root whole-slide scanning images into deep learning-ready image patches. We validated the robustness of our approach by analyzing eight bootstrap samples of a thousand images each classified as excellent, moderate, or bad quality by four independent experts. The goodness of fit, roughness, and irregularity of the splines were uniform across all quality levels, confirming our method's reliability for generating patches from gigapixel plant root images. These patches will be used as an input to deep learning based algorithms capable of detecting and classifying mycorrhized root segments and types of fungal structures presented.
In this paper, we address the problem of designing an experimental plan with both discrete and continuous factors under fairly general parametric statistical models. We propose a new algorithm, named ForLion, to search for locally optimal approximate designs under the D-criterion. The algorithm performs an exhaustive search in a design space with mixed factors while keeping high efficiency and reducing the number of distinct experimental settings. Its optimality is guaranteed by the general equivalence theorem. We present the relevant theoretical results for multinomial logit models (MLM) and generalized linear models (GLM), and demonstrate the superiority of our algorithm over state-of-the-art design algorithms using real-life experiments under MLM and GLM. Our simulation studies show that the ForLion algorithm could reduce the number of experimental settings by 25% or improve the relative efficiency of the designs by 17.5% on average. Our algorithm can help the experimenters reduce the time cost, the usage of experimental devices, and thus the total cost of their experiments while preserving high efficiencies of the designs.
Ranking, and inferences based on ranking of a set of entities, are important problems in numerous contexts. This is especially true in small area statistics where there may be only a limited amount of directly observed data from each entity or small area, while precise and accurate estimates of best or worst performing entities are needed for fund allocation, planning and policymaking, stakeholder advocacy, evaluation of welfare programs, and so on. However, ranks estimates constructed exclusively on point estimates of parameters lack uncertainty quantification, and may lead to imbalances and inequities when these are based on small sample sizes. We propose novel Bayesian approaches to address this problem. Our proposals result in partitions of the parameter space with posterior distribution driven partial ordering of the sets in a partition. This in turn translates to a coherent probability mass function over ranks for every entity, and a coherent probability mass function over entities for every rank. Our Bayesian algorithms significantly outperform the state-of-the-art non-Bayesian alternatives, and are amenable to inclusion of covariates in the model as well as borrowing strengths across small areas. We evaluate our proposed Bayesian algorithms in terms of accuracy and stability using a number of applications and a simulation study. Additionally, we develop a novel theoretical framework for inference and ranking problems involving a triangular array of Fay-Herriot models and data, and provide probabilistic guarantees of performances of the proposed Bayesian ranking algorithms.
A new type of experiment that aims to determine the optimal quantities of a sequence of factors is eliciting considerable attention in medical science, bioengineering, and many other disciplines. Such studies require the simultaneous optimization of both quantities and the sequence orders of several components which are called quantitative-sequence (QS) factors. Given the large and semi-discrete solution spaces in such experiments, efficiently identifying optimal or near-optimal solutions by using a small number of experimental trials is a nontrivial task. To address this challenge, we propose a novel active learning approach, called QS-learning, to enable effective modeling and efficient optimization for experiments with QS factors. QS-learning consists of three parts: a novel mapping-based additive Gaussian process (MaGP) model, an efficient global optimization scheme (QS-EGO), and a new class of optimal designs (QS-design). The theoretical properties of the proposed method are investigated, and optimization techniques using analytical gradients are developed. The performance of the proposed method is demonstrated via a real drug experiment on lymphoma treatment and several simulation studies.
The presence of Arbuscular Mycorrhizal Fungi (AMF) in vascular land plant roots is one of the most ancient of symbioses supporting nitrogen and phosphorus exchange for photosynthetically derived carbon. Here we provide a multi-scale modeling approach to predict AMF colonization of a worldwide crop from a Recombinant Inbred Line (RIL) population derived from Sorghum bicolor and S. propinquum. The high-throughput phenotyping methods of fungal structures here rely on a Mask Region-based Convolutional Neural Network (Mask R-CNN) in computer vision for pixel-wise fungal structure segmentations and mixed linear models to explore the relations of AMF colonization, root niche, and fungal structure allocation. Models proposed capture over 95% of the variation in AMF colonization as a function of root niche and relative abundance of fungal structures in each plant. Arbuscule allocation is a significant predictor of AMF colonization among sibling plants. Arbuscules and extraradical hyphae implicated in nutrient exchange predict highest AMF colonization in the top root section. Our work demonstrates that deep learning can be used by the community for the high-throughput phenotyping of AMF in plant roots. Mixed linear modeling provides a framework for testing hypotheses about AMF colonization phenotypes as a function of root niche and fungal structure allocations.
Computer simulators are often used as a substitute of complex real-life phenomena, which are either expensive or infeasible to experiment with. This paper focuses on how to efficiently solve the inverse problem for an expensive to evaluate time-series-valued computer simulator. The research is motivated by a hydrological simulator which has to be tuned for generating realistic rainfall–runoff measurements in Athens, Georgia, USA. Assuming that the simulator returns g(x, t) over L time points for a given input x, the proposed methodology begins with a careful construction of a discretization (time-) point set of size $$k \ll L$$ , achieved by adopting a regression spline approximation of the target response series at k optimal knot locations $$\{t_1^*, t_2^*,\ldots ,t_k^*\}$$ . Subsequently, we solve k scalar-valued inverse problems for simulator $$g(x,t_j^*)$$ via the contour estimation method. The proposed approach, named MSCE, also facilitates the uncertainty quantification of the inverse solution. Extensive simulation study is used to demonstrate the performance comparison of the proposed method with the popular competitors for several test-function-based computer simulators and a real-life rainfall–runoff measurement model.
With the help of Generalized Estimating Equations, we identify locally D -optimal crossover designs for generalized linear models. We adopt the variance of parameters of interest as the objective function, which is minimized using constrained optimization to obtain optimal crossover designs. In this case, the traditional general equivalence theorem could not be used directly to check the optimality of obtained designs. In this manuscript, we derive a corresponding general equivalence theorem for crossover designs under generalized linear models.
Models with ordinal outcomes are an important part of generalized linear models and design issues for them are less studied, especially when the model has discrete and continuous factors. We propose an effective and flexible Particle Swarm Optimization (PSO) algorithm for finding locally D-optimal approximate designs for experiments with ordinal outcomes modeled using the cumulative logit link. We apply our technique to obtain a locally D-optimal approximate design for an odor removal experiment with both discrete and continuous factors and show that this design is superior to the design obtained by discretizing the continuous factor. Additionally, we find a pseudo-Bayesian D-optimal approximate design for this problem and study the performance of both designs under a range of plausible parameter values. We also (i) demonstrate PSO's versatility by finding locally D-optimal approximate designs for a manufacturing example with surface defects and multiple continuous factors, and (ii) use PSO to find other optimal designs for estimating percentiles in a dose-response study.
This chapter presents a review of some state-of-the-art statistical techniques for analyzing real computer experiments which play a significant role in various scientific research and industrial applications. In computer experiments, emulators (i.e. surrogate models) are often used to rapidly approximate the outcomes and reduce the computational expense. Gaussian process (GP) models, also known as Kriging, are a common choice of emulators, and optimal experimental designs should be used to improve their accuracy. Specifically, space-filling designs are widely used in the literature, which proved to be efficient under GP models. In this chapter, we review different types of GP models as well as various kinds of space-filling designs. We further provide a practical tutorial on how to construct space-filling designs and fit GP emulators to analyze real computer experiments.
We consider locally D-optimal crossover designs for generalized linear models. Three different types of responses were recorded in a work environment experiment conducted at Booking.com . These responses follow Poisson, beta and gamma distributions. The responses from the same subjects are naturally correlated. To capture the dependence among these observations, we use six different types of correlation structures. The optimal allocations of subjects to each treatment sequence are obtained by minimizing the objective function, which is the variance of direct treatment effect estimates. We show that optimal allocations are reasonably robust to a different choice of correlation structures. Although uniform allocations are widely used in practice, we establish these designs are sub-optimal under certain conditions.
BackgroundBisphenol A (BPA) and Diethylhexyl Phthalate (DEHP) are common plastic‐derived environmental contaminants that act as endocrine disrupting chemicals (EDCs) and pose risks for expecting mothers and offspring. Research addressing the effect of a mixture of these EDCs is limited.ObjectiveTo assess the sex‐specific consequences of developmental exposure to individual as well as mixtures of environmentally relevant doses of BPA/DEHP in male and female Sprague‐Dawley rats.MethodsFrom gestational day 6 till 21, dams were orally administered either saline (control), BPA (5mg/Kg BW/day), low‐dose (LD) DEHP (5mg/Kg BW/day), high‐dose (HD) DEHP (7.5 mg/Kg BW/day), a combination of BPA and LD DEHP (B+LD‐D), or a combination of BPA and HD DEHP (B+HD‐D). Gestational weights and number of abortions were tracked. Litter size and weights, number of live births and stillbirths were counted. Morphometric parameters at birth were measured. Body weighs, food and water intakes were monitored weekly from postnatal weeks 3–12. Male and female offspring were sacrificed at 16–24 weeks of age and organs were dissected and weighed.ResultsThe abortion rate of dams exposed to HD DEHP and the mixtures, B+LD‐D and B+HD‐D were higher at 8, 14 and 23% respectively. Gestational index was reduced in these groups as well. Prenatal exposure to BPA or HD DEHP significantly decreased relative thymus weight in male but not female offspring at 16 weeks. Relative heart weight increased in B+HD‐D exposed male offspring compared to the other groups. No programming effects in relative organs weights were detected in female offspring.ConclusionThe results indicate that a mixture of BPA and DEHP, even at low doses, induced a pronounced effect on pregnancy outcomes. Male offspring appear to be more susceptible to the programming effects of these EDCs or their mixture. This suggests a great need to reconsider the possible additive, antagonistic or synergistic effects of EDC mixtures to which pregnant individuals are commonly exposed to.Support or Funding InformationSupported by UGA research foundation.
The objective of this study was to evaluate the effects of gestational exposure to low doses of bisphenol A (BPA), bisphenol S (BPS), and bisphenol F (BPF) on pregnancy outcomes and offspring development. Pregnant Sprague-Dawley rats were orally dosed with vehicle, 5 mu g/kg body weight (BW)/day of BPA, BPS and BPF, or 1 mu g/kg BW/day of BPF on gestational days 6-21. Pregnancy and gestational outcomes, including number of abortions and stillbirths, were monitored. Male and female offspring were subjected to morphometry at birth, followed by pre- and post-weaning body weights, post-weaning food and water intakes, and adult organ weights. Ovarian follicular counts were also obtained from adult female offspring. We observed spontaneous abortions in over 80% of dams exposed to 5 mu g/kg of BPF. BPA exposure increased Graafian follicles in female offspring, while BPS and BPF exposure decreased the number of corpora lutea, suggesting reduced ovulation rates. Moreover, BPA exposure increased male kidney and prostate gland weights, BPF decreased epididymal adipose tissue weights, and BPS had modest effects on male abdominal adipose tissue weights. Prenatal BPS exposure reduced anogenital distance (AGD) in male offspring, suggesting possible feminization, whereas both BPS and BPA induced oxidative stress in the testes. These results indicate that prenatal exposure to BPF affects pregnancy outcomes, BPS alters male AGD, and all three bisphenols alter certain organ weights in male offspring and ovarian function in female offspring. Altogether, it appears that prenatal exposure to BPA or its analogues can induce reproductive toxicity even at low doses. (C) 2021 Elsevier Ltd. All rights reserved.
We introduce statistical techniques required to handle complex computer models with potential applications to astronomy. Computer experiments play a critical role in almost all fields of scientific research and engineering. These computer experiments, or simulators, are often computationally expensive, leading to the use of emulators for rapidly approximating the outcome of the experiment. Gaussian process models, also known as Kriging, are the most common choice of emulator. While emulators offer significant improvements in computation over computer simulators, they require a selection of inputs along with the corresponding outputs of the computer experiment to function well. Thus, it is important to select inputs judiciously for the full computer simulation to construct an accurate emulator. Space-filling designs are efficient when the general response surface of the outcome is unknown, and thus they are a popular choice when selecting simulator inputs for building an emulator. In this tutorial we discuss how to construct these space filling designs, perform the subsequent fitting of the Gaussian process surrogates, and briefly indicate their potential applications to astronomy research.