Animal studies in pharmaceutical discovery and toxicology are not always statistically powered for estimation or hypothesis testing. Typically, only 3 to 5 animals are allocated per group, based on historical conventions or industry practice, particularly in early toxicology studies with several different types of controls and compounds at various concentrations. When we estimate means, variances, or other parameters under these conditions, often the confidence intervals generated will be of little practical use due to the small sample size. If, however, historical or even concurrent data with similar characteristics is available from comparable experiments, all data could be incorporated into the estimation by using an Empirical Bayesian approach. To implement this method, the existing data is used to determine prior distributions for the parameters of interest, which are then combined with the sample data of interest to produce posterior distributions. In our case study, we combined data from 30 different experiments to use as a basis for defining the prior distributions on the mean and standard deviation (SD). For practical reasons related to our application, we prefer to use the standard deviation instead of the variance or precision that are more commonly used in the Bayesian methodology. For the mean parameter, the prior distribution is approximated by a Normal distribution, covering the range of all samples. For SD, the prior distribution is approximated with a half-Normal, half-Cauchy, or Uniform with carefully chosen boundaries. An Empirical Bayes method is then applied, combining the selected prior distributions with observed data in each small experiment to obtain the posterior distribution for the mean and for the variance of that particular experiment. The strategy of using the combined data from multiple samples to develop a common prior distribution that borrows strength across all the available data reduces the variability of the estimates and improves the estimation of individual parameters. In effect, this method combines "borrowing strength" with "Empirical Bayes" in a way that suggests "Tukey meets Robbins"!
This paper introduces the novel methodology of differential projection pursuit and its applications to the analysis of large datasets. The method was applied to a cell flow cytometry dataset as an alternative approach to analyze this type of data. Multicolor cell flow cytometry is a well-established laboratory technique to identify cell subpopulations by measuring their physical and biochemical characteristics. Differential projection pursuit helps to find regions with maximal differences between two or more treatments or distributions. Data analysis in flow cytometry relies on gating, the process of manually selecting successive subpopulations of cells using two-dimensional plots. Plotting the variables only two at a time could mask the hidden structure present in the data, and manual selection makes the analysis inconsistent and arbitrary. The new methodology could automate flow cytometry analysis by utilizing the combination of projection pursuit, data nuggets, and factor analysis. When applied to flow cytometry data, differential projection pursuit allows researchers to quickly identify differences in cell populations exposed to different experimental conditions. This methodology could create a platform to explore differences in large datasets and improve the cell flow cytometry analysis clarity and reproducibility by considering the data in its true dimensional space and through automation, respectively.
We investigated drug-induced acute neuronal electrophysiological changes using Micro-Electrode arrays (MEA) to rat primary neuronal cell cultures. Data based on 6-key MEA parameters were analyzed for plate-to-plate vehicle variability, effects of positive and negative controls, as well as data from over 100 reference drugs, mostly known to have pharmacological phenotypic and clinical outcomes. A Least Absolute Shrinkage and Selection Operator (LASSO) regression, coupled with expert evaluation helped to identify the 6-key parameters from many other MEA parameters to evaluate the drug-induced acute neuronal changes. Calculating the statistical tolerance intervals for negative-positive control effects on those 4-key parameters helped us to develop a new weighted hazard scoring system on drug-induced potential central nervous system (CNS) adverse effects (AEs). The weighted total score, integrating the effects of a drug candidate on the identified six-pivotal parameters, simply determines if the testing compound/concentration induces potential CNS AEs. Hereto, it uses four different categories of hazard scores: non-neuroactive, neuroactive, hazard, or high hazard categories. This new scoring system was successfully applied to differentiate the new compounds with or without CNS AEs, and the results were correlated with the outcome of in vivo studies in mice for one internal program. Furthermore, the Random Forest classification method was used to obtain the probability that the effect of a compound is either inhibitory or excitatory. In conclusion, this new neuronal scoring system on the cell assay is actively applied in the early de-risking of drug development and reduces the use of animals and associated costs.
A complete workflow was presented for estimating the concentration of microorganisms in biological samples by automatically counting spots that represent viral plaque forming units (PFU) bacterial colony forming units (CFU), or spot forming units (SFU) in images, and modeling the counts. The workflow was designed for processing images from dilution series but can also be applied to stand-alone images. The accuracy of the methods was greatly improved by adding a newly developed bias correction method. When the spots in images are densely populated, the probability of spot overlapping increases, leading to systematic undercounting. In this paper, this undercount issue was addressed in an empirical way. The proposed empirical bias correction method utilized synthetic images with known spot sizes and counts as a training set, enabling the development of an effective bias correction function using a thin-plate spline model. Its application focused on the bias correction for the automated spot counting algorithm LoST proposed by Lin et al. Simulation results demonstrated that the empirical bias correction significantly improved spot counts, reducing bias for both fixed and random spot sizes and counts.
Class imbalance is a known issue in classification tasks that can lead to predictive bias toward dominant classes. This paper introduces a novel straightforward Bayesian framework that adjusts posterior probabilities to counteract the bias introduced by imbalanced data sets. Instead of relying on the mean posterior distribution of class probabilities, we propose a method that scales the posterior probability of each class according to their representation in the training data.
The identification of good surrogate endpoints is a challenging endeavor. This may, at least partially, be attributable to the fact that most researchers have focused on the identification of a single surrogate endpoint. It is thus implicitly assumed that the treatment effect on the true endpoint (T) can be accurately predicted based on the treatment effect on one surrogate endpoint (S) only. Given the complex nature of many diseases and the different therapeutic pathways in which a treatment can impact T, this assumption may be too optimistic. For example, in oncology, the effect of a treatment often depends on both the treatment's efficacy and its toxicity. In the present article, the meta-analytic framework of? is extended to the setting where multiple S are considered. To cope with potential model convergence issues that often arise in a meta-analytic framework, several simplified model fitting strategies are proposed. Further, simulation studies are conducted to evaluate the properties of the estimated surrogacy metrics, and the new methodology is applied on a case study in schizophrenia. An online Appendix that details how the analyses can be conducted in practice (using the R package Surrogate) is also provided.
Spatial heterogeneity of cells in liver biopsies can be used as biomarker for disease severity of patients. This heterogeneity can be quantified by non -parametric statistics of point pattern data, which make use of an aggregation of the point locations. The method and scale of aggregation are usually chosen ad hoc, despite values of the aforementioned statistics being heavily dependent on them. Moreover, in the context of measuring heterogeneity, increasing spatial resolution will not endlessly provide more accuracy. The question then becomes how changes in resolution influence heterogeneity indicators, and subsequently how they influence their predictive abilities. In this paper, cell level data of liver biopsy tissue taken from chronic Hepatitis B patients is used to analyze this issue. Firstly, Morisita-Horn indices, Shannon indices and Getis-Ord statistics were evaluated as heterogeneity indicators of different types of cells, using multiple resolutions. Secondly, the effect of resolution on the predictive performance of the indices in an ordinal regression model was investigated, as well as their importance in the model. A simulation study was subsequently performed to validate the aforementioned methods. In general, for specific heterogeneity indicators, a downward trend in predictive performance could be observed. While for local measures of heterogeneity a smaller grid -size is outperforming, global measures have a better performance with medium-sized grids. In addition, the use of both local and global measures of heterogeneity is recommended to improve the predictive performance.
Estimation of microorganism concentration in samples (bacterial cells or viral particles) has been a focal point in biomedical experiments for more than a century. Serial dilution of the samples is often used to estimate the target concentrations in immunology, virology, and pharmaceutical industry. A new methodology, called joint likelihood estimation (JLE), is proposed to estimate particles such as the number of microorganisms in a sample from counts obtained by serially diluting the sample. It models count data from the entire single dilution series rather than using only specific dilutions. The theoretical framework is based on the binomial and the Poisson distributions and is consistent with the actual experimental process. The estimator of the target concentration is obtained by MLE with derived joint likelihood functions of the observed counts including right‐censored values. Simulations demonstrated that the new JLE method significantly increases precision and accuracy of the estimate compared to the existing methods. It can be applied to a variety of studies with similar experimental designs, especially when the number of particles in the neat sample is very large.
Biological samples are routinely analyzed for microbe concentration. The samples are diluted, loaded onto established host cell cultures, and incubated. If infectious agents are present in the samples, they form circular spots that do not contain the host cells. Each spot is assumed to be originated from a single microbial unit such as a bacterial colony forming unit or viral plaque forming unit. The undiluted sample concentration is estimated by counting the spots and back-calculating. Counting the number of spots by trained technicians is currently the gold standard but it is laborious, subjective, and hard to scale. This paper presents a new automated algorithm for spot counting, Localized and Sequential Thresholding (LoST). Validation studies showed that LoST performance was comparable with manual counting and outperformed several existing tools on images with overlapping spots. The LoST algorithm employs sequential thresholding through a two-stage segmentation and borrows information across all images from the same dilution series to fine-tune the count and identify right censoring. The algorithm increases the efficiency of the spot counting and the quality of the downstream analysis, especially when coupled with an appropriate statistical serial dilution model to enhance the undiluted sample concentration estimation procedure.
The organization and interaction between hepatocytes and other hepatic non-parenchymal cells plays a pivotal role in maintaining normal liver function and structure. Although spatial heterogeneity within the tumor micro-environment has been proven to be a fundamental feature in cancer progression, the role of liver tissue topology and micro-environmental factors in the context of liver damage in chronic infection has not been widely studied yet. We obtained images from 110 core needle biopsies from a cohort of chronic hepatitis B patients with different fibrosis stages according to METAVIR score. The tissue sections were immunofluorescently stained and imaged to determine the locations of CD45 positive immune cells and HBsAg-negative and HBsAg-positive hepatocytes within the tissue. We applied several descriptive techniques adopted from ecology, including Getis-Ord, the Shannon Index and the Morisita-Horn Index, to quantify the extent to which immune cells and different types of liver cells co-localize in the tissue biopsies. Additionally, we modeled the spatial distribution of the different cell types using a joint log-Gaussian Cox process and proposed several features to quantify spatial heterogeneity. We then related these measures to the patient fibrosis stage by using a linear discriminant analysis approach. Our analysis revealed that the co-localization of HBsAg-negative hepatocytes with immune cells and the co-localization of HBsAg-positive hepatocytes with immune cells are equally important factors for explaining the METAVIR score in chronic hepatitis B patients. Moreover, we found that if we allow for an error of 1 on the METAVIR score, we are able to reach an accuracy of around 80%. With this study we demonstrate how methods adopted from ecology and applied to the liver tissue micro-environment can be used to quantify heterogeneity and how these approaches can be valuable in biomarker analyses for liver topology.
Drug-induced liver injury (DILI), believed to be a multifactorial toxicity, has been a leading cause of attrition of small molecules during discovery, clinical development, and postmarketing. Identification of DILI risk early reduces the costs and cycle times associated with drug development. In recent years, several groups have reported predictive models that use physicochemical properties or in vitro and in vivo assay endpoints; however, these approaches have not accounted for liver-expressed proteins and drug molecules. To address this gap, we have developed an integrated artificial intelligence/machine learning (AI/ML) model to predict DILI severity for small molecules using a combination of physicochemical properties and off-target interactions predicted in silico. We compiled a data set of 603 diverse compounds from public databases. Among them, 164 were categorized as Most DILI (M-DILI), 245 as Less DILI (L-DILI), and 194 as No DILI (N-DILI) by the FDA. Six machine learning methods were used to create a consensus model for predicting the DILI potential. These methods include k-nearest neighbor (k-NN), support vector machine (SVM), random forest (RF), Naïve Bayes (NB), artificial neural network (ANN), logistic regression (LR), weighted average ensemble learning (WA) and penalized logistic regression (PLR). Among the analyzed ML methods, SVM, RF, LR, WA, and PLR identified M-DILI and N-DILI compounds, achieving a receiver operating characteristic area under the curve of 0.88, sensitivity of 0.73, and specificity of 0.9. Approximately 43 off-targets, along with physicochemical properties (fsp3, log S, basicity, reactive functional groups, and predicted metabolites), were identified as significant factors in distinguishing between M-DILI and N-DILI compounds. The key off-targets that we identified include: PTGS1, PTGS2, SLC22A12, PPARγ, RXRA, CYP2C9, AKR1C3, MGLL, RET, AR, and ABCC4. The present AI/ML computational approach therefore demonstrates that the integration of physicochemical properties and predicted on- and off-target biological interactions can significantly improve DILI predictivity compared to chemical properties alone.
aJanssen R&D, Raritan, NJ; bPrinceton Data Analytics, Bridgewater, NJ; cRutgers University, New Brunswick, NJ; dPfizer Inc, Cambridge, MA; eJanssen R&D, Beerse, Belgium; fCMCStats, Wadsworth, Il; gPfizer Inc, Groton, CT; hIncyte Corp, Wilmington, DE; iPfizer Inc, Andover, MA; jAmgen Inc, Thousand Oaks, CA; kAstraZeneca, Gaithersburg, MD; lJanssen R&D, Prague, Czechia; mPDQ Research & Consulting, Birdsboro, PA; nRoche Diagnostics Gmbh, Penzberg, Germany; oCMC Sciences LLC, Germantown, MD; pPfizer Inc, Collegeville, PA; qCovance by Labcorp, Huntingdon, UK
The accurate prediction of binding affinity between protein and small molecules with free energy methods, particularly the difference in binding affinities via relative binding free energy calculations, has undergone a dramatic increase in use and impact over recent years. The improvements in methodology, hardware, and implementation can deliver results with less than 1 kcal/mol mean unsigned error between calculation and experiment. This is a remarkable achievement and beckons some reflection on the significance of calculation approaching the accuracy of experiment. In this article, we describe a statistical analysis of the implications of variance (standard deviation) of both experimental and calculated binding affinities with respect to the unknown true binding affinity. We reveal that plausible ratios of standard deviation in experiment and calculation can lead to unexpected outcomes for assessing the performance of predictions. The work extends beyond the case of binding free energies to other affinity or property prediction methods.
Pharmacokinetic (PK) studies are conducted to learn about the absorption, distribution, metabolism, and excretion processes of an externally administered compound by measuring its concentration in bodily tissue at a number of time points after administration. Two methods are available for this analysis: modeling and non-compartmental. When concentrations of the compound are low, they may be reported as below the limit of quantification (BLOQ). This article compares eight methods for dealing with BLOQ responses in the non-compartmental analysis framework for estimating the area under the concentrations versus time curve. These include simple methods that are currently used, maximum likelihood methods, and an algorithm that uses kernel density estimation to impute values for BLOQ responses. Performance is evaluated using simulations for a range of scenarios. We find that the kernel based method performs best for most situations. for this article are available online.
The most common objective for response adaptive clinical trials is to seek to ensure that patients within a trial have a high chance of receiving the best treatment available by altering the chance of allocation on the basis of accumulating data. Approaches which yield good patient benefit properties suffer from low power from a frequentist perspective when testing for a treatment difference at the end of the study due to the high imbalance in treatment allocations. In this work we develop an alternative pairwise test for treatment difference on the basis of allocation probabilities of the covariate-adjusted response-adaptive randomization with forward looking Gittins index rule (CARA-FLGI). The performance of the novel test is evaluated in simulations for two-armed studies and then its applications to multi-armed studies is illustrated. The proposed test has markedly improved power over the traditional Fisher exact test when this class of non-myopic response adaptation is used. We also find that the test's power is close to the power of a Fisher exact test under equal randomization.
Bayesian sequential integration is an appealing approach in drug development, as it allows to recursively update posterior distributions as soon as new data become available, thus considerably reducing the computation time. However, preclinical trials are often characterized by small sample sizes, which may affect the estimation process during the first integration steps, particularly when complex PK-PD models are used. In this case, sequential integration would not be practicable, and trials should be pooled together. This work is aimed at comparing simple Bayesian pooling with sequential integration through a simulation study. The two techniques are compared under several scenarios using linear as well as nonlinear models. The results of our simulation study encourage the use of Bayesian sequential integration with linear models. However, in the case of nonlinear models several caveats arise. This paper outlines some important recommendations and precautions in that respect.
Acetaminophen, a nonmutagenic compound as previously concluded from bacteria, in vitro mammalian cell, and in vivo transgenic rat assays, presented a good profile as a nonmutagenic reference compound for use in the international multilaboratory Pig-a assay validation. Acetaminophen was administered at 250, 500, 1,000, and 2,000 mg center dot kg(-1)center dot day(-1) to male Sprague Dawley rats once daily in 3 studies (3 days, 2 weeks, and 1 month with a 1-month recovery group). The 3-Day and 1-Month Studies included assessments of the micronucleus endpoint in peripheral blood erythrocytes and the comet endpoint in liver cells and peripheral blood cells in addition to the Pig-a assay; appropriate positive controls were included for each assay. Within these studies, potential toxicity of acetaminophen was evaluated and confirmed by inclusion of liver damage biomarkers and histopathology. Blood was sampled pre-treatment and at multiple time points up to Day 57. Pig-a mutant frequencies were determined in total red blood cells (RBCs) and reticulocytes (RETs) as CD59-negative RBC and CD59-negative RET frequencies, respectively. No increases in DNA damage as indicated through Pig-a, micronucleus, or comet endpoints were seen in treated rats. All positive controls responded as appropriate. Data from this series of studies demonstrate that acetaminophen is not mutagenic in the rat Pig-a model. These data are consistent with multiple studies in other nonclinical models, which have shown that acetaminophen is not mutagenic. At 1,000 mg center dot kg(-1)center dot day(-1), Cmax values of acetaminophen on Day 28 were 153,600 ng/ml and 131,500 ng/ml after single and repeat dosing, respectively, which were multiples over that of clinical therapeutic exposures (2.6-6.1 fold for single doses of 4,000 mg and 1,000 mg, respectively, and 11.5 fold for multiple dose of 4,000 mg) (FDA 2002). Data generated were of high quality and valid for contribution to the international multilaboratory validation of the in vivo Rat Pig-a Mutation Assay.
-In spite of medical and methodological advances, the identification of good surrogate endpoints has remained a challenging endeavor. This may, at least partially, be attributable to the fact that most researchers have only focused on univariate surrogates endpoints. In the present work, we argue in favor of using multivariate surrogates and introduce two new complementary metrics to assess their validity. The first one, the so-called individual causal association, quantifies the association between the individual causal treatment effects on the multivariate surrogate and true endpoints, while the second one quantifies the treatment-corrected association between the multivariate surrogate and the true endpoint outcomes. The newly proposed methodology is implemented in the R package Surrogate and a Web Appendix, detailing how the analysis can be conducted in practice, is provided. Supplementary materials for this article are available online.