Spatial heterogeneity of cells in liver biopsies can be used as biomarker for disease severity of patients. This heterogeneity can be quantified by non -parametric statistics of point pattern data, which make use of an aggregation of the point locations. The method and scale of aggregation are usually chosen ad hoc, despite values of the aforementioned statistics being heavily dependent on them. Moreover, in the context of measuring heterogeneity, increasing spatial resolution will not endlessly provide more accuracy. The question then becomes how changes in resolution influence heterogeneity indicators, and subsequently how they influence their predictive abilities. In this paper, cell level data of liver biopsy tissue taken from chronic Hepatitis B patients is used to analyze this issue. Firstly, Morisita-Horn indices, Shannon indices and Getis-Ord statistics were evaluated as heterogeneity indicators of different types of cells, using multiple resolutions. Secondly, the effect of resolution on the predictive performance of the indices in an ordinal regression model was investigated, as well as their importance in the model. A simulation study was subsequently performed to validate the aforementioned methods. In general, for specific heterogeneity indicators, a downward trend in predictive performance could be observed. While for local measures of heterogeneity a smaller grid -size is outperforming, global measures have a better performance with medium-sized grids. In addition, the use of both local and global measures of heterogeneity is recommended to improve the predictive performance.
Supplementary Figure 2 from Response prediction to a multitargeted kinase inhibitor in cancer cell lines and xenograft tumors using high-content tyrosine peptide arrays with a kinetic readout
The organization and interaction between hepatocytes and other hepatic non-parenchymal cells plays a pivotal role in maintaining normal liver function and structure. Although spatial heterogeneity within the tumor micro-environment has been proven to be a fundamental feature in cancer progression, the role of liver tissue topology and micro-environmental factors in the context of liver damage in chronic infection has not been widely studied yet. We obtained images from 110 core needle biopsies from a cohort of chronic hepatitis B patients with different fibrosis stages according to METAVIR score. The tissue sections were immunofluorescently stained and imaged to determine the locations of CD45 positive immune cells and HBsAg-negative and HBsAg-positive hepatocytes within the tissue. We applied several descriptive techniques adopted from ecology, including Getis-Ord, the Shannon Index and the Morisita-Horn Index, to quantify the extent to which immune cells and different types of liver cells co-localize in the tissue biopsies. Additionally, we modeled the spatial distribution of the different cell types using a joint log-Gaussian Cox process and proposed several features to quantify spatial heterogeneity. We then related these measures to the patient fibrosis stage by using a linear discriminant analysis approach. Our analysis revealed that the co-localization of HBsAg-negative hepatocytes with immune cells and the co-localization of HBsAg-positive hepatocytes with immune cells are equally important factors for explaining the METAVIR score in chronic hepatitis B patients. Moreover, we found that if we allow for an error of 1 on the METAVIR score, we are able to reach an accuracy of around 80%. With this study we demonstrate how methods adopted from ecology and applied to the liver tissue micro-environment can be used to quantify heterogeneity and how these approaches can be valuable in biomarker analyses for liver topology.
Supplementary Table 1 from Response prediction to a multitargeted kinase inhibitor in cancer cell lines and xenograft tumors using high-content tyrosine peptide arrays with a kinetic readout
Method: Liver tissue from C57/Bl6 and HBV-transgenic mouse strains (tg[Alb1HBV]44Bri [Alb/HBs], tg1.4HBV-s-mut and tg1.4HBV-s-rec [F1 generation Alb/HBs × tg1.4HBV-s-mut]) was analysed using transmission electron microscopy (TEM).HBV replication was characterized on RNA and protein level.All-in-one liver cell preparation enabled extraction of primary murine hepatocytes, liver sinusoidal endothelial cells and Kupffer cells and remaining non-parenchymal cells (rNPC).Responsiveness to poly (I:C) treatment (Tlr3) and transfection (Rig-I) was determined using quantitative RT-PCR and multiplex-based cytokine assay.Results: TEM visualised viral particles (tg1.4HBV-s-rec),nuclear spot formations (tg1.4HBV-s-mut and tg1.4HBV-s-rec) and malformation of the endoplasmic reticulum (Alb/HBs).Viral replication did not differ comparing tg1.4HBV-s-rec and tg1.4HBV-s-mut, except HBsAg levels.A distinct cell type-dependent and mouse strain-dependent gene expression pattern for interferon (Ifnb, Ifiti1), cytokine (Tnf, Il1b, Il10) and chemokine (Ccl2, Ccl5) was observed by quantitative RT-PCR and validated using a multiplex-based cytokine assay.Emphasizing that tg1.4HBV-s-rec-derived liver cells responded to poly (I:C) with significantly suppressed expression levels, compared to the other mouse strains.Contrary, rNPC of tg1.4HBVs-rec mice showed the highest induction for Il10, Tnf, Il1b and Ccl5 in response to poly (I:C).In all HBV-transgenic mice, the induction of Ccl2 was significant lower in PMH but increased in rNPC.Conclusion: Production of HBV particles and release of HBsAg in tg1.4HBV-s-rec mice led to suppressed hepatic Tlr3 and Rig-I signaling.However, rNPC seemed be more potent in producing poly (I:C)-induced inflammatory cytokines.
Phenomic profiles are high-dimensional sets of readouts that can comprehensively capture the biological impact of chemical and genetic perturbations in cellular assay systems. Phenomic profiling of compound libraries can be used for compound target identification or mechanism of action (MoA) prediction and other applications in drug discovery. To devise an economical set of phenomic profiling assays, we assembled a library of 1,008 approved drugs and well-characterized tool compounds manually annotated to 218 unique MoAs, and we profiled each compound at four concentrations in live-cell, high-content imaging screens against a panel of 15 reporter cell lines, which expressed a diverse set of fluorescent organelle and pathway markers in three distinct cell lineages. For 41 of 83 testable MoAs, phenomic profiles accurately ranked the reference compounds (AUC-ROC ≥0.9). MoAs could be better resolved by screening compounds at multiple concentrations than by including replicates at a single concentration. Screening additional cell lineages and fluorescent markers increased the number of distinguishable MoAs but this effect quickly plateaued. There remains a substantial number of MoAs that were hard to distinguish from others under the current study’s conditions. We discuss ways to close this gap, which will inform the design of future phenomic profiling efforts.
Most candidate drugs currently fail later-stage clinical trials, largely due to poor prediction of efficacy on early target selection(1). Drug targets with genetic support are more likely to be therapeutically valid(2,3), but the translational use of genome-scale data such as from genome-wide association studies for drug target discovery in complex diseases remains challenging(4-6). Here, we show that integration of functional genomic and immune-related annotations, together with knowledge of network connectivity, maximizes the informativeness of genetics for target validation, defining the target prioritization landscape for 30 immune traits at the gene and pathway level. We demonstrate how our genetics-led drug target prioritization approach (the priority index) successfully identifies current therapeutics, predicts activity in high-throughput cellular screens (including L1000, CRISPR, mutagenesis and patient-derived cell assays), enables prioritization of under-explored targets and allows for determination of target-level trait relationships. The priority index is an open-access, scalable system accelerating early-stage drug target selection for immune-mediated disease.
Broad toxicology profiling takes traditionally place at the interface between discovery and development when a potential drug candidate is selected. However, it would be both time- and cost-wise better if mechanism (target)-related toxicity and compound-chemistry-related toxicity are addressed earlier, when discussions on novel drug targets take place and compound series are identified and optimized. As the traditional in vivo and in vitro toxicity testing is rather low throughput, they cannot be used in these early stages of the drug discovery process. Therefore a paradigm shift in toxicity testing needs to take place to move to high-throughput cell-based assays to reveal key pathways and proteins linked with toxicity endpoints. In this proof-of-concept work, we explore both transcriptional profiling and imaging techniques to flag early potential toxicity issues. We provide findings of an in-depth exploration of a single chemical core structure. A subset of close analogs was identified via transcriptional profiling, which commonly downregulate multiple tubulin genes across cellular contexts, suggesting possible spindle poison effects. The biological relevance of the finding was confirmed by two different approaches: firstly, high content imaging using cells endogenously tagged with fluorescent tubulin protein genes flagged potential issues via the identification of a characteristic aggregate-formation phenotype. Secondly, a qualified toxicity assay (in vitro micronucleus test) showed increased micronucleus induction for selected compounds showing tubulin downregulation. SAR analysis triggered the synthesis of a new set of compounds and allowed us to extend the series showing the genotoxic effect. This indicates how medicinal chemistry can be guided to dial out undesired effects, thereby optimizing candidate selection. We conclude the chapter by discussing how this approach may be incorporated into state-of-the-art drug development strategies where data integration plays an important role.
By adding biological information, beyond the chemical properties and desired effect of a compound, uncharted compound areas and connections can be explored. In this study, we add transcriptional information for 31K compounds of Janssen's primary screening deck, using the HT L1000 platform and assess (a) the transcriptional connection score for generating compound similarities, (b) machine learning algorithms for generating target activity predictions, and (c) the scaffold hopping potential of the resulting hits. We demonstrate that the transcriptional connection score is best computed from the significant genes only and should be interpreted within its confidence interval for which we provide the stats. These guidelines help to reduce noise, increase reproducibility, and enable the separation of specific and promiscuous compounds. The added value of machine learning is demonstrated for the NR3C1 and HSP90 targets. Support Vector Machine models yielded balanced accuracy values ≥80% when the expression values from DDIT4 & SERPINE1 and TMEM97 & SPR were used to predict the NR3C1 and HSP90 activity, respectively. Combining both models resulted in 22 new and confirmed HSP90-independent NR3C1 inhibitors, providing two scaffolds (i.e., pyrimidine and pyrazolo-pyrimidine), which could potentially be of interest in the treatment of depression (i.e., inhibiting the glucocorticoid receptor (i.e., NR3C1), while leaving its chaperone, HSP90, unaffected). As such, the initial hit rate increased by a factor 300, as less, but more specific chemistry could be screened, based on the upfront computed activity predictions.
BACKGROUND:Alternative gene splicing is a common phenomenon in which a single gene gives rise to multiple transcript isoforms. The process is strictly guided and involves a multitude of proteins and regulatory complexes. Unfortunately, aberrant splicing events do occur which have been linked to genetic disorders, such as several types of cancer and neurodegenerative diseases (Fan et al., Theor Biol Med Model 3:19, 2006). Therefore, understanding the mechanism of alternative splicing and identifying the difference in splicing events between diseased and healthy tissue is crucial in biomedical research with the potential of applications in personalized medicine as well as in drug development.RESULTS:We propose a linear mixed model, Random Effects for the Identification of Differential Splicing (REIDS), for the identification of alternative splicing events. Based on a set of scores, an exon score and an array score, a decision regarding alternative splicing can be made. The model enables the ability to distinguish a differential expressed gene from a differential spliced exon. The proposed model was applied to three case studies concerning both exon and HTA arrays.CONCLUSION:The REIDS model provides a work flow for the identification of alternative splicing events relying on the established linear mixed model. The model can be applied to different types of arrays.
Slc17a5−/− mice represent an animal model for the infantile form of sialic acid storage disease (SASD). We analyzed genetic and histological time-course expression of myelin and oligodendrocyte (OL) lineage markers in different parts of the CNS, and related this to postnatal neurobehavioral development in these mice. Sialin-deficient mice display a distinct spatiotemporal pattern of sialic acid storage, CNS hypomyelination and leukoencephalopathy. Whereas few genes are differentially expressed in the perinatal stage (p0), microarray analysis revealed increased differential gene expression in later postnatal stages (p10–p18). This included progressive upregulation of neuroinflammatory genes, as well as continuous down-regulation of genes that encode myelin constituents and typical OL lineage markers. Age-related histopathological analysis indicates that initial myelination occurs normally in hindbrain regions, but progression to more frontal areas is affected in Slc17a5−/− mice. This course of progressive leukoencephalopathy and CNS hypomyelination delays neurobehavioral development in sialin-deficient mice. Slc17a5−/− mice successfully achieve early neurobehavioral milestones, but exhibit progressive delay of later-stage sensory and motor milestones. The present findings may contribute to further understanding of the processes of CNS myelination as well as help to develop therapeutic strategies for SASD and other myelination disorders.
The modern drug discovery process involves multiple sources of high-dimensional data. This imposes the challenge of data integration. A typical example is the integration of chemical structure (fingerprint features), phenotypic bioactivity (bioassay read-outs) data for targets of interest, and transcriptomic (gene expression) data in early drug discovery to better understand the chemical and biological mechanisms of candidate drugs, and to facilitate early detection of safety issues prior to later and expensive phases of drug development cycles. In this paper, we discuss a joint model for the transcriptomic and the phenotypic variables conditioned on the chemical structure. This modeling approach can be used to uncover, for a given set of compounds, the association between gene expression and biological activity taking into account the influence of the chemical structure of the compound on both variables. The model allows to detect genes that are associated with the bioactivity data facilitating the identification of potential genomic biomarkers for compounds efficacy. In addition, the effect of every structural feature on both genes and pIC50 and their associations can be simultaneously investigated. Two oncology projects are used to illustrate the applicability and usefulness of the joint model to integrate multi-source high-dimensional information to aid drug discovery.
The modern process of discovering candidate molecules in early drug discovery phase includes a wide range of approaches to extract vital information from the intersection of biology and chemistry. A typical strategy in compound selection involves compound clustering based on chemical similarity to obtain representative chemically diverse compounds (not incorporating potency information). In this paper, we propose an integrative clustering approach that makes use of both biological (compound efficacy) and chemical (structural features) data sources for the purpose of discovering a subset of compounds with aligned structural and biological properties. The datasets are integrated at the similarity level by assigning complementary weights to produce a weighted similarity matrix, serving as a generic input in any clustering algorithm. This new analysis work flow is semi-supervised method since, after the determination of clusters, a secondary analysis is performed wherein it finds differentially expressed genes associated to the derived integrated cluster(s) to further explain the compound-induced biological effects inside the cell. In this paper, datasets from two drug development oncology projects are used to illustrate the usefulness of the weighted similarity-based clustering approach to integrate multi-source high-dimensional information to aid drug discovery. Compounds that are structurally and biologically similar to the reference compounds are discovered using this proposed integrative approach.
The NIH-funded LINCS program has been initiated to generate a library of integrated, network-based, cellular signatures (LINCS). A novel high-throughput gene-expression profiling assay known as L1000 was the main technology used to generate more than a million transcriptional profiles. The profiles are based on the treatment of 14 cell lines with one of many perturbation agents of interest at a single concentration for 6 and 24 hours duration. In this study, we focus on the chemical compound treatments within the LINCS data set. The experimental variables available include number of replicates, cell lines, and time points. Our study reveals that compound characterization based on three cell lines at two time points results in more genes being affected than six cell lines at a single time point. Based on the available LINCS data, we conclude that the most optimal experimental design to characterize a large set of compounds is to test them in duplicate in three different cell lines. Our conclusions are constrained by the fact that the compounds were profiled at a single, relative high concentration, and the longer time point is likely to result in phenotypic rather than mechanistic effects being recorded.
With substantial numbers of breast tumors showing or acquiring treatment resistance, it is of utmost importance to develop new agents for the treatment of the disease, to know their effectiveness against breast cancer and to understand their relationships with other drugs to best assign the right drug to the right patient. To achieve this goal drug screenings on breast cancer cell lines are a promising approach. In this study a large-scale drug screening of 37 compounds was performed on a panel of 42 breast cancer cell lines representing the main breast cancer subtypes. Clustering, correlation and pathway analyses were used for data analysis. We found that compounds with a related mechanism of action had correlated IC50 values and thus grouped together when the cell lines were hierarchically clustered based on IC50 values. In total we found six clusters of drugs of which five consisted of drugs with related mode of action and one cluster with two drugs not previously connected. In total, 25 correlated and four anti-correlated drug sensitivities were revealed of which only one drug, Sirolimus, showed significantly lower IC50 values in the luminal/ERBB2 breast cancer subtype. We found expected interactions but also discovered new relationships between drugs which might have implications for cancer treatment regimens.