Copy number variations (CNVs) are major structural genomic variants that contribute to a wide range of human diseases. Accurate detection of CNVs from whole-exome sequencing (WES) data has been a long-sought goal for clinical and population genetic studies. Despite recent progress, existing WES-based CNV callers still suffer from high false-positive rates and reduced recall for short-length variants, and current deep learning methods have not fully used complementary information in region-level genomic features. Here we present CN-RNN, a deep learning-based CNV caller for WES data. The model combines a bidirectional long short-term memory (BiLSTM) branch that captures local depth changes and contextual dependencies across neighboring exons with a parallel multi-layer perceptron (MLP) branch that encodes region-level metadata such as GC content, mappability, and exon length. CN-RNN was trained on the Autism Sequencing Consortium (ASC) parent-child trio cohort using the Mendelian rule of inheritance to ensure high-quality training sets. It was evaluated across three independent datasets, in which we showed that CN-RNN outperformed existing WES-based CNV callers and deep learning methods. CN-RNN offers a scalable, accurate tool for CNV profiling in WES-based studies and supports broader application of CNV analysis in population and clinical research. CN-RNN is available at https://github.com/FeifeiXiao-lab/CN-RNN.
The emergence of spatial transcriptomic technologies has opened new avenues for investigating gene activities while preserving the spatial context of tissues. Utilizing data generated by such technologies, the identification of spatially variable (SV) genes is an essential step in exploring tissue landscapes and biological processes. Particularly in typical experimental designs, such as case-control or longitudinal studies, identifying SV genes between groups is crucial for discovering significant biomarkers or developing targeted therapies for diseases. However, current methods available for analyzing spatial transcriptomic data are still in their infancy, and none of the existing methods are capable of identifying SV genes between groups. To overcome this challenge, we developed SPADE for spatial pattern and differential expression analysis to identify SV genes in spatial transcriptomic data. SPADE is based on a machine learning model of Gaussian process regression with a gene-specific Gaussian kernel, enabling the detection of SV genes both within and between groups. Through benchmarking against existing methods in extensive simulations and real data analyses, we demonstrated the preferred performance of SPADE in detecting SV genes within and between groups. The SPADE source code and documentation are publicly available at https://github.com/thecailab/SPADE.
The availability of single cell sequencing (SCS) enables us to assess intra-tumor heterogeneity and identify cellular subclones without the confounding effect of mixed cells. Copy number aberrations (CNAs) have been commonly used to identify subclones in SCS data using various clustering methods, since cells comprising a subpopulation are found to share genetic profile. However, currently available methods may generate spurious results (e.g., falsely identified CNAs) in the procedure of CNA detection, hence diminishing the accuracy of subclone identification from a large complex cell population. In this study, we developed a CNA detection method based on a fused lasso model, referred to as FLCNA, which can simultaneously identify subclones in single cell DNA sequencing (scDNA-seq) data. Spike-in simulations were conducted to evaluate the clustering and CNA detection performance of FLCNA benchmarking to existing copy number estimation methods (SCOPE, HMMcopy) in combination with the existing and commonly used clustering methods. Interestingly, application of FLCNA to a real scDNA-seq dataset of breast cancer revealed remarkably different genomic variation patterns in neoadjuvant chemotherapy treated samples and pre-treated samples. We show that FLCNA is a practical and powerful method in subclone identification and CNA detection with scDNA-seq data.
Gene expression in mammalian cells is inherently stochastic and mRNAs are synthesized in discrete bursts. Single-cell transcriptomics provides an unprecedented opportunity to explore the transcriptome-wide kinetics of transcriptional bursting. However, current analysis methods provide limited accuracy in bursting inference due to substantial noise inherent to single-cell transcriptomic data. In this study, we developed BISC, a Bayesian method for inferring bursting parameters from single cell transcriptomic data. Based on a beta-gamma-Poisson model, BISC modeled the mean-variance dependency to achieve accurate estimation of bursting parameters from noisy data. Evaluation based on both simulation and real intron sequential RNA fluorescence in situ hybridization data showed improved accuracy and reliability of BISC over existing methods, especially for genes with low expression values. Further application of BISC found bursting frequency but not bursting size was strongly associated with gene expression regulation. Moreover, our analysis provided new mechanistic insights into the functional role of enhancer and superenhancer by modulating both bursting frequency and size. BISC also formulated a downstream framework to identify differential bursting (in frequency and size separately) genes in samples under different conditions. Applying to multiple datasets (a mouse embryonic cell and fibroblast dataset, a human immune cell dataset and a human pancreatic cell dataset), BISC identified known cell-type signature genes that were missed by differential expression analysis, providing additional insights in understanding the cell-specific stochastic gene transcription. Applying to datasets of human lung and colon cancers, BISC successfully detected tumor signature genes based on alterations in bursting kinetics, which illustrates its value in understanding disease development regarding transcriptional bursting. Collectively, BISC provides a new tool for accurately inferring bursting kinetics and detecting differential bursting genes. This study also produced new insights in the role of transcriptional bursting in regulating gene expression, cell identity and tumor progression.
Abstract Tumor tissues are heterogeneous with different cell types in tumor microenvironment, which play an important role in tumorigenesis and tumor progression. Several computational algorithms and tools have been developed to infer the cell composition from bulk transcriptome profiles. However, they ignore the tissue specificity and thus a new resource for tissue-specific cell transcriptomic reference is needed for inferring cell composition in tumor microenvironment and exploring their association with clinical outcomes and tumor omics. In this study, we developed SCISSOR™ (https://thecailab.com/scissor/), an online open resource to fulfill that demand by integrating five orthogonal omics data of >6031 large-scale bulk samples, patient clinical outcomes and 451 917 high-granularity tissue-specific single-cell transcriptomic profiles of 16 cancer types. SCISSOR™ provides five major analysis modules that enable flexible modeling with adjustable parameters and dynamic visualization approaches. SCISSOR™ is valuable as a new resource for promoting tumor heterogeneity and tumor–tumor microenvironment cell interaction research, by delineating cells in the tissue-specific tumor microenvironment and characterizing their associations with tumor omics and clinical outcomes.
Research ObjectiveThe Medicare Shared Savings Program (MSSP) encourages participating Accountable Care Organizations (ACOs) to deliver quality care at lower cost through shared savings models. While the MSSP model offers ACOs and the Centers for Medicare and Medicaid Services (CMS) an opportunity to reduce spending growth while promoting quality goals, ACOs need explicit guidance about those activities over which they have control (e.g. program offerings, care pathways) to improve upon cost and quality goals. Variation amongst geographically‐adjusted and risk‐adjusted total per beneficiary spending amongst ACOs suggests room for ACOs to consider further cost saving efforts. Existing studies do not clearly point to how to achieve those ends. This study presents a methodology for understanding MSSP cost trajectories and identifying those factors that may impact them.Study DesignThis study seeks to characterize and identify trajectories of health care spending for individual patients with type 2 diabetes who experience hospitalization for a major acute cardiovascular event (MACE) (e.g. myocardial infarction, stroke). This analysis uses a retrospective cohort design with the cohort created by an index event (MACE). We examine spending variability for at least six months following the acute event to estimate variability pre and post the event and to cluster patients into post‐event cost trajectories. We use a nonparametric random forest model to identify characteristics from prior to the index event that predict changes in cost trajectories post‐event. We test finding robustness using a RE‐EM model that allows for autocorrelation and combines the structure of mixed effects models for longitudinal data with the flexibility of tree‐based estimation methods. These models allow for testing a broad set of characteristics. These include patient characteristics (e.g. gender, age, percent in zipcode below federal poverty line), comorbid conditions (e.g. dementia, Charlson co‐morbidity index), service use (e.g. prescription count, eye exam, skilled nursing facility use), preventive activities (e.g. flu shot, mammogram) and clinical indicators (e.g. blood pressure, BMI).Population StudiedThis study uses claims and electronic medical record (EMR) data from 2015–2017 for an MSSP population from the largest ACO in South Carolina. In 2017, there were 58,472 attributed beneficiaries. 2015–2017 was one contract period for the ACO and an upside risk‐only arrangement (Track 1).Principal Findings4064 individuals met exclusionary criteria prior to the event (e.g. continuously enrolled from 2015–2017, no prior MACE, index event is non‐fatal). Fisher's F‐test finds strong evidence of variation (P value <2.2e‐16 [F = 19.516]) in costs for the sample before and after the index event. Further work to be presented will cluster patients by post‐event cost trajectory and examine the impact of characteristics (e.g eye exam, falls risk assessment) prior to the index event on changes in post‐event health care spending.ConclusionsVariability exists in our sample pre‐ and post‐index event suggesting reason to investigate drivers of differences in patient characteristics prior to their MACE across cost trajectory clusters.Implications for Policy or PracticeThis work contributes a methodology for examining patient cost trajectories that allows health systems and policy makers to point to specific services and programs for interventions that suggests ways for that cost trajectory to be intervened upon.Primary Funding SourceUniversity of South Carolina internal funding.
Clinical evidence shows that chronic pain and depression often accompany each other, but the underlying pathogenesis of comorbid chronic pain and depression remains mostly undetermined. Biotechnology is gradually revealing the phenotype and function of microglia, with great progress regarding microglia's role in neurodegeneration, depression, chronic pain, and other conditions. This article summarizes the role of microglia in chronic pain, depression, and comorbidities, which is conducive to finding new targets to treat chronic pain and depression.
The current spreading coronavirus SARS-CoV-2 is highly infectious and pathogenic. In this study, we screened the gene expression of three host receptors (ACE2, DC-SIGN and L-SIGN) of SARS coronaviruses and dendritic cells (DCs) status in bulk and single cell transcriptomic datasets of upper airway, lung or blood of COVID-19 patients and healthy controls. In COVID-19 patients, DC-SIGN gene expression was interestingly decreased in lung DCs but increased in blood DCs. Within DCs, conventional DCs (cDCs) were depleted while plasmacytoid DCs (pDCs) were augmented in the lungs of mild COVID-19. In severe cases, we identified augmented types of immature DCs (CD22+ or ANXA1+ DCs) with MHCII downregulation. In this study, our observation indicates that DCs in severe cases stimulate innate immune responses but fail to specifically present SARS-CoV-2. It provides insights into the profound modulation of DC function in severe COVID-19.
Motivation: Recent advancements in single-cell RNA sequencing (scRNA-seq) have enabled time-efficient transcriptome profiling in individual cells. To optimize sequencing protocols and develop reliable analysis methods for various application scenarios, solid simulation methods for scRNA-seq data are required. However, due to the noisy nature of scRNA-seq data, currently available simulation methods cannot sufficiently capture and simulate important properties of real data, especially the biological variation. In this study, we developed scRNA-seq information producer (SCRIP), a novel simulator for scRNA-seq that is accurate and enables simulation of bursting kinetics. Results: Compared to existing simulators, SCRIP showed a significantly higher accuracy of stimulating key data features, including mean-variance dependency in all experiments. SCRIP also outperformed other methods in recovering cell-cell distances. The application of SCRIP in evaluating differential expression analysis methods showed that edgeR outperformed other examined methods in differential expression analyses, and ZINB-WaVE improved the AUC at high dropout rates. Collectively, this study provides the research community with a rigorous tool for scRNA-seq data simulation.
Copy number variation has been identified as a major source of genomic variation associated with disease susceptibility. With the advent of whole-exome sequencing (WES) technology, massive WES data have been generated, allowing for the identification of copy number variants (CNVs) in the protein-coding regions with direct functional interpretation. We have previously shown evidence of the genomic correlation structure in array data and developed a novel chromosomal breakpoint detection algorithm, LDcnv, which showed significantly improved detection power through integrating the correlation structure in a systematic modeling manner. However, it remains unexplored whether the genomic correlation exists in WES data and how such correlation structure integration can improve the CNV detection accuracy. In this study, we first explored the correlation structure of the WES data using the 1000 Genomes Project data. Both real raw read depth and median-normalized data showed strong evidence of the correlation structure. Motivated by this fact, we proposed a correlation-based method, CORRseq, as a novel release of the LDcnv algorithm in profiling WES data. The performance of CORRseq was evaluated in extensive simulation studies and real data analysis from the 1000 Genomes Project. CORRseq outperformed the existing methods in detecting medium and large CNVs. In conclusion, it would be more advantageous to model genomic correlation structure in detecting relatively long CNVs. This study provides great insights for methodology development of CNV detection with NGS data.
MOTIVATION:Copy number variation plays important roles in human complex diseases. The detection of copy number variants (CNVs) is identifying mean shift in genetic intensities to locate chromosomal breakpoints, the step of which is referred to as chromosomal segmentation. Many segmentation algorithms have been developed with a strong assumption of independent observations in the genetic loci, and they assume each locus has an equal chance to be a breakpoint (i.e. boundary of CNVs). However, this assumption is violated in the genetics perspective due to the existence of correlation among genomic positions, such as linkage disequilibrium (LD). Our study showed that the LD structure is related to the location distribution of CNVs, which indeed presents a non-random pattern on the genome. To generate more accurate CNVs, we proposed a novel algorithm, LDcnv, that models the CNV data with its biological characteristics relating to genetic dependence structure (i.e. LD).RESULTS:We theoretically demonstrated the correlation structure of CNV data in SNP array, which further supports the necessity of integrating biological structure in statistical methods for CNV detection. Therefore, we developed the LDcnv that integrated the genomic correlation structure with a local search strategy into statistical modeling of the CNV intensities. To evaluate the performance of LDcnv, we conducted extensive simulations and analyzed large-scale HapMap datasets. We showed that LDcnv presented high accuracy, stability and robustness in CNV detection and higher precision in detecting short CNVs compared to existing methods. This new segmentation algorithm has a wide scope of potential application with data from various high-throughput technology platforms.AVAILABILITY AND IMPLEMENTATION:https://github.com/FeifeiXiaoUSC/LDcnv.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Abstract Background The article aims to compare the efficiency of minimax, optimal and admissible criteria in Simon’s and Fleming’s two-stage design. Methods Three parameter settings (p 1-p 0 = 0.25–0.05, 0.30–0.10, 0.50–0.30) are designed to compare the maximum sample size, the critical values and the expected sample size for minimax, optimal and admissible designs. Type I & II error constraints (α, β) vary across (0.10, 0.10), (0.05, 0.20) and (0.05, 0.10), respectively. Results In both Simon’s and Fleming’s two-stage designs, the maximum sample size of admissible design is smaller than optimal design but larger than minimax design. Meanwhile, the expected samples size of admissible design is smaller than minimax design but larger than optimal design. Mostly, the maximum sample size and expected sample size in Fleming’s designs are considerably smaller than that of Simon’s designs. Conclusions Whenever (p 0, p 1) is pre-specified, it is better to explore in the range of probability q, based on relative importance between maximum sample size and expected sample size, and determine which design to choose. When q is unknown, optimal design may be more favorable for drugs with limited efficacy. Contrarily, minimax design is recommended if treatment demonstrates impressive efficacy.
We sought to investigate safety of axitinib or sorafenib in renal cell carcinoma (RCC) patients and compare toxicity of these two vascular endothelial growth factor receptor inhibitors. Databases of PubMed and Embase were searched. We included phase II and III prospective trials, as well as retrospective studies, in which patients diagnosed with RCC were treated with axitinib or sorafenib monotherapy at a starting dose of 5 mg and 400 mg twice daily, respectively. The overall incidence of high grade hypertension, fatigue, gastrointestinal toxicity and hand-foot syndrome, along with their 95% confidence intervals (CI), were calculated using fixed- or random- effects model according to heterogeneity test results. A total of 26 trials, including 4790 patients, were included in our meta-analysis. Among them, 6 arms were related to axitinib and 22 were associated with sorafenib. The incidences of hypertension (24.9% vs. 7.9%), fatigue (8.2% vs. 6.6%), and gastrointestinal toxicity (17.6% vs. 11.3%) were higher in patients receiving axitinib versus those receiving sorafenib, while the incidence of hand-foot syndrome was lower in patients receiving axitinib versus those receiving sorafenib (9.5% vs. 13.3%). In conclusion, axitinib showed noticeably higher risks of toxicity versus sorafenib. Close monitoring and effective measures for adverse events are recommended during therapy.