Abstract Translating genome-wide association studies (GWAS) signals into trait-relevant cellular contexts remains challenging due to the complexity of the genomic regulatory code and linkage disequilibrium among associated variants. We present a novel computational framework that aggregates deep learning–based predictions of the functional effects of noncoding variants on transcriptional regulatory elements across GWAS loci and empirically evaluates their statistical significance. By organizing these aggregated signals within biological ontologies, our approach enables statistically calibrated interpretation of GWAS associations, highlighting relevant cell-type and tissue contexts across human traits.
Abstract Background Plasmids play a critical role in horizontal gene transfer and the spread of antibiotic resistance. However, recovering complete plasmid sequences from metagenomic samples remains highly challenging due to extensive repeat content, structural heterogeneity, and large variation in plasmid size. Existing methods typically identify plasmids from metagenome assemblies by exploiting coverage differences or detecting minimum-weight cycles in assembly graphs. While effective for dominant plasmids, these approaches often fail to recover low-abundance and long plasmids. Results Here we present PlasChain, a novel algorithm designed to improve plasmid assembly and identification from complex metagenomic data. Building upon the cycle-peeling strategy of SCAPP, PlasChain incorporates contig path information and a cycle-merging procedure to prevent long plasmids from being fragmented into multiple shorter cycles. In addition, PlasChain jointly leverages paired-end read alignments, sequence composition patterns, and coverage variation to filter out assembly artifacts and reduce false positives. We evaluated PlasChain against state-of-the-art plasmid assemblers, including SCAPP and metaplasmidSPAdes, using a diverse set of simulated and real metagenomic datasets. Across nearly all benchmarks, PlasChain demonstrates superior performance in recovering long plasmids while maintaining competitive accuracy in assembling short plasmids. Furthermore, analysis of real metagenomic samples shows that PlasChain is capable of assembling previously uncharacterized plasmids, including putative megaplasmids that are typically underrepresented in current plasmid databases. Conclusions PlasChain is a novel graph-based plasmid assembler that improves the recovery of long plasmids from short-read metagenomic data. Evaluation on diverse simulated and real metagenomic datasets demonstrates that PlasChain consistently outperforms existing plasmid assemblers, especially for long plasmid reconstruction. These results highlight the potential of PlasChain to facilitate comprehensive characterization of plasmid diversity, antimicrobial resistance, and horizontal gene transfer in complex microbial communities. The source code and testing data of PlasChain are freely available at https://github.com/SDU-ACG-Lab/PlasChain . Supplementary Information This manuscript is accompanied by a supplementary file.
Polygenic risk scores (PRSs), which quantify inherited susceptibility to complex traits and diseases, have emerged as valuable tools for risk stratification and precision medicine. Despite their promise, PRS developed on European cohorts often demonstrate substantially reduced predictive accuracy in non-European populations, due to differences in genetic architecture. The disproportionate representation of European ancestry cohorts in genome-wide association studies (GWAS) leads to inequitable deployment of PRS technologies across diverse populations. Here, we introduce PRANA (Polygenic Risk Adaptation via Neural-network Architecture), a deep learning framework that adapts an existing PRS developed on one population to other ancestries. Unlike methods that require large-scale GWAS in the target population, PRANA leverages pre-trained PRS models derived from European cohorts and adapts them using modestly sized cohorts from the target population. We evaluated PRANA on seven complex traits in South Asian, East Asian and Ashkenazi Jewish populations, as well as in selected smaller East Asian subpopulations where the scarcity of training data poses a particular challenge. PRANA mostly improved predictive performance of the baseline PRS models by 5%-20% in terms of effect size (β) and Nagelkerke's R2, and, in most cases, outperformed existing cross-ancestry multi-PRS approaches. These results highlight PRANA as a scalable and practical strategy to reduce disparities in genomic risk prediction and advance the equitable application of PRS in diverse populations.
Parkinson’s disease (PD) is a highly heterogeneous condition with symptoms spanning motor and non-motor domains. Clinical scales like the Movement Disorder Society’s Unified Parkinson’s Disease Rating Scale (MDS-UPDRS) are standard in clinical trials where disease progression is monitored. They rely on summing item values, assuming uniform item importance and score increments. Here, we propose a novel data-driven approach to optimize weights for such scales–so that total scores better reflect the underlying disease severity. In a retrospective observational analysis of longitudinal cohort data from the Parkinson’s Progression Markers Initiative (PPMI), our methods identified which items (and value increments) most strongly indicate PD progression, down-weighting or excluding less informative items. The learned weights substantially improve the monotonic relationship between total scores and clinical progression. We validated our weights using both held-out PPMI data and an independent dataset (BeaT-PD), demonstrating their robustness. Applying such weights in clinical trials may increase power and reduce the required sample size1.
Recent studies have shown that the tumor mycobiome may have prognostic and diagnostic significance in cancer patients. We aimed to gain a better understanding of how patient characteristics (age, sex, body mass index [BMI], and race) influence the composition of the tumor mycobiome, using the data of these studies. We first tested the data in view of recent critiques of tumor microbiome data processing procedures and concluded that the batch correction and transformation used on it may produce false signals. Instead, we explored 14 combinations of data transformation and batch correction methods on data of 224 fungal species across 13 cancer types. Propensity scores were utilized to adjust for potential confounders such as histological type and tumor stage. To minimize false outcomes, we identified as positive results only those fungi species that showed significant difference in abundance across a demographic factor within a particular cancer type, using data normalized according to all 14 combinations. We observed significant differences in 24 fungal species abundance within tumors for certain demographic characteristics. A total of 20 of these differences were among races in specific cancers. The findings indicate that there are intricate interactions between the mycobiome, cancer type, and patient demographics. Our study highlights the need to account for race in order to understand the role of the mycobiome in cancer development and treatment response. The study also underscores the importance of data processing techniques.IMPORTANCEThis study analyzes the demographic-dependent variability of the intratumor mycobiome, providing a novel understanding of fungal abundance across different cancer types and patient demographics. By analyzing over 5,000 tumor samples from The Cancer Genome Atlas, the research identified 24 fungal species with significant abundance variations linked to demographic factors such as race, age, sex, and body mass index. These findings underscore the complexity of the tumor microenvironment and the importance of accounting for demographic diversity in cancer research. The study emphasizes the necessity of using robust data normalization and batch correction techniques to avoid spurious associations in order to ensure the reliability of mycobiome analysis. This work highlights the mycobiome as a new frontier in precision oncology and paves the way for future personalized cancer diagnostics and treatments that account for the influence of demographic factors on tumor biology.
Abstract Background Insomnia is one of the most common sleep disorders, affecting up to 30% of the adult population worldwide. Sleep disorders can have a marked impact on immune functions and inflammation, including in patients with inflammatory bowel diseases (IBD) in both active and inactive disease states. Whether insomnia preceding IBD flares affects clinical outcomes is unknown. We examined the association between insomnia and IBD-related outcomes in adults diagnosed with of Crohn’s disease (CD) or ulcerative colitis (UC). Methods A retrospective, real-world, cohort study utilizing the Lynx MD’s electronic medical record dataset, provided by a US-based community gastroenterology practice. Deidentified patient charts were analyzed to identify patients with a confirmed diagnosis of IBD (either CD or UC). Eligibility criteria included: ≥3 diagnoses of IBD, ≥1 diagnoses of IBD and: histopathologic diagnosis of IBD; or IBD related surgery; or ≥ 3 months of treatment with standard accepted therapies for IBD; and ≥3 years of follow-up. IBD-related events of hospitalization and/or surgery were recorded. The cohort was classified into two groups based on insomnia diagnosis, prior to IBD related hospitalizations or surgeries. We implemented Cox Proportional Regression models to estimate Hazard Ratios for the IBD-related outcomes. Results Among 20,011 patients with IBD, 10,775 (54%) were women, at a mean age of 48.2 ± 18.2 years old, 11,842 (59%) were diagnosed with UC and 8,169 (41%) with CD. Over 3,100 patients (16%) were diagnosed with insomnia prior to the first IBD related hospitalization or surgery. In an unadjusted model, insomnia was significantly associated with IBD-related hospitalizations - hazard ratio (HR) of 1.43 (95% CI 1.28, 1.59) and IBD related surgeries - HR of 1.42 (95% CI 1.12, 1.81). After adjustment for age, sex, race, body mass index, anxiety and type of IBD (CD vs. UC), IBD-affected patients with insomnia had a higher HR of hospitalization and surgery: HR of 1.35 (95% CI 1.2,1.51) and 1.29 (95% CI 1.01,1.66), respectively, compared to those without insomnia. Conclusion In a large real-life dataset of patients with IBD, we found higher rates of IBD-related hospitalizations and surgeries among patients with comorbid insomnia, suggesting that insomnia may be associated with a more severe course of IBD.
Plasmids are influential drivers of bacterial evolution, facilitating horizontal gene transfer and shaping microbial communities. Current knowledge on plasmid persistence and mobilization in natural environments is derived from community-level studies, neglecting the single-cell level, where these dynamic processes unfold. Pinpointing specific plasmids within their natural environments is essential to unravel the dynamics between plasmids and their bacterial hosts. Here, we overcame the technical hurdle of natural plasmid detectability in single cells by developing SPEci-FISH (Short Probe EffiCIent Fluorescence In Situ Hybridization), a novel molecular method designed to detect and visualize plasmids, regardless of their copy number, directly within bacterial cells, enabling their precise identification at the single-cell level. To complement this method, we created ProFiT (PRObe FInding Tool), a program facilitating the design of sequence-based probes for targeting individual plasmids or plasmid families. We have successfully applied these methods, combined with high-resolution microscopy, to investigate the dispersal and localization of natural plasmids within a clinical isolate, revealing various plasmid spatial patterns within the same bacterial population. Importantly, bridging the technological gap in linking plasmids to hosts in native complex microbial environments, we demonstrated that our method, when combined with fluorescence-activated cell sorting (FACS), can track plasmid-host dynamics in a human fecal sample. This approach identified multiple potential bacterial hosts for a conjugative plasmid that we assembled from this fecal sample's metagenome. Our integrated approach offers a significant advancement toward understanding plasmid ecology in complex microbiomes.
Aneuploidy is a hallmark of cancer, yet the genes driving recurrent chromosome-arm losses remain largely unknown. We present a systematic framework integrating mutation, copy number, and gene expression data to identify candidate driver genes of cancer type-specific recurrent chromosome-arm losses across 20 cancer types, using ∼7,500 tumors from The Cancer Genome Atlas. By analyzing focal deletions and point mutations that co-occur, or are mutually exclusive, with chromosome-arm losses, we pinpoint 322 candidate drivers associated with 159 recurring events. Our approach identifies known aneuploidy drivers such as TP53 and PTEN, while revealing multiple additional candidates, including tumor suppressors not previously linked to aneuploidy. We leverage expression changes associated with chromosome-arm losses to propose cancer-promoting pathway-level alterations. Integrating these findings highlights key candidate drivers that underlie the observed expression alterations, reinforcing their biological relevance. We provide a comprehensive catalog of candidate driver genes for recurrently lost chromosome-arms in human cancer.
Abstract Background Anti-Saccharomyces cerevisiae antibodies (ASCA) have been associated with a more aggressive Crohn’s disease (CD) phenotype. However, the effect of ASCA serology on the response to biologic therapy, specifically in children on anti-TNF therapy, is unknown. Our goal was to assess whether ASCA levels are associated with clinical outcomes in pediatric CD patients treated with anti-TNF antibodies. Methods A single center retrospective study. Demographic, clinical and laboratory data were collected from pediatric CD patients (aged 2.2-17.9 years) treated with anti-TNF therapy, who were tested for IgG ASCA levels upon diagnosis, between 2010-2023. The cut-off value for a positive ASCA result was defined as 30 EU/mL. Clinical outcomes included durability of anti-TNF therapy, corticosteroid-free survival (CSFS) at week 52 of therapy, IBD-associated hospitalizations and IBD related surgery. Results One hundred seventy-two patients with CD (69 ,39% females) with a median age at diagnosis of 13.5 (11.1-15.4) years were included. Ninety-two (52%) patients had positive ASCA (> 30 EU/mL). ASCA positivity was associated with female sex (p=0.02) and ileocolonic disease location at diagnosis (p=0.02). Sixty-six (37%) patients were treated with infliximab and 111 (63%) with adalimumab. The median time of follow-up in our cohort was 115 (65-174) weeks. Time to discontinuation of anti-TNF was comparable between patients with ASCA positive and negative levels (97.8% vs. 97.4% and 89.3% vs. 87.4% at 1 and 3 years, respectively). In addition, time to IBD-associated hospitalization and surgery were also similar. Interestingly, patients with positive ASCA achieved CSFR more often, compared to patients with ASCA negative (P=0.04). Sub-analysis of patients with high ASCA values, the top 20%, (>80 EU/mL) demonstrated significantly longer durability of anti-TNF, compared to the patients with lower ASCA levels (<80 EU/mL,p= 0.02). Conclusion ASCA positivity is not associated with worse outcomes in pediatric patients with CD treated with TNF antagonists.
To date, most studies explored changes in 3D-genome organization between different tissues or during differentiation, which involve massive reprogramming of transcriptional programs. Much fewer studies examined alterations in genome organization in response to cellular stress, which involves less pervasive transcriptional modulation. Here, we examined associations between spatial chromatin organization and gene expression in two different biological contexts: transcriptional programs determining cell identity and transcriptional responses to stress, using p53 activation as a model. We selected 10 cell lines of diverse tissues, and in each performed micro-C, RNA-seq, and p53 ChIP-seq, before and after p53 induction. In the comparison between cell types, we delineated marked correlations between gene expression and spatial genome organization and identified hundreds of active enhancer-promoter loops associated with the expression of cell-type marker genes. In contrast, within each cell type, no such links were observed for expression changes induced by p53 activation, even for enhancers and promoters activated by p53 binding. Our analysis points to a fundamental difference between chromatin interactions that define cell identity and those that are established in response to cellular stress. Our results on p53-induced transcriptional responses support the recently proposed TF activity gradient model, which speculated a contact-independent mechanism for enhancer-promoter communication.
Randomized Controlled Trials (RCTs) are the gold standard for evaluating the effect of new medical treatments. Treatments must pass stringent regulatory conditions in order to be approved for widespread use, yet even after the regulatory barriers are crossed, real-world challenges might arise: Who should get the treatment? What is its true clinical utility? Are there discrepancies in the treatment effectiveness across diverse and under-served populations? We introduce two new objectives for future clinical trials that integrate regulatory constraints and treatment policy value for both the entire population and under-served populations, thus answering some of the questions above in advance. Designed to meet these objectives, we formulate Randomize First Augment Next (RFAN), a new framework for designing Phase III clinical trials. Our framework consists of a standard randomized component followed by an adaptive one, jointly meant to efficiently and safely acquire and assign patients into treatment arms during the trial. Then, we propose strategies for implementing RFAN based on causal, deep Bayesian active learning. Finally, we empirically evaluate the performance of our framework using synthetic and real-world semi-synthetic datasets.
Antimicrobial resistance is a rising global health threat, leading to ineffective treatments, increased mortality and rising healthcare costs. In ICUs, inappropriate empiric antibiotic therapy is often given due to treatment urgency, causing poor outcomes. This study developed a machine learning model to predict the appropriateness of empiric antibiotics for ICU-acquired bloodstream infections, using data from the MIMIC-III database. To address missing values and dataset imbalances, novel computational methods were introduced. The model achieved an AUROC of 77.3% and AUPRC of 40.4% on validation, with similar results on external datasets from MIMIC-IV and Rambam Hospital. The model also predicted mortality risk, identifying a 30% mortality rate in high-risk patients versus 16.8% in low-risk groups. External validation on the eICU database showed a comparable gap, with mortality rates at 24% for high-risk and 7.7% for low-risk groups. Our study demonstrates the potential of machine learning models to predict inappropriate empiric antibiotic treatment.
As Molecular Systems Biology marks its 20th anniversary, we take this moment to reflect on two decades of discovery and innovation. Since its launch, the journal has stood at the forefront of integrating quantitative biology, computational modeling, and systems science—helping to shape how we understand complex biological systems.
Abstract Background The intestinal mucosa regenerates every few days, and cells are continuously shed into the gut lumen. These cells, primarily epithelial and immune cells, have recently been shown to remain viable after shedding and hold critical information for assessing bowel pathology. Analysis of endoscopic washes of luminal material during colonoscopies suggests that the human transcriptome in fecal material accurately reflects degree of histologic inflammation in patients with inflammatory bowel diseases (IBD). We aimed to define whether transcriptomic analysis of stool samples can reflect degree of intestinal inflammation in patients with IBD. Methods Bulk RNA sequencing was performed on stool samples to obtain fecal human transcriptomes from 82 IBD patients, as well as from healthy controls. Each sample received an inflammation score derived from a group of pro-inflammatory transcripts. Inflammatory score was compared to fecal calprotectin as well as endoscopic evaluation. Results The quality of human RNA isolated from fecal content was high, with a median of 2,218 human genes detected per sample. The correlation between fecal calprotectin protein levels and mRNA expression was high (r=0.75). An inflammatory score was generated based on a group of differentially expressed genes. The inflammatory score provided high sensitivity and specificity (both 90%) in predicting endoscopic inflammatory activity when compared to colonoscopic findings. Conclusion Fecal shed cell transcriptomics provided a novel non-invasive tool to measure inflammatory activity at the endoscopic level in patients with IBD. Ongoing studies will address whether this technology can assist in defining degree and location of inflammation, as well as predict response to different therapies.
Introduction Recently, Narunsky-Haziza et. al. showed that fungi species identified in a variety of cancer types may have prognostic and diagnostic signficane. We used that data in order to better understand the effects of demographic factors (age, sex, BMI, and race) on the intratumor mycobiome composition. Materials and Methods We first tested the data in view of recent critiques of microbiome data processing procedures, and concluded that the batch correction and transformation used on it may produce false signals. Instead, we explored 14 combinations of data transformation and batch correction methods on data of 224 fungal species across 13 cancer types. Propensity scores were utilized to adjust for potential confounders such as histological type and tumor stage. To minimize false outcomes, we identified as positive results only those fungi species that showed significant difference in abundance across a demographic factor within a particular cancer type, using data normalized according to all 14 combinations. Results and Discussion We observed significant differences in fungal species abundance within tumors for certain demographic characteristics. Most differences were among races in specific cancers. The findings indicate that there are intricate interactions among the mycobiome, cancer types, and patient demographics. Our study highlights the need for accounting for potential confounders in order to further understanding of the mycobiome’s role in cancer, and underscores the importance of data processing techniques. ### Competing Interest Statement The authors have declared no competing interest.
Abstract Background Intensification of adalimumab (ADL) dosing to weekly 40 mg injections, in response to low drug levels, has been shown to provide beneficial outcomes in pediatric patients with Crohn’s disease (CD). Our study aimed to evaluate the safety and efficacy of weekly 80 mg ADL administration in children with CD. Methods In this retrospective cohort study conducted across five Israeli centers, we reviewed the medical records of pediatric CD patients who received a high dose of ADL 80 mg weekly injections between 2016 and 2023. Collected data included demographic characteristics, disease features, laboratory studies, and treatment outcomes. Results Thirty-two children with CD were included: mean age 15.8 (±1.7) years at intensification, 21 male (66%), 18 (56%) with L3 phenotype. The median time to ADL 80 mg intensification from ADL induction was 48.4 weeks (IQR 23.1-122.5). The mean weighted Pediatric Crohn's Disease Activity Index (wPCDAI) was 28.5 (±16.9) at the time of intensification and the median calprotectin and C-reactive protein levels were 937 μg/g (IQR 540-1410) and 1.3 mg/dL (IQR 0.6-5.2), respectively. Clinically active disease was the main reason for ADL intensification (30, 94%). Baseline ADL levels were available in 30 patients (94%) with a median of 3.8 μg/mL (IQR 2.4-7.2). Among these, 23 children (77%) failed to achieve a target of ≥ 7.5 μg/mL. The median follow-up duration from the intensification dose was 91.2 weeks (IQR 53.6-149.1). Corticosteroids were required in 5/32 (15.6%) children, with a median time of 32.3 weeks (IQR 10.7-51.3) from intensification. There was no statistically significant difference in steroid utilization rates between children who achieved a target baseline drug level of ≥ 7.5 μg/mL and those who did not (p=0.934). Thirteen individuals (41%) discontinued ADL treatment, within a median of 24.7 weeks (IQR 11.7-57.1) from intensification. CD-related exacerbation, hospitalization, and surgery rates were 9 (28%), 4 (13%), and 2 (6%), respectively. No statistically significant differences were found in exacerbation (p=0.406), hospitalization (p=0.322), or surgery rates (p=0.427) between children who achieved a baseline ADL trough concentration level of ≥ 7.5 μg/mL and those who did not. Overall, ADL intensification was safe, but 3 (9%) patients developed new-onset psoriasis. Conclusion Our findings support the safety and efficacy of administering ADL at a weekly dosage of 80 mg as maintenance therapy in pediatric patients with moderate to severe CD.
MOTIVATION:Polygenic risk scores (PRSs) predict individuals' genetic risk of developing complex diseases. They summarize the effect of many variants discovered in genome-wide association studies (GWASs). However, to date, large GWASs exist primarily for the European population and the quality of PRS prediction declines when applied to other ethnicities. Genetic profiling of individuals in the discovery set (on which the GWAS was performed) and target set (on which the PRS is applied) is typically done by SNP arrays that genotype a fraction of common SNPs. Therefore, a key step in GWAS analysis and PRS calculation is imputing untyped SNPs using a panel of fully sequenced individuals. The imputation results depend on the ethnic composition of the imputation panel. Imputing genotypes with a panel of individuals of the same ethnicity as the genotyped individuals typically improves imputation accuracy. However, there has been no systematic investigation into the influence of the ethnic composition of imputation panels on the accuracy of PRS predictions when applied to ethnic groups that differ from the population used in the GWAS. RESULTS:We estimated the effect of imputation of the target set on prediction accuracy of PRS when the discovery and the target sets come from different ethnic groups. We analyzed binary phenotypes on ethnically distinct sets from the UK Biobank and other resources. We generated ethnically homogenous panels, imputed the target sets, and generated PRSs. Then, we assessed the prediction accuracy obtained from each imputation panel. Our analysis indicates that using an imputation panel matched to the ethnicity of the target population yields only a marginal improvement and only under specific conditions. AVAILABILITY AND IMPLEMENTATION:The source code used for executing the analyses is this paper is available at https://github.com/Shamir-Lab/PRS-imputation-panels.
The laminar microstructure of the cerebral cortex has distinct anatomical characteristics of the development, function, connectivity, and even various pathologies of the brain. In recent years, multiple neuroimaging studies have utilized magnetic resonance imaging (MRI) relaxometry to visualize and explore this intricate microstructure, successfully delineating the cortical laminar components. Despite this progress, T1 is still primarily considered a direct measure of myeloarchitecture (myelin content), rather than a probe of tissue cytoarchitecture (cellular composition). This study aims to offer a robust, whole-brain validation of T1 imaging as a practical and effective tool for exploring the laminar composition of the cortex. To do so, we cluster complex microstructural cortical datasets of both human ( N = 30) and macaque ( N = 1) brains using an adaptation of an algorithm for clustering cell omics profiles. The resulting cluster patterns are then compared to established atlases of cytoarchitectonic features, exhibiting significant correspondence in both species. Lastly, we demonstrate the expanded applicability of T1 imaging by exploring some of the cytoarchitectonic features behind various unique skillsets, such as musicality and athleticism.
Roded Sharan合作论文数Tel-Aviv University;School of Computer Science23
Gideon Dror合作论文数Google;School of Computer Science, The Academic College of Tel Aviv Yaffo7