Pathology foundation models (PFMs) have demonstrated strong potential across clinical and scientific applications, yet their performance is often hindered by batch effects, which are non-biological variations across tissue source institutions (TSIs) that distort learned feature representations and impair generalization. Conventional mitigation strategies, such as stain normalization, offer limited success in addressing these high-dimensional, complex artifacts. We present GLMP (General-purpose LLM-Mediated Pathology model), a novel framework that generates robust numerical embeddings from histology image patches through an intermediate textual representation. By leveraging pretrained general-purpose multimodal large language models (MLLMs) and text encoders, GLMP effectively prioritizes biologically meaningful signals over TSI-specific artifacts, thereby improving cross-institutional generalization. To our knowledge, GLMP is the first pathology model to use text descriptions of histological features as an intermediate representation for generating numerical embeddings from histology images. Our results highlight the untapped potential of broad-domain, non-specialized MLLMs in computational pathology and introduce a new paradigm for building versatile, generalizable, and robust pathology models.
The DNA repair protein RAD18 activates 'Y-family' Trans-Lesion Synthesis (TLS) DNA polymerases that are DNA damage-tolerant and potentially error-prone. RAD18 is also frequently overexpressed and pathologically activated in cancer cells. However, the extent to which RAD18 shapes cancer genomes and impacts tumorigenesis is unclear. Therefore, we tested the effect of Rad18 status on chemically-induced and oncogene-driven tumorigenesis. In a chemically-induced oral carcinogenesis model, acute (2-16 days) 4NQO-treatment induces expression of Rad18 and TLS polymerase mRNAs in mouse oral epithelial cells prior to emergence of oral squamous cell carcinomas (OSCCs). Chronic (8 week) 4NQO-treatment leads to onset of oral tumors that is accelerated in Rad18 -/- mice when compared with Rad18 +/+ animals. Analysis of OSCC exomes reveals increased levels of G(C)>T(A) transversions in Rad18 -/- tumors when compared with Rad18 +/+. Therefore, Rad18 promotes error-free bypass of 4NQO-induced DNA lesions and suppresses 4NQO-induced oral carcinogenesis. In a Kras G12D -induced lung carcinogenesis model, Rad18-deficiency did not affect rates or incidence of oncogene-induced lung tumors or mutations. Taken together, we demonstrate that Rad18 has context-specific tumor-suppressive activity. Given the prevalence of 4NQO-like environmental exposures, RAD18 is highly likely to shape human cancer genomes and perhaps influence other aspects of the tumorigenic process.
BRAF-mutant (BRAF-MT) colorectal cancer (CRC) represents a clinically aggressive subtype characterized by distinct biological features and significantly worse prognosis compared to BRAF wild-type (BRAF-WT) CRC, with median survival reduced by approximately 40
Oral inflammatory diseases affect nearly half of all humans, yet mechanisms underlying rapidly-destructive inflammation remain poorly understood. We compared peri-implantitis with moderate- and high-grade periodontitis using integrated microbial and single-cell sequencing (>967,169-cells; single-cell RNA-seq, spatial proteotranscriptomics). Laser capture microdissection with compartmental microbiome analysis revealed reduced bacterial load and diversity in peri-implantitis. Expansion of the Human Periodontal Atlas with peri-implantitis single-cell RNA-seq data (36-samples; 121,395 cells) identified CD34+ vascular endothelial cell (VEC) rarefaction and oxidative stress, hypoxia, and NAD⁺ metabolism-associated transcriptional programs enriched in a TNFRSF6B⁺/ICAM1⁺ post-capillary venule (PC-VEC) subpopulation. NAD⁺-consuming ectoenzyme CD38 was selectively enriched and orthogonally confirmed by spatial transcriptomics (6-samples; 283,377-cells) and proteomics (23-samples; 562,397-cells). Spatial neighborhood analyses demonstrated CD38⁺-high PC-VEC expansion, closer proximity, and higher IL16-CD4 T cell signaling in peri-implantitis. Matched high-grade periodontitis biopsies confirmed spatially restricted CD38⁺-VECs despite similar microbial burden, identifying endothelial vasculopathy underlying rapidly advancing oral inflammation and a potential therapeutic axis.
INTRODUCTION:High-risk chronic atrophic gastritis (CAG; OLGA/OLGIM Ⅲ-Ⅳ) carries significant gastric cancer (GC) risk yet lacks reliable gastric stem cell (GSC)-based biomarkers. We evaluated GSC markers LGR5 (proliferative) and TFF2 (protective) for risk stratification. METHODS:TCGA/GEO bioinformatics analysis preceded immunohistochemical validation in 60 clinical samples. Protein co-expression (Wnt/β-catenin, Ki67, Bax) was assessed. Diagnostic/prognostic power was tested via ROC and Kaplan-Meier analyses. Functional networks were deciphered through GO/KEGG enrichment. RESULTS:High-risk CAG and GC tissues showed LGR5 upregulation and TFF2 downregulation (p < 0.001). IHC confirmed these patterns, with concurrent Wnt activation (β-catenin↑, cyclin D1↑) and proliferation-apoptosis imbalance (Ki67↑, Bax↓). TFF2 outperformed LGR5 in diagnosing high-risk CAG (AUC: 0.842 vs. 0.681). Poor GC prognosis correlated with high LGR5/low TFF2 (p < 0.05). Co-expression networks linked LGR5 to metabolic genes (CPS1, ADH6) and TFF2 to mucosal defense (GKN1, PGC). CONCLUSION:The coordinated assessment of LGR5 and TFF2 offers a promising approach to identifying high-risk CAG. This biomarker pair captures a homeostatic imbalance in GSCs linked to Wnt/β-catenin signaling, establishing a novel molecular framework for early detection and future targeted strategies.
Ovarian cancer is an aggressive disease characterized by intraperitoneal dissemination and a distinctive microenvironment. By generating metastatic cohorts encompassing approximately 60 pairs of whole-genome and RNA sequencing, 100 single-cell samples, and 2.5 million spatial transcriptomics (ST) spots, we delineate site-specific tumor-host colocalization patterns. Utilizing our STARLETS framework, we elucidate a Darwinian evolutionary trajectory in which hypoxia and immune pressures select for clones that eventually metastasize. High-resolution ST and ultimate dimensional imaging of solvent-cleared organs (uDISCO) imaging further identify a tripartite ensemble comprising MMP11+ myCAFs, epithelial cells, and SPP1+ macrophages in ascites and metastases, which can be modulated via SPP1-CD44 inhibition. SPP1+ macrophages predict therapeutic responses in clinical trials, including oncolytic virus and poly(ADP-ribose) polymerase inhibitor treatments. Collectively, our study advances insights into spatial dynamics that hold promise for therapeutic approaches in ovarian cancer.
Background/Objectives: Head and neck cancer (HNC) represents the seventh most common cancer diagnosis globally, yet current treatments, including surgery, radiation, and immunotherapy, have shown limited improvement in outcomes. Drug repurposing offers a cost-effective strategy to identify new therapeutic options by leveraging existing medications with known safety profiles. Within this study, we developed the GARD pipeline (Genomic Alteration-based Repurposing for Drugs), designed to uncover repurposing candidates for HNC using genomic and network-based approaches. Methods: GARD integrates multi-omics data from The Cancer Genome Atlas (TCGA), including copy number variation (CNV) and somatic mutations (SOM). The cohort was stratified by human papillomavirus (HPV) status. Risk-associated genes were identified and then expanded via high-confidence protein-protein interaction (PPI) networks. Top candidate genes were filtered through comprehensive analysis of publicly available literature data in PubMed using LLMs to validate the relationship between the identified genes and HNC. The top risk genes and their network-expanded neighbors were mapped against DrugBank, and through statistical significance testing and literature validation, established significant drug-gene associations. Results: Significant genes associated with HNC, inferred by genomics alteration, were identified across HPV-positive and HPV-negative subgroups, such as PIK3CA, SOX2, TP53, EIF4G1, TLR7, CLDN1, PRKCI, and EPHA2. Further expansion through the PPI network identified other targetable genes such as EGFR, ERBB2, and the FGFRs. Literature-based validation efforts ensured confidence in the gene-disease association. Drug-gene mapping revealed candidates spanning those already in clinical trials for HNC (e.g., Afatinib, Cabozantinib, Dasatinib, Brigatinib, Lenvatinib, Capivasertib, and Erdafitinib) and emerging or repurposing candidates (Amuvatinib, XL765 (Voxtalisib), Golotimod, Artenimol, Quercetin, and Acetylsalicylic Acid), offering opportunities for precision repurposing. Conclusions: The GARD pipeline demonstrates a genomics-driven, network-informed framework for systematic drug repurposing in HNC. HPV stratification enhances precision, literature-based validation strengthens confidence, and integrated drug mapping enables refinement of existing therapies and discovery of novel candidates for personalized treatment strategies. Code Availability: The full implementation of the GARD pipeline, including preprocessing scripts, statistical analysis modules, and visualization tools, is publicly available on GitHub.
The human microbiome is a foundational and dynamic foundation for several health-related functions and disease processes. Advances in microbiome sequencing have enabled the characterization of microbial communities in several niches. Longitudinal microbiome studies further strive to discover clinically informative microbial community trajectories. However, these data are fraught with dropout events, high noise, and irregular sampling that limit and prevent the use of many available longitudinal analysis tools. To address these challenges, we introduce Bidirectional GRU-ODE-Bayes (BGOB), a deep learning framework developed for longitudinal microbiome interpolation. BGOB combines bidirectional information flow and ODE-based continuous modeling to jointly interpolate and smoothen trends across individual participants, providing uniform, denoised time intervals across patients. BGOB enables vastly improved performance in differential abundance testing and time-to-event analysis, and makes possible longitudinal analyses requiring uniformity, such as lead-lag detection and temporal clustering. After interpolation, previously low-powered datasets are able to broadly recapitulate known microbiology and elucidate interacting microbial communities. We highlight several associations between microbial taxa and disease, including novel species associated with Early Childhood Caries and disruption of key healthy gut microbiota in Inflammatory Bowel Disease. The BGOB package is publicly available at https://github.com/Rachel-Lyu/BGOB\_n\_test ### Competing Interest Statement The authors have declared no competing interest.
Oral inflammatory diseases affect nearly half of the global population. Among them, newly defined peri-implantitis and high-grade periodontitis represent rapidly advancing inflammatory disease types, marked by relatively rapid tissue destruction. Despite their prevalence, the cell mechanisms and spatial architecture driving this severity remain poorly understood. Focusing first on peri-implantitis versus low- and moderate-grade periodontitis, we applied microbial profiling, single-cell RNA sequencing (scRNA-seq), and spatial proteomics (sp-proteomics) to uncover shared pathogenic programs linked to accelerated niche breakdown. Furthermore, to preserve spatial fidelity, each tissue was anatomically orientated along the tooth- or implant-epithelial interface, analogous sites of disease origination. Laser capture microdissection followed by microbiome analysis of unique tissue compartments revealed reduced bacterial load and diversity in peri-implantitis stroma. We then expanded our version-1 Human Periodontal Atlas by integrating newly generated peri-implantitis scRNAseq data (36-total samples; 121395-cells), revealing widespread transcriptional alterations, including oxidative stress, hypoxic, and NAD+ metabolism-associated signatures, primarily in a subpopulation of TNFRSF6B +/ICAM1 + post-capillary venules. We then performed high-resolution sp-proteomics (15-total samples; 337260-cells) and analyzed VEC states and associated neighborhoods via AstroSuite using newly developed tri-wise spatial analysis. This revealed CD34+-VEC loss and CD38+-VEC expansion almost exclusively in peri-implantitis. We extended this analysis to high-grade periodontitis. Mucosal biopsies from four lesion-affected and four unaffected sites within the same individuals (1:1 matched; 8-samples; 225137-cells) again demonstrated spatially restricted CD38+-VEC remodeling exclusively in affected tissues, with similar vasculopathy front patterning. The findings nominate spatially distinct vasculopathy patterning as a hallmark of rapidly advancing oral inflammation and a targetable therapeutic axis.
MOTIVATION:Omics features, often measured by high-throughput technologies, combined with clinical features, significantly impact the understanding of many complex human diseases. Integrating key omics biomarkers with clinical risk factors is essential for elucidating disease mechanisms, advancing early diagnosis, and enhancing precision medicine. However, the high dimensionality and intricate associations between disease outcomes and omics profiles present substantial analytical challenges. RESULTS:We propose a high-dimensional feature importance test (HiFIT) framework to address these challenges. Specifically, we develop an ensemble data-driven biomarker identification tool, Hybrid Feature Screening (HFS), to construct a candidate feature set for downstream machine learning models. The pre-screened candidate features from HFS are further refined using a computationally efficient permutation-based feature importance test employing machine learning methods to flexibly model the potential complex associations between disease outcomes and molecular biomarkers. Through extensive numerical simulation studies and practical applications to microbiome-associated weight changes following bariatric surgery, as well as the examination of gene-expression-associated kidney pan-cancer survival data, we demonstrate HiFIT's superior performance in both outcome prediction and feature importance identification. AVAILABILITY AND IMPLEMENTATION:An R package implementing the HiFIT algorithm is available on GitHub (https://github.com/BZou-lab/HiFIT).
AIM:The COVID-19 pandemic affected practice in endodontic offices. Same-day endodontic emergencies are cases with moderate or severe self-reported pain who request an unscheduled visit on the day they contact the office. The aims of this observational study were to: (A) analyse the rate of same-day endodontic emergencies in two endodontists' private offices, with respect to their demographic, aetiologic, diagnostic and procedural data; and (B) investigate the changes in characteristics of same-day emergencies between March 16 and May 31 annually over five years: 2019-2023. METHODOLOGY:Records of 5795 patients were reviewed and 892 same-day emergencies were identified. Overall and year-to-year comparisons of proportions of same-day emergencies, as well as demographic, aetiologic, diagnostic and procedural data were performed using chi-square test of independence followed by adjustments for multiple testing using the Benjamini-Hochberg method. RESULTS:The rate of same-day endodontic emergencies significantly increased during the initial outbreak of COVID-19 in 2020 and remained high in 2021 (p < .05; Q < .05). The rate of same-day emergencies in 2022 subsided to levels comparable to 2019 (p > .05). Year-to-year comparisons of aetiologic factors (caries, restorative, persistent infection and cracks) showed a significant increase only in the rate of cracks in 2020, 2021and 2022 compared with 2019 (p < .05), but this increase did not reach the significance level after adjusting for multiple comparisons throughout the 5 years (Q > .05). CONCLUSIONS:The COVID-19 pandemic was associated with a significant increase in the rate of same-day endodontic emergencies for 2 years. The spike in endodontic emergencies associated with the COVID-19 pandemic lasted well beyond the initial period of the outbreak. Further national and international studies are recommended to better understand the long-term impacts of pandemics of respiratory diseases on the public's oral health.
The DNA damage response (DDR) mechanisms that allow cells to tolerate DNA replication stress are critically important for genome stability and cell viability. Using an unbiased genetic screen we identify a role for the RING finger E3 ubiquitin ligase RNF25 in promoting DNA replication stress tolerance. In response to DNA replication stress, RNF25-deficient cells generate aberrantly high levels of single-stranded DNA (ssDNA), accumulate in S-phase and show reduced mitotic entry. Using single-molecule DNA fiber analysis, we show that RNF25 protects reversed DNA replication forks generated by the fork remodeler HLTF from nucleolytic degradation by MRE11 and CtIP. Mechanistically, RNF25 interacts with the replication fork protection factor REV7 and recruits REV7 to nascent DNA after replication stress. The role of RNF25 in protecting replication forks is fully separable from its canonical functions in ubiquitin conjugation. This work reveals the RNF25-REV7 signaling axis as an important protective mechanism in cells experiencing replication stress.
Semi-continuous data frequently arise in clinical practice. For example, while many surgical patients suffer from varying degrees of acute postoperative pain (POP) post surgery (i.e., POP score > 0), others experience none (i.e., POP score = 0), indicating the existence of two distinct data processes at play. Existing parametric or semi-parametric two-part modeling methods for this type of semicontinuous data can fail to appropriately model these two underlying data processes as such methods rely heavily on (generalized) linear additive assumptions. However, many factors may interact to jointly influence the experience of POP non-additively and non-linearly. Motivated by this challenge and inspired by the flexibility of deep neural networks (DNN) to accurately approximate complex functions universally, we derive a DNN-based two-part model by adapting the conventional DNN methods by adding two additional components: a bootstrapping procedure along with a filtering algorithm to boost the stability of the conventional DNN, an approach we denote as sDNN. To improve the interpretability and transparency of sDNN, we further derive a feature importance testing procedure to identify important features contributing to the outcome measurements of the two data processes, denoting this approach fsDNN. We show that fsDNN not only offers a valid feature importance test but also that using the identified features can further improve the predictive performance of sDNN. The proposed sDNN- and fsDNN-based twopart models are applied to the analysis of real data from a POP study, in which application they clearly demonstrate advantages over the existing parametric and semi-parametric two-part models. Further, we conduct extensive numerical studies to demonstrate that sDNN and fsDNN consistently outperform the existing two-part models regardless of the data complexity. An R package implementing the proposed methods has been developed and deposited on GitHub ( https://github.com/SkadiEye/fsDNN ).
BACKGROUND:Between 6 and 22% of children are affected by dental anxiety. Dental anxiety is a significant barrier to dental care and is associated with dental avoidance and negative oral health outcomes. Pharmacological methods of anxiety management are costly, carry risks of adverse outcomes, and may not be acceptable to some families. Alternative non-pharmacological methods are needed for the safe and effective delivery of dental care. Although there is an abundance of literature regarding animal-assisted therapy (AAT) in medicine, only preliminary studies on AAT exist in dentistry. To identify optimal outcome measures for evaluating AAT in pediatric dental contexts, a randomized controlled trial protocol was developed. METHODS:A prospective randomized controlled trial protocol was developed to examine the impact of AAT on objective (heart rate, salivary stress and pain markers, and observational coding) and subjective self-reported measures of anxiety, pain, and dental expectations in pediatric patients. The study is designed to enroll 180 pediatric patients (4-8 years old), randomized into three arms (n = 60 per arm) with stratification by age (< 6.5 vs ≥ 6.5) and gender (block size = 4). Two therapy protocols (+ Short AAT and + Long AAT exposures) will be compared relative to an active control (coloring a dog picture) during a diagnostic dental visit consisting of an oral exam, dental cleaning, and simulated bitewing intraoral radiographs. DISCUSSION:This study will provide information on optimal outcome measures to evaluate the impact of AAT on dental anxiety and behavior in pediatric dental patients. Determining the effects of AAT in pediatric dental care may provide a safe, non-pharmacological method of anxiety and behavior management, with broad translational impact. TRIAL REGISTRATION:This trial was registered on ClinicalTrials.gov with number NCT05464888, on 15 July 2022 (first submitted to ClinicalTrials.gov) and 19 July 2022 (first posted to ClinicalTrials.gov).
Background: The Bacille Calmette–Guérin (BCG) vaccine is part of the Extended Programme on Immunization (EPI) and as such is generally administered at birth. The global introduction of BCG not only protected many vaccinated infants against severe complications of tuberculosis but also resulted in markedly reduced overall childhood mortality. Studies in human adults determined that BCG vaccination induces epigenetic reprogramming of innate immune cells (also known as trained immunity) and can also enhance T cell responses to both mycobacterial and non-mycobacterial antigens. Goal and Methods: The current study tested the hypothesis that BCG immunization similarly impacts the functionally distinct infant immune system. Towards this goal, we applied RNA sequencing to assess transcriptome changes in circulating CD4+ T cells of Malawian infants prior to and 2 to 13 weeks after BCG immunization. Results: In the first three months of life, transcriptome changes of infant CD4 T cells implied a functional shift towards T helper 1 and Th17 immunity. Vaccination with BCG resulted in additional modulation of the CD4 T cell transcriptome and differentially expressed genes could be linked to metabolomic function. Conclusions: These findings are consistent with data reported in BCG vaccinated adults and contribute to the understanding of molecular changes in infant CD4 T cells that may explain the improved capacity of the infant immune system to respond to pathogens after BCG vaccination.
The Cancer Testis Antigens (CTAs) are a group of germ cell proteins that are absent from normal somatic cells yet aberrantly expressed in many cancer cells. When mis-expressed in cancer cells, many CTAs promote tumorigenic characteristics including genome instability, DNA damage tolerance and therapy resistance. Here we highlight some of the CTAs for which their roles in genome maintenance in cancer cells are well established. We consider three broad CTA categories: (1) Melanoma Antigens (MAGEs) (2) Mitotic CTAs and (3) CTAs with roles in meiotic homologous recombination. Many cancer cells rely on CTAs to tolerate intrinsic and therapy-induced genotoxic stress. Therefore, CTAs represent molecular vulnerabilities of cancer cells and may provide opportunities for therapy. Owing to their high-level expression in tumors and absence from normal somatic cells, CTA-directed therapies could have a high level of specificity and would likely be devoid of side-effect toxicity.
Information generated from longitudinally sampled microbial data has the potential to illuminate important aspects of development and progression for many human conditions and diseases. Identifying microbial biomarkers and their time-varying effects can not only advance our understanding of pathogenetic mechanisms, but also facilitate early diagnosis and guide optimal timing of interventions. However, longitudinal predictive modeling of highly noisy and dynamic microbial data (e.g. metagenomics) poses analytical challenges.To overcome these challenges, we introduce a robust and interpretable machine-learning-based longitudinal microbiome analysis framework, LP-Micro, that encompasses (i) longitudinal microbial feature screening via a polynomial group lasso, (ii) disease outcome prediction implemented via machine learning methods (e.g. XGBoost, deep neural networks), and (iii) interpretable association testing between time points, microbial features, and disease outcomes via permutation feature importance. We demonstrate in simulations that LP-Micro can not only identify incident disease-related microbiome taxa, but also offers improved prediction accuracy compared with existing approaches. Applications of LP-Micro in two longitudinal microbiome studies with clinical outcomes of childhood dental disease and weight loss following bariatric surgery yield consistently high prediction accuracy. Moreover, LP-Micro highlights critical time points and associated microbial changes: oral microbial changes, including Streptococcus mutans, are most informative for predicting childhood dental disease at around 39 months of age, while gut microbial changes shortly after bariatric surgery strongly predict future weight loss. These findings are both informative and aligned with clinical expectations. The tool LP-Micro can be seen at https://github.com/IV012/LPMicro.
The regeneration of periodontal, periapical, and pulpal tissues is a complex process requiring the direct involvement of cells derived from pluripotent stem cells in the periodontal ligament and dental pulp. Dental pulp stem cells (DPSCs) and periodontal ligament stem cells (PDLSCs) are spatially distinct with the potential to differentiate into similar functional and phenotypic cells. We aimed to identify the cell heterogeneity of DPSCs and PDLSCs and explore the differentiation potentials of their specialized organ-specific functions using single-cell transcriptomic analysis. Our results revealed 7 distinct clusters, with cluster 3 showing the highest potential for differentiation. Clusters 0 to 2 displayed features similar to fibroblasts. The trajectory route of the cell state transition from cluster 3 to clusters 0, 1, and 2 indicated the distinct nature of cell differentiation. PDLSCs had a higher proportion of cells (78.6%) at the G1 phase, while DPSCs had a higher proportion of cells at the S and G2/M phases (36.1%), mirroring the lower cell proliferation capacity of PDLSCs than DPSCs. Our study suggested the heterogeneity of stemness across PDLSCs and DPSCs, the similarities of these 2 stem cell compartments to be potentially integrated for regenerative strategies, and the distinct features between them potentially particularized for organ-specific functions of the dental pulp and periodontal ligament for a targeted regenerative dental tissue repair and other regeneration therapies.