Despite the intrinsic risk-awareness of Large Language Models (LLMs), current defenses often result in shallow safety alignment, rendering models vulnerable to disguised attacks (e.g., prefilling) while degrading utility. To bridge this gap, we propose SafeThinker, an adaptive framework that dynamically allocates defensive resources via a lightweight gateway classifier. Based on the gateway's risk assessment, inputs are routed through three distinct mechanisms: (i) a Standardized Refusal Mechanism for explicit threats to maximize efficiency; (ii) a Safety-Aware Twin Expert (SATE) module to intercept deceptive attacks masquerading as benign queries; and (iii) a Distribution-Guided Think (DDGT) component that adaptively intervenes during uncertain generation. Experiments show that SafeThinker significantly lowers attack success rates across diverse jailbreak strategies without compromising utility, demonstrating that coordinating intrinsic judgment throughout the generation process effectively balances robustness and practicality.
Comprehensive characterization of cellular states requires simultaneous measurements of transcriptomes and proteomes at single-cell resolution. However, current technologies either measure only limited protein panels or quantify thousands of proteins at extremely low throughput. As a result, obtaining large-scale paired transcriptomic–proteomic measurements at single-cell resolution remains challenging. Here we present msInfer, a computational framework that integrates unpaired scRNA-seq and single-cell mass spectrometry (scMS) proteomics data to enable large-scale proteome inference for individual transcriptomic cells. To address the weak correlation between mRNA and protein abundance, msInfer replaces traditional anchor-based integration with a cell type–guided contrastive learning strategy for cross-omics alignment and employs an unsupervised weight generation module to infer protein abundances. Across extensive computational benchmarking and experimental validation, msInfer shows strong concordance between inferred and experimentally measured protein expression. msInfer facilitates the exploration of drug-induced molecular changes, supports the construction of single-cell multi-omics atlas and improves cell subtype annotation. Overall, msInfer provides a scalable and robust framework for bridging transcriptomic and proteomic measurements and enables comprehensive multi-omics characterization of cellular states.
Disease progression is often accompanied by alterations in tissue architecture. Although spatial omics technologies enable direct characterization of tissue organization, their limited tissue coverage makes the observed spatial architecture highly sensitive to sectioning position and limits the characterization of tissue-wide structural alterations. In addition, the high cost and limited scalability of current spatial omics platforms hamper their application in large clinical cohorts, limiting the discovery of reproducible architecture-associated disease patterns. Here, we present NicheDECODE, a deep learning framework that enables virtual spatial profiling of large-scale omics cohorts by reconstructing niche-level tissue architecture from non-spatial molecular measurements. Through graph-based integration of spatial continuity, gene-specific spatial dependencies, and functional coordination, NicheDECODE recovers tissue organizational states that were previously inaccessible from conventional omics data. Across benchmarking scenarios including cross-section, cross-donor, cross-dataset, cross-technology, and multi-omics settings, NicheDECODE demonstrates strong generalizability, robustness, and scalability. Furthermore, when applied to 832 dorsolateral prefrontal cortex (DLPFC) samples from the ROSMAP Alzheimer's disease cohort, it successfully reveals layer-resolved cortical gray matter atrophy and relative white matter shifts, with niche-level compositions more strongly associated with disease progression than gene expression or cell-type abundances. In breast cancer cohorts, NicheDECODE-derived spatial architecture profiles identify reproducible prognostic architectural states, improve risk stratification and survival prediction beyond conventional clinical variables, and generalize across independent patient cohorts. By bridging large-scale cohort data with tissue spatial architecture, NicheDECODE enables virtual spatial profiling for population-scale biomarker discovery, disease stratification, and low-cost exploration of disease-associated tissue remodeling.
Long-read sequencing has enabled comprehensive exploration of human genome at an unprecedented scale, particularly enhancing our understanding of structural variants (SVs). Phasing, a powerful approach for assigning haplotypes to sequencing reads, enables the generation of haplotype-aware call sets without requiring whole-genome assembly and provides a new direction for SV detection. Herein, we present cuteHap, a haplotype-aware SV detection method designed for phased long-read sequencing data. cuteHap fully leverages phased alignments and automatically selects a self-adaptive clustering strategy or a cluster credibility-prioritized beam search algorithm to achieve accurate haplotype-resolved SV calls. In addition, cuteHap incorporates a mosaic detection module to resolve somatic mosaicism. cuteHap achieved 6% and 3% higher F1-scores on Pacific Biosciences High-Fidelity (PacBio HiFi) and Oxford Nanopore Technologies (ONT) datasets, respectively, and detected a greater diversity of low-frequency SVs in tumor datasets. Its robust and high-performance SV detection facilitates the generation of high-quality haplotype-resolved call sets and advancing global genomic and genetic research.
Motivation Spatial transcriptomics techniques capture gene expression data and spatial coordinates, while simultaneously correlating them with tissue section images. This advantage makes Spatial transcriptomics data highly valuable for research, such as investigating disease mechanisms and cancer prognosis. However, the extended time and high cost of spatial transcriptomic sequencing currently limit further advancements in this field. The development of numerous deep learning methods aimed at predicting spatial transcriptomics from histology images has advanced significantly. However, these approaches often lack the ability to effectively integrate histology images with spatial transcriptomic data. Here, we propose GR2ST, a deep learning model that learns the underlying connections between image features and gene expression to predict spatial transcriptomics.Results GR2ST leverages a large pre-trained pathology model to extract high-level histological features. We designed a dual-branch graph architecture, consisting of a dynamic threshold-based functional graph and a radius-constrained spatial graph, to capture complex spot interactions within heterogeneous tissues. The model aligns histology images with gene expression representations through a multimodal contrastive learning framework. It achieves adaptive gene expression generation via a Cell-Type Guided Multi-Branch Regression Head supervised by a context-aware weighting network, which is further integrated with cross-sample retrieval to construct an ensemble prediction. The performance of the model is evaluated on three cancer-related spatial transcriptomics datasets, including cutaneous squamous cell carcinoma and two human breast cancer cohorts, to demonstrate its effectiveness and robustness.Availability https://github.com/zjl1109294570/GR2ST.
Plasma proteomics can provide a dynamic molecular readout of human health, but models that learn generalizable protein-expression patterns in population cohorts remain limited. Here we show that ProLM, a BERT-based plasma proteomics model pretrained on 15,499 relatively healthy UK Biobank participants, captures baseline protein-expression relationships and supports prediction of 16 common chronic diseases. After disease-specific fine-tuning, the ProLM-derived proteomic risk score outperformed the Age+Sex model for all 16 diseases, the cardiovascular disease (ASCVD) risk equation for 14 diseases and a 35-variable clinical PANEL score for 11 diseases. Model interpretation highlighted proteins including GDF15 whose expression changed more than 15 years before clinical diagnosis, and key findings were externally evaluated in the China Kadoorie Biobank. These results support plasma proteomics pretrained models as tools for early chronic-disease risk stratification, while prospective validation is needed before clinical implementation.
Functional remodeling of tissues is driven not only by changes in cellular composition, but also by shifts in how individual cell types engage biological programs across disease states, biological transitions and therapeutic contexts. Single-cell transcriptomics provides a direct view of cell-type functional heterogeneity, but its limited cohort scale restricts population-level association analyses, and conventional single-cell profiling lacks the spatial context needed to localize functional changes within tissues. Conversely, bulk RNA-seq cohorts and spatial transcriptomic profiles offer large-scale or spatially resolved measurements, yet their mixed expression signals obscure the cell-type origins of functional program activity. Here, we present FuncDECODE, a computational framework that infers the relative contributions of different cell types to functional programs from mixed transcriptomic profiles. By learning from single-cell references, FuncDECODE converts bulk samples and spatial spots into interpretable cell type–functional program features, enabling functional remodeling to be assigned to its likely cellular sources. Across bulk and spatial benchmarks, FuncDECODE consistently recovered cell-type-resolved functional contribution patterns and remained robust across donors, sequencing technologies, biological states and alternative functional program definitions. Across bulk and spatial transcriptomic applications, FuncDECODE revealed biologically and clinically relevant cell type–program remodeling. These features resolved immune-state transitions after vaccination, identified prognostic functional programs in breast cancer, and localized treatment-associated microenvironmental remodeling within spatial tissue architecture. Together, FuncDECODE provides a general framework for connecting single-cell functional heterogeneity with cohort-scale disease variation and spatial tissue organization.
Deconvolution algorithms estimate cell-type abundances from tissue-level data, enabling systematic cellular analysis of large cohorts. However, most deconvolution algorithms are specifically designed for single-omics data, thereby limiting their generalizability and scalability for various omics data from different cohorts. Here we present DECODE, a universal deconvolution framework for both cell types and cell states that can be applied to transcriptomic, proteomic and metabolomic data, and that seamlessly integrates diverse multiomics tissue datasets at the cellular level. DECODE fills the gap in metabolomics deconvolution and significantly outperformed state-of-the-art methods on different omics data across donors, disease conditions, healthy states, datasets and measurement platforms. In addition, DECODE exhibits high robustness in scenarios that are closer to real applications so it can accurately deconvolve known cell types even when the reference single-cell data are incomplete. DECODE will serve as a powerful tool for the fully extending multiomics cohort data into cellular level.
Somatic structural variants (SVs) are the predominant source of cancer driver mutations and play a critical role in oncogenesis. Comprehensive characterization of somatic SVs is critical for elucidating the mechanisms underlying tumorigenesis and for identifying biomarkers with diagnostic and therapeutic potential. However, their accurate detection remains challenging, primarily because most existing SV detection algorithms were originally developed for germline variants and are not well-suited to addressing the high heterogeneity of somatic mutations. In recent years, although several tools specifically designed for somatic SVs have emerged, their detection performance has not yet been rigorously validated. To bridge this gap, we conducted a comprehensive benchmarking of four leading somatic SV detection tools, namely, Sniffles2, Nanomonsv, Savana, and Severus, on the HG008 genome from Genome in a Bottle Consortium (GIAB). Their outputs were evaluated against the HG008 clonal somatic SV draft benchmark to assess overall performance. We further integrated the somatic SV callsets from multiple tools and compared them with the benchmark set, thereby establishing a multi-tool ensemble strategy for SV detection to achieve more accurate and comprehensive identification of somatic SVs.
How genetic variation regulates long non-coding RNA (lncRNA) expression across brain cell types remains poorly understood. A major barrier is the resolution–power trade-off between bulk and single-nucleus expression quantitative trait locus (eQTL) studies. Here, we quantify gene expression from cortical RNA-seq of 2443 individuals using an expanded transcriptome annotation and perform transcriptome-wide interaction eQTL mapping with deconvolution-derived cell-type proportions. Among 17,541 analyzed lncRNAs, we identify 3763 lncRNAs with cellular-context-dependent effects, of which 2,783 (74%) are not annotated by GENCODE. Colocalization with genome-wide association studies (GWAS) of brain-related traits identifies 118 lncRNA–trait colocalization events, approximately two-thirds of which are detectable only after modeling cellular context. Our study establishes a map of cellular-context-dependent genetic regulation in the human brain, providing a basis for nominating lncRNAs with potential roles in mediating genetic risk for brain disorders. Genetic regulation of long non-coding RNAs (lncRNAs) in the human brain remains poorly understood. Here the authors map cell-type-dependent genetic regulation of lncRNAs, highlighting their potential roles in brain disorders.
Motivation Although deep learning has significantly advanced the field of protein function prediction, current approaches are limited by their reliance on a narrow set of modalities. Specifically, they primarily rely on sequence patterns and treat protein domain data and functional labels merely as categorical tags. Consequently, they fail to capitalize on the semantic richness embedded within their textual definitions. These constraints hinder their ability to generalize to novel labels. To tackle this issue, we present MZSGO, a multimodal zero-shot framework that fuses evolutionary signals from protein language models with semantic features derived from large language models (LLMs). By employing an adaptive gated fusion mechanism, MZSGO effectively aligns sequence-based and text-based modalities to enable robust predictions for unseen labels.Results By unifying protein representations and functional annotations, we bridge the semantic gap that limits current approaches. Results indicate that while our model remains competitive on supervised benchmarks, it demonstrates a marked advantage over existing methods in zero-shot tasks. It specifically excels at recognizing previously unseen long-tail and novel Gene Ontology (GO) terms.Availability and implementation The source code and datasets are available at https://github.com/toxic-byte/MZSGO.
Accurate prediction of synthetic lethality (SL) is important for guiding the development of cancer drugs and therapies. SL prediction faces significant challenges in the effective fusion of heterogeneous multi-source data. Existing multimodal methods often suffer from "modality laziness" due to disparate convergence speeds, which hinders the exploitation of complementary information. This is also one reason why most existing SL prediction models cannot perform well on both pan-cancer and single-cancer SL pair prediction. In this study, we propose SynLeaF, a dual-stage multimodal fusion framework for SL prediction across pan- and single-cancer contexts. The framework employs a VAE-based cross-encoder with a product of experts mechanism to fuse four omics data types (gene expression, mutation, methylation, and CNV), while simultaneously utilizing a relational graph convolutional network to capture structured gene representations from biomedical knowledge graphs. To mitigate modality laziness, SynLeaF introduces a dual-stage training mechanism employing featurelevel knowledge distillation with adaptive uni-modal teacher and ensemble strategies. In extensive experiments across eight specific cancer types and a pancancer dataset, SynLeaF achieves superior performance in 17 out of 19 scenarios. Ablation studies and gradient analyses further validate the critical contributions of the proposed fusion and distillation mechanisms to model robustness and generalization. To facilitate community use, a web server is available at https://synleaf.bioinformatics-lilab.cn.
SUMMARY:Nanopore sequencing technology enables real-time sequencing and is widely used in rapid detection applications. However, in clinical scenarios, existing structural variant (SV) detection tools typically separate sequencing from computation, limiting their timeliness for clinical applications. To address this, we introduce cuteSV-OL, a novel framework designed for real-time SV discovery, which can be embedded within nanopore sequencing instruments to analyze data concurrently with its generation. Additionally, cuteSV-OL features a real-time SV detection rate evaluation module, allowing users to terminate sequencing early when appropriate, thereby reducing time and cost. Experimental results show that on a standard desktop computer, cuteSV-OL can perform real-time analysis during sequencing and complete SV calling within min after sequencing ends, achieving performance comparable to offline methods. This approach has the potential to enhance rapid clinical diagnostics. AVAILABILITY AND IMPLEMENTATION:cuteSV-OL is released under the MIT license and is available at https://github.com/gwmHIT/cuteSV-OL. It can also be installed via Bioconda or accessed through https://doi.org/10.5281/zenodo.17777436.
Structural variants (SVs) are a major source of genomic diversity, yet their discovery remains challenging due to repetitive genomic contexts, alignment ambiguity, and the trade-off between sequencing cost and read length. Here we introduce HitSV, which substantially improves SV discovery by implementing repetitiveness and signature density aware breakpoint recognition coupled with precise haplotype-resolved local assembly, thereby enabling base-resolution SV reconstruction and genotyping across various sequencing technologies. HitSV is 12-68% (long-read), 3%-36% (short-read) and 13% (hybrid-sequencing), respectively, more accurate than state-of-the-art SV callers across different coverages. Applying HitSV to the 1KGP Phase 4 cohort, we identified 31.5% more SVs, substantially reshaping allele-frequency landscapes. Notably, analysis of a large Chinese long-read cohort uncovers tandem repeat–mobile element composite arrays as a prevalent and multi-allelic class of complex SVs, highlighting composite repeat architectures as a fundamental hallmark of human genomes.
Long non-coding RNAs (lncRNAs) play essential roles in the pathogenesis of neurological disease, but the cellular heterogeneity of their genetic regulation remains poorly understood, with a key barrier being the resolution–power trade-off between bulk and single-nucleus eQTL studies. Here, we profiled 17,541 lncRNAs including both GENCODE-annotated and newly assembled transcripts in the dorsolateral prefrontal cortex (DLPFC) of 2,443 individuals and performed interaction eQTL (ieQTL) mapping based on computational deconvolution. We identified 3,763 interaction-effect lncRNAs (ieLncRNAs) across brain cell types, including 2,783 (74%) previously unannotated in GENCODE and 3,262 (86%) not captured in current single-nucleus eQTL resources. Integration of lncRNA-ieQTLs with GWAS data of brain-related diseases indicated that nearly two-thirds of colocalization signals were missed by bulk eQTL analyses, underscoring the added interpretive resolution afforded by modeling cellular context. We further illustrate this with representative loci, including AC096667.1 and lncRNAs overlapping NRGN, which show cell-context–dependent colocalization with neuroticism and schizophrenia risk, respectively, refining the cellular interpretation of genetic risk. Collectively, this resource advances our understanding of lncRNA contributions to brain disease by uncovering regulatory effects that are obscured in bulk tissues and not captured by single-nucleus eQTL resources.
Cognitive performance has been found to be associated with the complex structure of human cerebral cortex. However, due to the limitations of previous cortical parcellation atlases, the cortical genetic patterns determining cognitive performance remain unknown. Here, we utilized the latest Human Connectome Project Multi-Modal Parcellation (HCP-MMP) atlas to divide the cerebral cortex into 180 regions per hemisphere. We investigated the shared genetic architecture between four types of magnetic resonance imaging (MRI)-derived cortical phenotypes and cognitive performance using large-scale genome-wide association studies (N for cortical phenotypes = 36,843; N for cognitive performance = 257,828). We observed extensive genetic overlap between cortical surface area, volume, and local gyrification index (LGI) with cognitive performance, particularly the subregions in the insula, cingulate cortex, and ventromedial prefrontal cortex, many of which were novel findings. However, the thickness of some prefrontal regions was negatively correlated with cognitive performance. We identified 18 and 312 shared genetic loci for global and regional cortical phenotypes with cognitive performance, respectively. These genetic loci were involved in a substantial number of biological processes related to neuronal development, cell growth, and neuronal death or apoptosis. The cortical patterns defined by these shared loci were established entirely along the sensorimotor-association (S-A) axis. These findings provide new insights into the genetic relationship between cognitive performance and the human cerebral cortex under a more refined multimodal cortical parcellation scheme.
DNA methylation constitutes the primary epigenetic language mediating organismal phenotypic plasticity. Establishing a cohort-level genomic methylation landscape featuring wide geographical diversity is fundamental for dissecting its genetic and environmental attributes. Leveraging nanopore sequencing's strength in genome-methylome co-sequencing, we generated a whole-genome, haplotype-resolved methylation atlas for 106 individuals from 19 provinces across China. The atlas identified 27,609,354 CpG sites genome-wide, with notably more informed gene proximal regions and CpG islands compared to whole-genome bisulfite sequencing. Detailed analyses revealed genomic structural variants as a pervasive covariate of DNA methylation, with a remarkable 2-fold compensation effect found in genome-wide heterozygous deletions. On the other hand, habitat altitude is found to be a strong environmental determinant of DNA methylation. We established a quantitative relationship between altitude and methylation states and identified a gene set strictly responsive to altitude differences, revealing epigenetically regulated genes such as PRDM16, EPHB2 and WNT7A. The methylation atlas provides a reference resource to facilitate further explorations into human epigenetics.
Accurate detection of single-nucleotide variants (SNVs) and small insertions/deletions (indels) from second-generation sequencing (NGS) data is essential for clinical applications such as cancer diagnostics, infectious disease monitoring, and rapid genetic screening. However, conventional variant calling pipelines, such as GATK, decouple analysis from sequencing, deferring detection until sequencing is fully completed. We introduce RVC, a real-time variant calling framework tailored for cycle-based NGS workflows. RVC incrementally processes partially sequenced reads and continuously updates variant evidence using a scanline-based alignment algorithm and a lightweight binomial scoring model. This design enables progressive, low-latency SNVs and indels detection during sequencing, without disrupting the sequencing pipeline. In benchmark experiments using the HG002 dataset, RVC completed variant calling within tens of minutes after sequencing, significantly outperforming GATK in runtime. By tightly integrating analysis with sequencing output, RVC bridges the gap between sequencing speed and clinical responsiveness, offering a scalable and practical solution for real-time genomic diagnostics.
Deconvolution algorithm enables estimation of cell type abundances from tissue-level data, providing a crucial way for exploring plentiful cohort data at the cellular level. However, most deconvolution algorithms are specifically designed for single-omics data, thereby limiting their generalizability and scalability for multiomics data from different cohorts. A deconvolution algorithm applicable to various omics data can use cell abundance as a bridge to improve the comparability of different cohorts. Here, we developed DECODE, a universal deconvolution framework of both cell type and cell state designed for transcriptomics, proteomics and metabolomics data, which seamlessly integrates diverse multiomics tissue datasets in cellular level. DECODE fills the gap in metabolomics deconvolution and significantly outperformed state-of-the-art methods on different omics data across donors, disease conditions, healthy states, and measurement platforms. In addition, DECODE exhibits high robustness in scenarios that are closer to real applications, that it can accurately deconvolve known cell types even when the reference single-cell data incomplete all cell types of target tissue. DECODE will serve as a powerful tool for the fully extending multiomics cohort data into cellular level.
BACKGROUND:The development of long-read sequencing is promising for the high-quality and comprehensive de novo assembly for various species around the world. However, it is still challenging for assemblers to handle thousands of genomes, tens of gigabase-level assembly sizes, and terabase-level datasets efficiently, which is a bottleneck to large-scale de novo sequencing studies. A major cause is the read overlapping graph construction that state-of-the-art tools usually have to cost terabyte-level RAM space and tens of days for large genomes. Such lower performance and scalability are not suited to handle the numerous samples being sequenced. FINDINGS:Herein, we propose xRead, a novel iterative overlapping graph construction approach that achieves high performance, scalability, and yield simultaneously. Under the guidance of its coverage-based model, xRead converts read-overlapping to heuristic read-mapping and incremental graph construction tasks with highly controllable RAM space and faster speed. It enables the processing of very large datasets (such as the 1.28 Tb Ambystoma mexicanum dataset) with less than 64 GB RAM and obviously lower time costs. Moreover, benchmarks suggest that it can produce highly accurate and well-connected overlapping graphs, which are also supportive of various kinds of downstream assembly strategies. CONCLUSIONS:xRead is able to break through the major bottleneck to graph construction and lays a new foundation for de novo assembly. This tool is suited to handle a large number of datasets from large genomes and may play important roles in many de novo sequencing studies.