
Paired spatial multi-omics provides a supervised basis for learning RNA-protein correspondence in situ, but predicting protein abundance from spatial transcriptomic data alone remains challenging across tissue contexts and protein panels. Here, we present DPAS-Graph, an adaptive relation-learning framework for spatial RNA-to-protein prediction. Rather than directly merging spatial proximity and transcriptomic similarity as fixed graph priors, DPAS-Graph represents them as two relation channels on a shared edge support and updates their contributions during representation learning for protein prediction. Its Niche-Coupled Field Encoder combines layer-wise edge-relation modeling, intra-branch relation refinement, and cross-branch residual correction to learn spot representations for protein abundance prediction. In a leave-one-dataset-out benchmark across seven paired spatial multi-omics datasets, DPAS-Graph achieved lower aggregate prediction errors and improved spot-level agreement of protein expression profiles, with gains mainly reflected in error-based metrics and PCC-Spot. Spatial autocorrelation and protein-derived domain agreement analyses were further used to characterize the spatial behavior of the predicted protein maps. When applied to external RNA-only spatial sections, DPAS-Graph generated qualitatively interpretable marker-level virtual protein maps, illustrating its use as a complementary tool for protein-level interpretation of transcriptomics-only spatial data.
The evolution of molecular visualization software represents a pivotal yet often underappreciated dimension of computational structural biology and bioinformatics. As the scale and complexity of structural datasets increase, spanning from cryo-electron microscopy reconstructions to proteome-wide AI-predicted models, visualization tools have become essential for data interpretation, hypothesis generation, and effective scientific communication. Widely used platforms such as PyMOL, UCSF Chimera, and emerging virtual reality environments reflect decades of progressive innovation driven by advancements in computational hardware, graphics rendering, and algorithm development. This review traces the historical development of molecular graphics software, from early vector-based applications like FRODO and O, through interactive tools including RasMol and VMD, to modern real-time 3D visualization frameworks. We highlight key technological milestones and design principles that continue to shape contemporary software architectures, underscoring the interdisciplinary collaborations among chemists, structural biologists, and computer scientists that have propelled the field. Given the current integration of artificial intelligence, immersive visualization, and cloud computing in structural bioinformatics workflows, revisiting this topic provides critical context for guiding the design of more intuitive, scalable, and accessible visualization platforms. By documenting and critically analyzing this trajectory, the present work aims to honor the foundational contributions to molecular graphics while positioning the field as a dynamic and integral component of future bioinformatics research and computational discovery.
Pleiotropic genetic loci have been increasingly reported in cancer, and identifying genetic variants with pleiotropic associations can reveal shared biological pathways influencing multiple cancers. Using summary statistics from genome-wide association studies for 37 cancer types (N = 433 836), we identified extensive genome-wide and local genetic correlations among cancers. Through pairwise pleiotropic analysis, we identified 75 243 significant pleiotropic single nucleotide polymorphisms (SNPs) across 372 cancer pairs, among which 3472 were lead SNPs with potential regulatory functions. Using FUMA and MAGMA, we identified 2527 pleiotropic risk loci and 4272 candidate pleiotropic genes. Notably, genes such as TERT (5p15.33), POU5F1B (8q24.21), and FANCA (16q24.3) exhibited widespread pleiotropy across multiple cancer types. Pathway enrichment analysis highlighted the critical roles of pigment synthesis, metabolism, and apoptosis in skin-related cancers, while cross-cancer enrichment analysis emphasized pathways related to apoptosis, chromatin structure, and intermediate filaments. We also identified 33 novel functional genes harboring previously unreported cancer risk variants. Drug-gene interaction analysis revealed several repositionable FDA-approved drugs. Importantly, drug sensitivity assays demonstrated that bosutinib and cobimetinib exhibited promising therapeutic potential in breast cancer cell lines. Finally, we developed the PleioCancer database (https://gonglab.hzau.edu.cn/PleioCancer/), providing a comprehensive resource for cancer pleiotropy research. These findings have important implications for carcinogenesis cancer, prevention and treatment.
Predicting biomolecular interactions is fundamental to understanding cellular mechanisms and advancing drug discovery. However, biomolecular interactions exhibit immense diversity across multiple dimensions. Most existing computational methods are designed to handle one specific task or data modality, which limits their applicability and generalization capability in broader scenarios. To address this methodological rigidity, we propose a flexible framework for multi-modal feature fusion in biomolecular interaction prediction (FlexBIP). The core of FlexBIP lies in its modular architecture, which decouples intrinsic molecular features from complex graph topologies, enabling the adaptive integration of node attributes, edge properties, and auxiliary graph information. The flexible fusion methodology breaks through the limitations of task-specific models. This design enables FlexBIP to adaptively process and integrate biological data of different types and from various sources, including homogeneous interactions between molecules of the same type, heterogeneous interactions between different molecular classes, as well as qualitative binary, multi-class, and quantitative regression prediction tasks. Our research has yielded exciting results. In extensive testing across 15 benchmark datasets, covering 8 major categories of biomolecular associations, FlexBIP's performance comprehensively surpasses that of 25 state-of-the-art specialized models. Crucially, in data-scarce "cold-start" scenarios that simulate the discovery of new molecules, FlexBIP continues to demonstrate remarkable robustness and predictive accuracy. Furthermore, FlexBIP provides robust and reliable interpretability for various downstream analysis tasks.
Large-scale single-cell and spatial transcriptomic atlases enable the study of developmental processes at high resolution. However, most datasets capture only static snapshots of cells, making it difficult to infer continuous biological time from transcriptomic profiles. Existing temporal inference methods often show limited robustness across heterogeneous datasets, and recent single-cell foundation models, although powerful for representation learning, are not designed to capture continuous temporal relationships. We present Gene-Chronos, a parameter-efficient framework for developmental time inference built on a frozen pretrained Geneformer backbone. The model introduces learnable temporal prompt tokens and a temporal contrastive objective to extract time-informative signals and encourage temporally coherent organization of cell representations. Across multiple benchmark datasets spanning diverse species and developmental stages, Gene-Chronos outperforms existing approaches and demonstrates strong generalization to previously unseen samples. Attention-based analyses further identify genes associated with developmental progression, providing interpretable insights into temporal gene expression dynamics.
Spatial transcriptomics (ST) has emerged as a powerful approach for profiling gene expression in spatial tissue context; yet, its quantitative accuracy remains substantially compromised by amplification bias introduced during the complex library preparation process. These biases arise at multiple stages and accumulate throughout the experimental workflow, distorting transcript abundance, reducing detection sensitivity, and ultimately confounding downstream spatial analyses. This review systematically analyzes amplification bias in ST. We examine how input templates, oligonucleotide components characteristics, enzymatic properties, and experimental conditions collectively contribute to amplification bias, and discuss how these factors propagate through the workflow to generate systematic distortions in data. We further review and critically compare existing strategies for mitigation, encompassing both experimental optimizations and computational approaches and propose a practical decision framework for selecting amplification-bias mitigation strategies according to platform type, sample quality, and RNA input levels. Finally, we outline key challenges and future directions, emphasizing the need for integrative solutions that jointly consider experimental design and computational modeling. This work provides practical guidance for improving data fidelity and interpretation in ST.
Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue Cα-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.
Genome-wide association studies have successfully identified many genetic variants associated with complex traits. However, most existing methods target autosomes rather than X chromosome, and several existing X chromosome-wide association studies (XWAS) at quantitative trait loci (QTL) largely focus on unrelated individuals, with limited attention to general pedigrees or mixture of general pedigrees and additional unrelated individuals (called the mixed data for brevity). In this study, we propose nine novel methods for XWAS at QTL in the mixed data (${\mathrm{MQX}}_{\mathrm{cat}}$, ${\mathrm{MQZ}}_{\mathrm{max}}$, ${\mathrm{MT}}_{\mathrm{plinkw}}$, ${\mathrm{MT}}_{\mathrm{chenw}}$, $\mathrm{MwM}3\mathrm{VNA}$, ${\mathrm{MQMVX}}_{\mathrm{cat}}$, ${\mathrm{MQMVZ}}_{\mathrm{max}}$, $\mathrm{MpMV}$, and $\mathrm{McMV}$), also applicable to general pedigrees alone. The first four methods test for mean differences across genotypes; the latter four test for differences in both means and variances; $\mathrm{MwM}3\mathrm{VNA}$ tests for variance differences only. All mean-based and mean-variance-based methods incorporate X chromosome inactivation information, and all nine methods consider genetic relatedness in pedigrees. Simulation studies confirm well-controlled type I error rates, and inclusion of pedigrees significantly improves statistical power. Note that there has been no study focusing on X chromosome for the mixed data or general pedigrees from UK Biobank database, so we apply our proposed methods to this dataset, which identify five total cholesterol (TC)-associated and 13 low-density lipoprotein cholesterol (LDL-C)-associated single nucleotide polymorphisms (SNPs). Linkage disequilibrium (LD) analysis reveals that these SNPs fall into three distinct LD blocks. Functional annotation and gene ontology enrichment analysis reveal 16 and 28 enriched pathways for TC-associated and LDL-C-associated genes, respectively. These methods provide robust and powerful tools for XWAS at QTL in both mixed data and general pedigrees.
Human RNA sequencing (RNA-seq) data originally generated for human transcriptome profiling are overwhelmingly dominated by host sequences, yet they often contain a small fraction of non-human reads that can be exploited for microbial detection. When such datasets are repurposed for secondary microbiome-oriented analyses, extracting and accurately classifying this weak microbial signal becomes technically challenging, and no ready-to-use pipeline currently exists. In this study, we evaluate computational strategies for filtering host reads and classifying microbial transcripts in host-dominated RNA sequencing data. We compare assembly-based approaches similar to those used in a previous study focusing on microbial translocation with state-of-the-art assembly-free methods, and assess their respective strengths and limitations using simulated datasets reflecting low microbial abundance. Our results show that assembly-based methods yield accurate taxonomic predictions but struggle at low read depth, whereas assembly-free methods are more robust in sparse settings at the cost of reduced precision. To leverage the complementarity of both approaches, we propose a hybrid pipeline that integrates assembly-based and assembly-free classification. On simulated data, this hybrid strategy improves microbial classification performance compared with either approach alone. Application to a real human metatranscriptomic dataset analyzed in a microbial translocation context illustrates the broader microbial signal captured by the hybrid approach, despite intrinsic challenges related to the absence of reliable ground truth and the risk of host read misclassification. Our work provides a framework for extracting microbial signals from host-dominated human metatranscriptomes, enabling the reuse of existing transcriptomic datasets for microbiome-related analyses, including but not limited to microbial translocation studies.
DNA methylation alterations are early and stable hallmarks of cancer and represent promising biomarkers for non-invasive detection using circulating cell-free DNA (cfDNA). However, current computational approaches often model DNA sequence and methylation features separately and struggle to capture complex read-level methylation architecture in heterogeneous, low-signal liquid biopsy data. Here, we present DNAmBERT, a Transformer-based deep learning framework designed to jointly model DNA sequence context and read-level methylation haplotype structure from cfDNA methylation sequencing data. DNAmBERT integrates k-mer-encoded DNA sequences with methylation haplotype tokens using a unified representation and masked language modelling objective, enabling context-aware learning of sequence-epigenetic dependencies through self-attention. We evaluated DNAmBERT across multiple cfDNA methylation platforms (RRBS, cfRRBS, and cfMethyl-seq) and cancer types, including colorectal cancer, lung adenocarcinoma and hepatocellular carcinoma. In binary classification tasks, the model achieved high performance across platforms (AUC up to 0.99-1.00) and outperformed conventional machine learning and existing deep learning approaches. Aggregation of read-level predictions enabled quantitative tumour probability estimation at the sample level. Beyond binary detection, DNAmBERT supported multi-cancer and stage-aware classification, including early-stage disease, with multiclass AUC values up to 0.99. The framework further demonstrated effective cross-cancer transfer learning, maintaining robust performance under limited data availability. These results indicate that integrated sequence-haplotype representation learning provides an accurate and scalable approach for cfDNA-based multi-cancer detection.
Antimicrobial resistance (AMR) threatens microbiology and microbiome bioinformatics because resistance phenotypes are shaped by interactions among genes, mobile genetic elements, and functional environments across microbial communities. Prioritizing resistance determinants requires models that reason across knowledge graphs (KGs) linking genes, proteins, pathways, drugs, and microbial phenotypes. Existing graph-based methods compress this evidence into scalar scores, whereas large language models can produce explanations not grounded in structured evidence. We developed AResKGLM (Antimicrobial Resistance Knowledge Graph Language Model), a graph-grounded language-model framework for interpretable microbial AMR bioinformatics that serializes breadth-first-search-retrieved multi-hop paths and per-entity biomedical descriptions into a structured Context-Path-Question prompt. Llama-3-8B and DeepSeek-R1-7B are adapted with QLoRA to produce binary link predictions and concise reasoning traces. On the KIDs benchmark, AResKGLM (Llama-3-8B) achieved F1 = 0.8482, outperforming KG-BERT (0.7213), NBFNet (0.5260), and ULTRA (0.2541) (paired Wilcoxon $p = 1.2 \times 10^{-7}$). Its advantage increased with reasoning depth: F1 decreased from 0.9197 at 2 hops to 0.8148 at 6 hops, whereas KG-BERT dropped from 0.8110 to 0.6716. Counterfactual path corruption produced an apparent F1 of 0.000, mechanically forced by the probe label assignment; the operative diagnostic is the per-sample flip rate (0.04-0.16), consistent with sensitivity to supplied biological evidence rather than reliance on pretrained priors alone. Cross-species evaluation yielded F1 = 0.81-0.88 with Matthews correlation coefficient (MCC) = 0.35-0.54 on Mycobacterium tuberculosis, Pseudomonas aeruginosa, and Staphylococcus aureus. Temporal ranking of 81 post-2022 gene-drug associations achieved Precision@20 = 100% and AUC-PR = 0.855. AResKGLM offers an interpretable, reproducible framework for multi-hop AMR reasoning, linking candidate prioritization with mechanism-oriented hypothesis generation.
Comprehensive two-dimensional gas chromatography-mass spectrometry (GC × GC-MS) is a powerful tool for analysing complex mixtures, but its wide use is limited by the lack of robust and transparent inter-sample peak alignment workflows. Commercial solutions often rely on proprietary algorithms, while user-friendly open-source alternatives remain scarce. Here, we introduce jAligner4GCxGC, a new open-access Julia package for inter-sample GC × GC-MS peak alignment and benchmark its performance within practical GC × GC-MS data processing workflows against Guineu (open-source) and ChromaTOF Sync 2D (commercial). The package performs data preprocessing, alignment, and postprocessing including feature merging, library searching, and compound data retrieval. Three indoor dust samples, including certified reference material, were spiked with >150 reference compounds at concentrations of 0.5, 5.0, and 50 pg/μl and analysed using workflows optimized for each application. Performance was evaluated based on recovery of correctly aligned spiked compounds and workflow processing time. Because the software packages differ in peak detection and preprocessing, comparisons involving Sync 2D should be interpreted at workflow rather than alignment algorithm level. At 50 pg/μl, Guineu and jAligner4GCxGC detected 94% and 92% of compounds, respectively, outperforming Sync 2D (82%). At 5 and 0.5 pg/μl, Sync 2D detected the highest proportions (79% and 71%), whereas Guineu and jAligner4GCxGC performed identically (68% and 33%). Alignment-related processing required seconds for Guineu, 4 min for jAligner4GCxGC, and 9 min for Sync 2D. However, Sync 2D provided the quickest end-to-end workflow. These results highlight practical trade-offs between commercial and open-source workflows and demonstrate that jAligner4GCxGC provides a transparent and reproducible solution for GC × GC-MS peak alignment.
Genomic prediction has become a central paradigm in biology, enabling quantitative inference of genetic contributions to complex traits across humans, animals, and plants. Although genomic research in human genetics and animal breeding shares a highly homologous methodological foundation, significant barriers persist in their analytical paradigms and application scenarios. This study aims to promote cross-disciplinary integration by introducing human-derived polygenic scores (PGS) algorithms into animal genomic selection (GS) and proposing a PGS-GS framework with a preliminary weighting-based implementation. We systematically benchmarked the predictive performance and computational efficiency of 20 algorithms, including classical linear models, machine learning, PGS, and PGS-GS using both array and whole-genome sequencing (WGS) data across four major agricultural species: beef cattle, sheep, pigs, and chickens. Our results demonstrate that PGS and PGS-GS algorithms achieve predictive accuracy competitive with genomic best linear unbiased prediction (GBLUP) while offering markedly higher computational efficiency. Moreover, incorporating PGS-derived prior information into weighted linear and non-linear models outperformed conventional weighted GBLUP. The results provide empirical evidence to inform algorithm selection and highlight the potential of integrating human-derived PGS methodologies into animal genomic prediction frameworks.
Spatial transcriptomics (ST) has enabled direct interrogation of cell-cell communication (CCC) within intact tissues, providing critical spatial context that is lost in single-cell RNA-sequencing-based inference and allowing more accurate identification of physically plausible and spatially organized interactions. A rapidly expanding community of computational tools has emerged to decode CCC from ST data. Here, we provide a comprehensive review of the conceptual evolution and methodological landscape of spatial CCC inference, classifying existing approaches into two major trajectories. One trajectory, spatial pattern-based methods, assumes CCC events manifest as identifiable spatial patterns, such as colocalization, coordinated spatial signals, or higher-order spatial organization captured by deep learning models. The other trajectory, expression modulation-based approaches, assumes that CCC events influence the transcriptomic state of receiver cells. We systematically dissect their biological assumptions, statistical and deep learning frameworks, strengths, and limitations, and highlight emerging challenges in validation, benchmarking, multimodal integration, and tissue-specific modeling. Finally, we outline future directions toward achieving dynamic, multilayered reconstruction of inter- and intracellular communication, de novo signaling, and integrative multi-omics modeling.
T cell receptor (TR) genes are essential components of the adaptive immune receptor repertoire, and growing evidence links TR germline variants to immune-related diseases. However, their highly allelic diversity and sequence homology make them challenging dark regions of the human genome. Current TR genotyping tools have limited support for population-scale studies using standard-depth (30×) whole-genome sequencing (WGS), leaving a critical gap. We present germline Adaptive Immune Receptor Repertoire (gAIRR)-wgs, the first highly resource-efficient workflow specifically designed for high-resolution TR allele typing from short-read WGS. Benchmarking against 44 assembly-validated Human Pangenome Reference Consortium Release 1 subjects showed high overall accuracy performance across TR loci (mean F1/accuracy: 0.996/0.997), with comparable results in an independent cohort of 182 Release 2 individuals (0.984/0.988). Applying gAIRR-wgs to 1492 Taiwan Biobank (TWB) participants, we identified 450 novel TR alleles absent from the international ImMunoGeneTics (IMGT) information system database, accounting for 57.5% of all identified TR alleles and representing an ~102% expansion of the current IMGT TR repertoire-277 of which were cross-validated in non-East Asian cohorts-and 109 novel TR V alleles with allele frequencies >1% in the TWB. Notably, the tool uncovered population-specific structural polymorphisms, including T cell receptor gamma variable (TRGV) genes (TRGV4/TRGV5 deletions) and T cell receptor beta variable (TRBV) genes (TRBV3-2/TRBV4-3 insertion/deletion), which were overlooked by Illumina Dynamic Read Analysis for GENomics (DRAGEN). Furthermore, we identified 34 TR genes exhibiting significant allelic divergence between Taiwanese and global populations. By enabling accurate TR genotyping from 30× WGS data, gAIRR-wgs effectively unlocks the immunogenomic potential of massive biobank resources, bridging the gap between standard genomic surveys and adaptive immune repertoire analysis.
The KEGG Orthology (KO) system links DNA and protein sequences to biological functions and pathways, providing a curated, fundamental, and consistent annotation framework across all domains of life. While accurate, traditional sequence alignment-based annotation methods are computationally expensive, which severely limits their application in large-scale datasets. To address this challenge, we introduce Deep KEGG Orthology and Links Annotation (DeepKOALA), a deep learning approach based on Gated Recurrent Units (GRU), which frames KO annotation as an open-set recognition task. This design reduces false positives arising from out-of-scope sequences and, together with a lightweight GRU backbone, enables high-throughput annotation. The GRU-based model was benchmarked against four other deep learning architectures and showed the best balance between speed and accuracy. We then trained a GRU-based model, DeepKOALA, and performed a cross-species evaluation against existing KO annotation tools. In this comparison, DeepKOALA achieved a F1 of 83.37%, which is comparable to existing alignment-based tools. Meanwhile, the speed of DeepKOALA was 36.5-fold faster than Blast KEGG Orthology and Links Annotation (BlastKOALA). We also provide a specialized fragment model for handling incomplete sequences and an optional multi-domain mode. Together, these features make DeepKOALA a scalable and efficient option for high-throughput function annotation.
Single-cell DNA methylation (scDNAm) profiling is revolutionizing our understanding of epigenetic control of gene expression, but its accurate analysis is severely hindered by extreme data sparsity. While imputation methods have undergone remarkable development in recent years, a rigorous benchmark to guide method selection remains absent. We established the first systematic benchmarking framework for scDNAm imputation, subjecting five state-of-the-art methods to a comprehensive evaluation across 13 published experimental scDNAm datasets. Performance was systematically assessed across seven critical dimensions: accuracy, sensitivity to data characteristics, scalability, robustness to data splitting strategies, inter-dataset generalizability, convergence behavior, and computational efficiency. Through rigorous statistical analysis, we dissected the influence of intrinsic data attributes and model architectures on the fidelity of scDNAm imputation to provide guidance for selecting appropriate methods for given scenarios. Furthermore, based on the benchmark-identified limitations, we proposed a dual-view strategy to address the performance bottlenecks of existing methods: at the model view, we developed BridgeCpG, an ensemble strategy to integrate complementary modeling strengths to overcome single-model limitations; at the data view, we introduced an adaptive divide-and-conquer strategy to partition highly heterogeneous datasets into several homogeneous subsets amenable to accurate imputation, followed by aggregating the sub-results. This integrated framework, spanning both model and data views, delivers quantitative analyses, scenario-aware selection guidelines, and targeted innovative strategies, establishing a rigorous, enabling foundation for accurate, high-throughput, and scalable next-generation single-cell epigenomic analysis.
Network-based analyses of omics data are widely used and, while many of these methods have been adapted to single-cell scenarios, they often remain memory- and space-intensive. As a result, they are better suited to batch data or smaller datasets. Furthermore, the application of network-based methods in multi-omics often relies on similarity-based networks, which lack structurally discrete topologies. This limitation may reduce the effectiveness of graph-based methods that were initially designed for topologies with better defined structures. We propose Subset-Contrastive multi-Omics Network Embedding (SCONE), a method that employs contrastive learning techniques on large datasets through a scalable subgraph contrastive approach. By exploiting the pairwise similarity basis of many network-based omics methods, we transformed this characteristic into a strength, developing an approach that aims to achieve scalable and effective analysis. Our method demonstrates synergistic omics integration for cell type clustering in single-cell data. Additionally, we evaluate its performance in a bulk multi-omics integration scenario, where SCONE performs comparable with the state-of-the-art despite utilizing limited views of the original data. We anticipate that our findings will motivate further research into the use of subset contrastive methods for omics data.
Phenotypic variation is shaped by genotype, environment, and their interactions. Accurately predicting crop performance across diverse environments therefore requires models capable of capturing these complex and context-dependent relationships. Here, we developed GE-BiFormer, an explainable multimodal deep learning framework for genotype-by-environment prediction. GE-BiFormer integrates genomic and enviromic information through dual-path feature disentanglement, tokenized bidirectional cross-attention, and mixture-of-experts routing, enabling fine-grained modeling of genetic effects, environmental responses, and their interplay. We evaluated GE-BiFormer using the Genomes to Fields maize dataset. After preprocessing, the dataset contained approximately 360,000 non-missing genotype-environment-trait observations across six traits. The evaluation focused on three breeding-relevant scenarios: predicting known genotypes in unseen environments, predicting novel genotypes in known environments, and predicting novel genotypes in entirely untested environments. Across these scenarios, GE-BiFormer consistently outperformed GBLUP, classical machine learning methods, and recent deep learning approaches. External validation on an independent winter wheat dataset further demonstrated the broad applicability and cross-dataset robustness of GE-BiFormer across crop species. In addition, SHAP analysis identified biologically interpretable environmental drivers, while mixture-of-experts routing revealed trait-specific computational specialization. Together, these results demonstrate that GE-BiFormer provides a practical and interpretable framework for environment-aware genomic selection and climate-adaptive crop breeding.
De novo mutations (DNMs) play a crucial role in the pathogenesis and clinical interpretation of genetic diseases. However, existing pathogenicity prediction methods either uniformly handle all variants or focus on specific variant types, lacking a systematic prediction framework for DNMs and failing to effectively integrate functional annotation with clinical evidence. In this study, we present DeNovoSeer, a deep learning pathogenicity prediction framework for coding-region DNMs. Built upon a labeling system that combines high-confidence ClinVar annotations with complementary phenotype information from Gene4Denovo, the method integrates multi-source functional annotations and clinical evidence derived from the ACMG/AMP guidelines. It employs a semi-supervised convolutional-dilated convolution hybrid network architecture, enabling joint representation learning across multiple tasks under limited labeled data. On the Gene4Denovo test set, DeNovoSeer demonstrated stable performance across 10 independent random splits, achieving an AUC of 0.876 ± 0.007 and an AP of 0.881 ± 0.008, while outperforming existing tools. SHAP-based analysis provides feature-level attribution of model predictions, revealing biologically and clinically meaningful evidence patterns and offering interpretable support for variant assessment. This study proposes a systematic framework for pathogenicity prediction of coding-region DNMs, integrating robust label construction, semi-supervised representation learning, and clinical evidence integration. It provides new methodological support for molecular diagnosis and the interpretation of disease mechanisms.