
Systematic identification of transcriptional and epigenetic regulators (TERs) remains a challenge in myeloid leukemia. Current methods for TER identification typically rely on single data types and show limited power for long-range regulatory interactions. Here we present TERfinder, a deep learning framework that integrates multi-omics features to predict enhancer–promoter interactions (EPIs) and characterize transcriptional regulatory programs in myeloid leukemia. TERfinder achieved AUC 0.9644 and AUPRC 0.9584 on held-out chromosomes, exceeding baselines without autoencoder or histone features (Table S7; DeLong test, P < 0.01). Motif enrichment identified C/EBP and ETV family TFs as candidate regulators. Single-cell regulon analysis confirmed their activity in AML progenitor populations. Single-cell analysis showed SPI1- and CEBPA-centered regulatory networks active in AML blasts, and their activity was associated with poor overall survival. A four-gene expression signature (SPI1, CEBPA, MYC, PTPN6) stratified AML patients into high- and low-risk groups (log-rank P < 0.01). TERfinder provides a framework for multi-omics regulatory inference and candidate TF identification in myeloid leukemia.
Pathway activity scoring is a widely used step in single-cell RNA-seq (scRNA-seq) analysis, yet where it is applied, method choice is rarely guided by systematic, multi-criterion evidence in the pseudobulk case-control regime that now dominates applied single-cell disease studies. Prior single-cell pathway-scoring benchmarks have focused on perturbation ground truth at cell-level resolution, leaving the donor-level pseudobulk setting—and the question of why methods disagree—largely unaddressed. We present PathwayBench, a benchmark comparing five widely used pathway scoring methods (ssGSEA, GSVA, z-score, AUCell, UCell) across eight pseudobulked scRNA-seq case-control datasets spanning five tissues (brain, heart, kidney, lung, blood) and 682 donors. Methods are evaluated against five criteria covering biological relevance (direction accuracy, AUROC, effect size) and four robustness axes (aggregation, outlier, normalization, sample-size stability). We identify rank-window truncation as a two-sided mechanism by which top-window rank methods (AUCell, UCell) lose biological signal: they go blind when pathway genes fall outside the fixed top-rank window, and they attenuate or invert when non-pathway competitor genes invade that window. A 49-condition simulation sweep shows the effect is broad across the parameter grid: AUCell returns a null score in 33 of 49 conditions when pathway genes fall below its default window, and both rank methods invert sign under heavy competitor burden. A window-parameter experiment on real data shows that widening the rank window recovers most of the lost signal at an intermediate setting, confirming that the default single-cell window is mis-set for dense pseudobulk profiles. Real-data illustration comes from extracellular matrix (ECM) remodeling in chronic kidney disease, where top-window rank methods produce near-zero or wrong-direction effects across three normalizations. Discretizing per-criterion performance as Good/Intermediate/Poor indicates that no single method satisfies all five criteria under the primary thresholds. Because this classification is sensitive to threshold and weighting choices, we report the underlying continuous scores as the primary evidence and use the discrete scoreboard only as a summary aid. Rank-window truncation provides a mechanistic explanation for systematic divergence between whole-distribution methods and top-window truncated-rank methods on scRNA-seq data. We adopt a three-family taxonomy—cross-sample standardized magnitude (z-score), rank-enrichment hybrids (ssGSEA, GSVA), and top-window truncated rank (AUCell, UCell)—which is more accurate than a two-way magnitude-versus-rank split, since ssGSEA is itself a rank-based enrichment statistic and should guide method selection, particularly for fibrotic, inflammatory, or otherwise broadly remodeled transcriptomes. PathwayBench is delivered as a versioned, extensible benchmark with per-dataset scores, an interactive advisor application, and complete reproducibility infrastructure to support evidence-based method selection and community extension.
Accurate modelling of allele frequency trajectories requires incorporation of both genetic mechanisms such as selection and mutation, as well as realistic population demography. Accounting for a non-constant demography within a Wright–Fisher diffusion framework induces a time-inhomogenous drift coefficient, a regime falling outside the scope of existing exact simulation routines. To address this gap, we introduce EWF 2.0, an exact simulation algorithm that accommodates time-varying demography within Wright–Fisher diffusions whilst retaining all the functionality of previous EWF versions. We validate correctness using distributional tests (Kolmogorov–Smirnov, QQ plots), confirming agreement with theoretical expectations. In spite of its greater generality, EWF 2.0 retains the same runtime as in previous versions, ensuring computational efficiency and scalability. All software is available at https://github.com/JaroSant/EWF. EWF 2.0 is particularly valuable for bridge simulation, where existing methods cannot handle time-varying mutation and selection rates. For a specified demographic history, mutation parameters, selection function and sampling times, EWF 2.0 generates exact draws from the law of the corresponding Wright–Fisher diffusion or diffusion bridge.
Immune checkpoint therapy (ICT) targeting PD-1/PD-L1 fails in nearly half of lung adenocarcinoma (LUAD) patients. Tumor-cell-intrinsic FOXP3 directly regulates CD274 (PD-L1) transcription. I hypothesized that gain-of-function (GOF) mutant p53 (mutp53) physically sequesters FOXP3, preventing CD274 promoter occupancy and enabling unchecked PD-L1 expression in a lineage-specific manner. Six-layer multi-scale computational analysis: (1) clinical survival meta-analysis across three independent LUAD cohorts (TCGA PanCancer Atlas n = 510, OncoSG n = 181, CPTAC n = 110; N = 670); (2) FIMO transcription factor occupancy analysis (JASPAR 2024 MA0850.1) on the CD274 promoter; (3) ATAC-seq chromatin accessibility profiling across LUAD/HNSC lineages; (4) AlphaFold 3 mutp53-FOXP3 complex prediction with PAE analysis; (5) AMBER ff19SB/TIP3P molecular dynamics (2 ns NPT, 310 K, 0.15 M NaCl, A100 GPU); (6) pan-cancer transcriptomic analysis (8 TCGA cohorts, N = 4,205; DESeq2/Spearman) from preprint [36]. Disrupted FOXP3-PD-L1 axis significantly predicted inferior OS (log-rank p = 6 × 10−4,N = 670; HR = 1.48, 95
The development of RNA therapeutics requires formulations that achieve efficient delivery while limiting unwanted immune activation and off-target effects. Computational approaches can help explore these competing objectives before experimental testing, but mechanistic interpretation, formulation prioritization, and biological consistency assessment are often addressed separately. We developed a modular computational framework with two complementary analytical branches. The candidate-prioritization branch combines a heterogeneous virtual-patient cohort, a 500-formulation RNA–lipid nanoparticle (RNA-LNP) design space, five synthetic efficacy- and safety-related endpoints, a composite Safety-by-Design score, a Random Forest surrogate, and Pareto-based multi-objective analysis. In the biological-consistency branch, descriptors derived from previously published Universal Immune System Simulator (UISS) trajectories were condensed into five Mechanistically Informed Immune Signatures (MIIS), used to constrain synthetic multi-omics profiles, and compared at the immune-program level with the independent human single-cell RNA-sequencing dataset GSE171964. MIIS and synthetic multi-omics variables were not used as predictors, endpoint inputs, or Pareto objectives. In formulation-grouped five-fold cross-validation, the Random Forest reproduced the composite score within the synthetic design space (pooled out-of-fold R² = 0.887; RMSE = 0.055; MAE = 0.044). The primary Pareto analysis identified 21 non-dominated formulations. Ranking was generally robust across 3,876 alternative Safety-by-Design weight combinations, and a complete balanced analysis of all 100,000 patient–formulation pairs produced a ranking closely matching the primary analysis (Spearman ρ = 0.998; top-20 overlap 18/20). LNP_0462 ranked first globally and within the Low-, Intermediate-, and High-inflammatory-risk strata. Mean BPCI across seven post-vaccination days was 0.760 and exceeded the exact shuffled-label null (p = 0.025), although day-specific evidence and donor-level uncertainty were heterogeneous. This proof-of-concept framework supports transparent exploration of efficacy–safety trade-offs and prioritization of RNA-LNP formulations within a fully synthetic design space. The additional sensitivity analyses indicate that the leading candidates are robust to alternative score weights, sampling balance, and inflammatory-risk stratification. The MIIS/BPCI branch provides a separate assessment of immune-program organization rather than validation of individual formulations. Experimental testing and evaluation in additional human datasets are required before drawing conclusions about biological or clinical performance.
Protein contacts often serve as prerequisites for characterizing molecular interactions and bonds. Traditional implementations remain computationally expensive, posing significant scalability and usability challenges. We present COCaDA-web, an interactive and user-friendly Web Server that allows users to dynamically explore and visualize protein contact data, building upon the COCaDA (COntact search pruning by C α Distance Analysis) command-line tool, which is aimed for large-scale protein interatomic contact detection. Users can submit their own queries, or explore precomputed data for over 240,000 proteins in the Protein Data Bank. COCaDA-web includes novel features like pH value customization for protonation-aware analysis and biological assembly support, and takes less than a second to process proteins up to 1000 residues. Detailed results are provided for each entry in a dynamic table, alongside an interactive 3D visualization, annotated PyMOL sessions, and step-by-step usage guides. We demonstrate our tool by performing two comparative analyses of the interatomic contacts of Transthyretin and 5-Hydroxyisourate Hydrolase, two evolutionarily and structurally related proteins yet known to perform distinct biological functions associated with amyloidogenesis and gout. COCaDA-web reveals important differences, especially in their central cavities, that help explain the functional divergence of these proteins. At the moment, the COCaDA-web database contains approximately 746 million contacts, divided into 690 million intra-chain and 55 million inter-chain. COCaDA-web provides a fast, interactive, and scalable platform for protein interatomic contact analysis, combining efficient large-scale contact detection with intuitive visualization and exploration tools. COCaDA-web is freely accessible at https://bioinfo.dcc.ufmg.br/cocada-web, and there are no login requirements.
Machine learning models for toxicity prediction are routinely evaluated using random train/test splits, allowing structurally similar compounds to appear in both sets and inflating reported performance metrics. A rigorous benchmark quantifying this overestimation alongside calibration, uncertainty, and applicability domain analyses is needed to guide practitioners in selecting and trusting toxicity prediction models. We present ToxBench, a leakage-audited benchmark comprising three toxicology datasets (Tox21, 7,538 compounds, 12 tasks; ClinTox, 1,379 compounds, 2 tasks; SIDER, 1,350 compounds, 27 tasks) processed through a transparent standardization pipeline with explicit reporting of conflicting-label removals. Four model classes were evaluated: Random Forest, XGBoost, MLP, and Graph Neural Network (GNN), each trained under random and Bemis-Murcko scaffold-based splits across five independent seeds (120 experimental conditions). Analyses included post-hoc probability calibration, ensemble-based uncertainty quantification, nearest-neighbor applicability domain analysis, and scaffold-level error analysis. Scaffold splitting consistently reduced AUROC by 0.057–0.079 points across all model classes on Tox21 (mean drop: 0.070) and by 0.031–0.035 points on SIDER for three of four models, demonstrating systematic performance overestimation under random splitting. ClinTox showed reversed performance ordering due to small dataset size and extreme class imbalance, with high seed-to-seed variance (± 0.085–0.160) confirming results are dominated by sampling noise. Post-hoc calibration reduced ECE by 67–68
Channelrhodopsin variants with desired photocurrent properties are commonly developed through iterative experimental screening. Computational models are increasingly being used to prioritize candidate variants, but many existing approaches rely primarily on sequence-derived features. Because protein properties arise from physicochemical interactions in three-dimensional space, structure-derived representations may offer a useful perspective for modeling and interpreting variant–property relationships. We therefore developed Foldinsight, a structure-derived framework for modeling quantitative properties of channelrhodopsin variants and visualizing property-associated spatial regions. Foldinsight converts amino acid sequences into AlphaFold2-predicted structures, aligns the structures in a common coordinate system, and calculates van der Waals and electrostatic molecular fields on a shared three-dimensional grid. These fields provide fixed-length descriptors suitable for regression modeling. When applied to a published channelrhodopsin variant dataset, the molecular-field descriptors captured predictive signals across the measured photocurrent properties. Mapping the fitted regression coefficients back onto the molecular-field grid highlighted candidate spatial regions associated with the modeled properties and enabled their relationship to established functional and structural features to be examined. Foldinsight provides a workflow for connecting predicted protein structures with quantitative property modeling and spatial interpretation. AlphaFold-derived molecular fields offer an alternative spatial representation for channelrhodopsin variant analysis and may help generate hypotheses for future mutational experiments.
Accurate classification of enzymes and non-enzymes from protein sequences is fundamental to understanding plant metabolism, with applications in crop improvement and biotechnology. While deep learning has shown promise for protein function prediction, most approaches train from raw sequences requiring extensive computational resources. Transfer learning using pre-trained protein embeddings offers an efficient alternative, yet its application to plant enzyme classification across multiple species remains unexplored. We present a transfer learning framework combining UniProt-derived protein embeddings with attention-enhanced and baseline deep neural networks for enzyme classification in four major plant species: Arabidopsis thaliana, Brassica species, Oryza sativa, and Triticum aestivum. Our curated dataset comprises 22,267 unique protein sequences with 1024-dimensional embeddings. To prevent data leakage, we implemented homology-aware splitting using CD-HIT clustering at 60
Accurate classification of protein–protein interaction (PPI) categories is critical for elucidating intracellular signaling pathways and disease mechanisms. Reliable computational prediction methods can substantially reduce the cost of high-throughput wet-lab screening. However, existing multimodal PPI predictors generalize poorly, primarily for two reasons. First, most multi-category PPI predictors integrate only one or two of the three key feature modalities: evolutionary sequence signatures, three-dimensional (3D) structural profiles, and Gene Ontology (GO) functional annotations. This incomplete modal coverage prevents them from building comprehensive protein representations. Second, even models that cover all three modalities typically rely on shallow fusion strategies, which fail to capture the deep complementary relationships among orthogonal biological signals. This limitation leads to performance degradation on low-homology proteins and rare PPI categories. To address these limitations, we propose Tri-modal Chained Cross-Attention Protein–Protein Interaction (TriCCA-PPI), a framework with two key contributions. First, a fully decoupled feature extraction pipeline independently generates ESM-2 sequence, ESM-IF1 structural, and GO-anc2vec functional embeddings. This modular architecture supports independent replacement and upgrading of each modality’s feature extractor, and compensates for the incomplete biological characterization inherent in single- or dual-modal inputs. Second, a chained pairwise cross-attention module performs three rounds of progressive modal alignment to capture layered cross-modal complementary relationships. We further employ a global–local dual-channel Graph Isomorphism Network (GIN) with Jumping Knowledge aggregation to enhance graph topological representation. Asymmetric loss (ASL) is applied to mitigate class imbalance. Evaluations on the SHS27K and SHS148K datasets across Random, BFS, and DFS splits show that TriCCA-PPI outperforms state-of-the-art methods, with substantial gains observed on low-homology proteins and rare PPI categories. Ablation experiments validate the contribution of each modality and the advantage of chained cross-attention over shallow fusion. TriCCA-PPI addresses two key limitations of existing multimodal PPI predictors. Its decoupled, replaceable feature extraction pipeline integrates multi-dimensional biological cues into comprehensive protein representations, while chained cross-attention enables deep progressive fusion that leverages cross-modal complementary relationships. The method improves generalization for low-homology proteins and rare PPI categories, and its modular design offers a practical, interpretable, and extensible multimodal fusion framework for large-scale multi-category PPI prediction.
Accurate phasing of genomic sequences is necessary for knowing genetic variation and its role in human genomics. Formal phasing methods, such as Mendel Impute, Eagle, and Beagle, require large reference panels, which curtail scalability and limit applicability to undersampled populations. To this, we propose RefFree-Phaser, a reference-free deep learning model for genotype phasing assembled on a BigBird transformer architecture. The BigBird model architecture, founded on sparse attention, is suitable for genomic sequencing as it is computationally less expensive than full-attention-based transformers while modeling long-range dependencies. The RefFree-Phaser’s framework operates with positional embeddings and tokenizes input sequences, processes them through a 12-layer BigBird transformer with sparse attention, and derives contextualized hidden states that are mapped into final hidden state, logits, confidence scores, and binary haplotype predictions aligned with unphased genotypes and true labels. We adopted a lexicographical haplotype-pairing strategy, in which all possible haplotypes were sorted and systematically tagged according to their input genotypes. We evaluated RefFree-Phaser across two cohorts, one from the 1000 Genomes Project subset and the other from Omni2.5 M common-variants dataset. Experiments across diverse populations showed strong performance for European (EUR) and other superpopulations, while accuracy was slightly lower for African (AFR) individuals, indicating higher phase-switch, genetic diversity, and limited representation in training data. RefFree-Phaser attains average accuracies of 92.60
Lysine acetylation is a pervasive post-translational modification with critical regulatory roles, yet computational prediction of acetylation sites remains hampered by unreliable benchmarking practices including dataset redundancy and non-independent test set evaluation. We report a systematic investigation demonstrating that three independent sources of metric inflation, including near-ubiquitous sequence redundancy, non-independent test sets, and distribution-locked dimensionality reduction collectively reduced an apparent accuracy of 99.38
Transcriptomics arrived with the prospect of mechanistic deduction and quantification of active molecular pathways, providing insights into both state and function of cells. However, it remains a challenging task to interpret functional analysis of single samples of RNA-sequencing data, because of inconsistencies between methods, and inherent reliance on a priori gene signatures encoding known biological knowledge. Initially, methods were designed for experimental setups with known conditions and biological replicates, but for many use-cases, particularly in clinical diagnostics, the problem is N-of-1, where each single sample presents a unique and independent case, without replicates or baseline for comparison. Here, we are comparing the performance of 17 single-sample GSEA methods on a diverse set of datasets, with known biology, and test them on a matched collection of gene signatures, with relevance to both experimental research and clinical diagnostics. We find that truly single-sample methods perform well and that the methods that we adapted to work on individual samples, by introducing a large data collection as a universal reference, are generally able to correctly classify different samples of known phenotypes. Single-sample GSEA remains a challenging task where the tool, query data, gene signature, and potential baseline influence the results. Here, we have systematically assessed a range of tools using a designed truth set of data and gene sets. Z-score-based and truly single-sample rank-based methods were consistently competitive, but method choice must be validated for each dataset–gene set context. The benchmark is implemented in the OmniBenchmark framework, inviting continuous additions of relevant tools and datasets.
Abstract Background Protein complexes constitute fundamental functional modules within cells and play a crucial role in regulating many biological processes. Detecting protein complexes from protein-protein interaction (PPI) networks has therefore become a central problem in computational systems biology. However, many existing computational approaches struggle to accurately identify overlapping complexes where proteins participate in multiple functional modules simultaneously. In addition, large-scale PPI networks are inherently noisy and incomplete due to experimental limitations, which significantly affects the reliability of complex detection methods. Results In this work, we propose TOMOC, a topology-driven multi-objective evolutionary framework for robust detection of overlapping protein complexes in noisy PPI networks. The novelty of TOMOC lies in the integration of a topology-driven bi-objective formulation, an edge-based evolutionary representation that naturally supports overlapping memberships, and a topology-aware structural refinement mechanism within a unified framework for protein complex detection in noisy PPI networks. The proposed framework introduces an edge-based evolutionary representation that models candidate solutions at the interaction level, allowing overlapping memberships to emerge naturally during decoding. It further optimizes two complementary structural objectives by minimizing average conductance and triangle-density loss, enabling the algorithm to balance boundary quality and internal structural density. In addition, a topology-aware structural overlap refinement (SOR) operator is designed to improve structural coherence and robustness against noisy interactions through boundary-aware repair, triangle-closure expansion, and triangle-support pruning. Extensive experiments conducted on three benchmark PPI networks (Yeast-D1, Yeast-D2, and Collins) demonstrate that TOMOC achieves competitive performance compared with several state-of-the-art methods in terms of precision, recall, and F1-score. Conclusions The proposed TOMOC framework provides an effective and scalable topology-driven approach for detecting overlapping protein complexes directly from PPI network topology. By integrating multi-objective evolutionary optimization with topology-aware refinement mechanisms, TOMOC effectively captures the structural characteristics of protein complexes and demonstrates strong robustness when applied to large and noisy biological interaction networks.
MicroRNAs (miRNAs) are a class of small noncoding RNAs that inhibit the translation of target messenger RNAs (mRNAs). Given that a single miRNA can regulate the translation of many mRNAs, miRNAs have emerged as critical regulators of physiological processes. MiRNAs have been linked to the development and progression of cancers, neurodegenerative and other diseases, most recently using high-throughput miRNA “miRNome” sequencing. As miRNome sequencing represents a newer ‘omics application, limited guidance is available for how to analyze this data. Existing interfaces that enable non-computational users to interpret and perform comprehensive secondary analysis on their own miRNome data are limited in functionality and/or interactivity. Therefore, we developed MiRQuery to address this need. MiRQuery is an RShiny application which features common visualization methods for high-throughput sequencing data, such as multidimensional scaling, stacked column charts, heatmaps, and boxplots to compare expression across groups for a user-specified miRNA of interest. MiRQuery further provides support for differential miRNA and gene expression analysis. Unique to miRNome sequencing data analysis, users may retrieve predicted gene targets of differentially expressed miRNA and follow up with pathway overrepresentation analysis of the gene targets. Finally, if users upload paired bulk mRNA sequencing data, they may identify differentially expressed genes and negatively correlated miRNA-gene pairs. By providing access to sophisticated bioinformatics tools through a user-friendly interface, MiRQuery empowers both scientists new to bioinformatics and bioinformaticians new to the field to extract insights rapidly and reproducibly from their sequencing data. MiRQuery can be accessed through PositConnect at https://julianneyang-mirquery.share.connect.posit.cloud/ , and alternatively is available by user local installation via instructions on the Github project homepage.
Codon optimization is a routine yet high-impact step in de novo gene design for heterologous expression and serves as an important tool in synthetic biology. Protein expression output depends on multiple sequence-level determinants, such as global codon-usage, codon-pair context, initial mRNA folding, and motif/repeat content. However, many existing tools still provide limited transparency or restricted design flexibility. We developed Codon Design Online (CoDOn), a publicly accessible web application that formulates codon optimization as an explicit multi-objective design task and uses the non-dominated sorting genetic algorithm II (NSGA-II) to generate diverse Pareto-optimal coding sequences. CoDOn supports various design criteria that capture both host-level codon usage and local translation context, including individual codon usage (ICU), codon context (CC), codon adaptation index (CAI) and hidden stop codons. Users can also impose sequence-level constraints, e.g., GC/GC3 composition targets, 5′-proximal mRNA folding energy, and motif/repeat exclusion. Importantly, the platform features a modular interface that enables intuitive, interactive parameter configuration and result exploration through Pareto plots, codon-usage radar charts and tables, codon-pair heatmaps, and exportable reports. We showcase the utility of the platform using representative design cases in E. coli and S. cerevisiae under realistic practical constraints. CoDOn turns codon optimization from a black-box, single-output procedure into a transparent, user-configurable decision process based on Pareto trade-offs, supporting customizable gene design across hosts and applications. The CoDOn web service is available at https://codondesign.com/.
Predicting a protein’s binding sites helps to understand the functional mechanisms of protein interactions, which provides insights into drug discovery. Although experimentally determining the protein complex’s structure can accurately identify the binding residues, the process is labor-intensive and expensive. With the recent advances in the protein language model (PLM) and geometric deep learning, we introduce S2Site (Sequence and Structure based Binding Site Prediction), an end-to-end framework that incorporates the geometric deep learning model with the PLM to identify the protein binding sites. S2Site consistently outperforms various state-of-the-art methods in the protein binding site prediction of three different interactions, including protein-protein, antigen-antibody, and protein-peptide binding sites. Compared to methods based on multiple sequence alignments, PLM allows S2Site to predict protein binding sites on a large scale efficiently. Our experiments also show that both sequence and structural features contribute to the performance of binding site prediction. Overall, S2Site is a robust and practical model for efficiently identifying binding residues of various protein-ligand interactions. The source code and model can be accessed at https://github.com/LW-21/S2Site.
Enhancers are distal cis-regulatory elements that play essential roles in gene regulation, development, and disease. Although numerous computational methods have been developed for enhancer identification, commonly used DNA encoding strategies, such as one-hot and k-mer representations, typically treat DNA as a linear symbolic sequence. This simplification ignores the intrinsic geometry of the DNA double helix, where nucleotides are organized along a periodic three-dimensional structure with a helical repeat of approximately 10–11 base pairs. To address this limitation, we propose HelixPR (Helix-aware Periodic Representation), a biologically motivated DNA sequence representation for enhancer identification. HelixPR augments per-nucleotide one-hot encoding with a sinusoidal periodic encoding using a period of T=10 base pairs. It further incorporates a global nucleotide-composition token to capture sequence-level compositional information complementary to local nucleotide patterns. Together, these components form a compact, interpretable, and information-preserving representation of DNA sequences. Combined with a lightweight one-dimensional convolutional neural network, HelixPR achieved competitive performance across benchmark datasets. On Liu’s benchmark, HelixPR achieved an accuracy of 0.8434 and a Matthews correlation coefficient of 0.6874 for enhancer identification. For enhancer strength classification, HelixPR achieved an accuracy of 0.8400 and an AUC of 0.9346, outperforming baseline methods. On Basith’s benchmark, HelixPR obtained a mean balanced accuracy of 0.9778, substantially outperforming existing methods. Further analysis showed that HelixPR supports closed-form decoding and that HelixPR-highlighted regions are enriched for known transcription-factor binding motifs, supporting both the interpretability and biological relevance of the proposed representation. These results suggest that incorporating the periodic structure of the DNA double helix can improve enhancer modeling. By combining nucleotide identity, helical periodicity, and global compositional information, HelixPR offers a biologically grounded framework for enhancer identification and related regulatory sequence prediction tasks.
The Gene Ontology (GO) is a public resource that describes gene functions and characteristics through a structured vocabulary of standardised terms. It currently contains annotations for over 1.5 million gene products, each linked to one or more GO terms. In this study, we propose integrating this GO-based semantic structure into machine learning systems for medical diagnostics. This approach serves a dual purpose: first, to prioritise genes that are semantically relevant to a given clinical task, thereby refining model input; and second, to enable the analysis of biologically predefined gene sets, which may reveal novel mechanisms underlying disease. Evaluated across 16 benchmark data sets spanning diverse medical domains, our GO term-informed gene selection method generally outperformed models trained on full gene sets. Further analysis of individual GO terms not only enhanced classification performance but also identified high-performing, task-specific gene subsets that were overlooked during initial gene selection. Our findings demonstrate that Gene Ontology can be effectively leveraged for semantics-aware gene selection in clinical machine learning. Moreover, systematically evaluating individual GO terms offers a scalable strategy to uncover new, testable biological hypotheses–revealing gene functions that might otherwise remain hidden when examining only broadly selected gene combinations.