Domain-specific large language models (LLMs) can improve access to complex bioinformatics resources, but their development is often limited by fragmented source material, inconsistent data preparation, evaluation leakage, and incomplete reporting of training decisions. We present a reproducible nine-step protocol for converting heterogeneous bioinformatics knowledge sources into quality-controlled question–answer (QA) datasets and parameter-efficiently fine-tuned language models. The workflow covers task definition, source registration, content acquisition, text canonicalisation, source-grounded question–answer generation, filtering, semantic deduplication, group-preserving data partitioning, low-rank adaptation (LoRA)-based fine-tuning, and comparative evaluation against zero-shot, retrieval-augmented, and four-bit deployment baselines. We demonstrate the protocol using PRSGPT, a domain-specific assistant derived from resources covering 32 polygenic risk score (PRS) tools, and BioStarsGPT, a corresponding assistant derived from BioStars community discussions, yielding 28,378 and 154,282 retained QA pairs, respectively. In zero-shot evaluation, Qwen2.5-7B achieved the highest Recall-Oriented Understudy for Gisting Evaluation (ROUGE) scores for PRSGPT, while Mistral-7B achieved the highest or joint-highest performance on several other measures. Qwen2.5-7B was selected for the principal fine-tuned and deployment comparisons in both applications. For PRSGPT, the fine-tuned model exceeded QA-pair retrieval on eight of 13 non-exact-match metrics. Descriptive BioStarsGPT comparisons indicated that retrieval remained competitive in the more heterogeneous forum-derived setting. On an independent 176-question benchmark, PRSGPT answered 109 questions correctly (61.9%), compared with 108 for Google Gemini (61.4%) and 126 for Claude (71.6%). Four-bit deployment caused substantial, dataset-dependent performance loss. The protocol emphasises provenance, leakage-aware evaluation, independent factual assessment, and transparent reporting, with code, datasets, models, prompts, and evaluation outputs openly available.
Abstract Endosomal trafficking is a major pathway that delivers cell-surface proteins, including glutamate receptors, to support neurotransmission and normal brain functions. Activity-dependent insertion of glutamate receptors is essential for synaptic plasticity, learning and memory. Copine-6 is a neuronal-specific calcium (Ca 2+ ) binding protein that mediates activity-induced exocytosis of α-amino-3-hydroxy-5-methyl-4-isoxazole propionic acid (AMPA)-type glutamate receptors. The activation of N -methyl- D -aspartate (NMDA) receptors triggers Ca 2+ -dependent translocation of Copine-6 to intracellular endosomal compartments. However, the mechanisms underlying the activity-dependent accumulation of Copine-6 in endosomes remain unknown. Here, we show that Copine-6 exhibits Ca 2+ -dependent binding to phosphatidylinositol-3-phosphate (PI(3)P) through the C2B domain and displays enhanced interaction with active Rab11a in a Ca 2+ -independent manner via the vWA domain. Mutations in the C2B that inhibit binding to PI(3)P not only block the activity-induced translocation of Copine-6 to early endosomes, but it also causes an aberrant accumulation of Copine-6 in recycling endosomes. Consequently, loss of Copine-6 expression impairs the efficient coupling of early and recycling endosomes and blocks activity-dependent delivery of both AMPA and NMDA receptors onto the neuronal plasma membrane. These defects can be restored by re-expressing wild-type Copine-6, but not the C2B phospholipid-binding mutant. Together, our findings establish Copine-6 as a molecular bridge that enhances coupling between the early and recycling endosomal membranes, thereby facilitating the activity-dependent forward trafficking of glutamate receptors to the neuronal plasma membrane to maintain synaptic potentiation. Significance Statement Activity-dependent trafficking of glutamate receptors is essential for synaptic plasticity, learning, and memory, but the mechanisms coordinating receptor transport through endosomal compartments remain unclear. This study identifies the neuronal calcium-binding protein Copine-6 as a molecular bridge that couples early and recycling endosomes during receptor trafficking. Copine-6 integrates calcium-dependent phospholipid binding and Rab11-mediated endosomal interactions to promote activity-dependent delivery of both AMPA and NMDA receptors to the neuronal surface. Loss of Copine-6 or disruption of its phospholipid-binding activity uncouples endosomal trafficking and impairs receptor insertion during synaptic potentiation. These findings reveal a key mechanism linking neuronal activity, endosomal organization, and glutamate receptor delivery to support synaptic function.
Abstract Missense mutations in TERT , the gene encoding the human telomerase catalytic subunit hTERT, are associated with Telomere Biology Disorders (TBDs). Experimentally elucidating the effects of all possible missense variants would be time-consuming and technically challenging. Moreover, current computational predictors are not hTERT-specific and primarily rely on sequence information, failing to capture the complex biological and structural context of the telomerase enzyme. In this work, we developed three machine learning models integrating both sequence- and structure-based features to account for the biological mechanisms of hTERT. Compared to state-of-the-art methods, our best-performing models achieved a higher Matthew’s Correlation Coefficient of 0.88 on ClinVar and gnomAD curated variants and demonstrated robust sensitivity (0.75) on a dataset curated according to guidelines from the American College of Medical Genetics and Genomics and Association for Molecular Pathology (ACMG/AMP). Feature interpretation highlighted hTERT residue conservation and changes in hydrophobic and weak polar interactions as critical determinants of pathogenicity. Finally, in silico saturation mutagenesis was performed to present a mutational landscape of TERT , available in a user-friendly web server, CharacTERT, which could offer valuable insights into the molecular mechanisms driving TBDs, aid in early diagnosis, as well as guide personalized treatment strategies. CharacTERT is freely available at https://biosig.lab.uq.edu.au/charactert/ .
The spectrum of congenital malformations in VACTERL association varies among patients and can be differentially diagnosed with CHARGE syndrome, Fanconi anaemia, and others (reviewed in Solomon 2011). Despite overlapping clinical findings, the genetic causes of these diseases are distinct. In this context, unbiased whole genome sequencing can assist in differential diagnoses, as well as identify new gene-disease associations. In this report, we demonstrate that whole genome sequencing of a proband with suspected VACTERL association revealed two gene variants in the ribosomal genes RPL18 (NM_000979.4 c.397 G > C p.(Gly133Arg)) and RPS6 (NM_001010.3: c.370 C > G p.(Leu124Val)). Mutations in ribosomal genes are associated with Diamond-Blackfan anaemia, a condition that shares phenotypic similarities with Fanconi anaemia. Our modelling and functional assessment of the identified variants strongly indicate pathogenicity of the RPL18:p.(G133R) variant as it is novel, displays reduced expression and stability and abnormal intracellular distribution, and interferes with protein synthesis in cultured cells. The RPS6:p.(L124V) variant reduced protein expression and altered cytoplasmic distribution but does not interfere with protein synthesis in cultured cells. Overall, our study indicates the significant advantage of using unbiased whole genome sequencing for the examination of patients with complex congenital malformations.
Predicting complex human traits from genetic data is challenging because different genetic, clinical, and molecular data sources often contain different parts of the signal. Here, we present EFGPP, a reproducible framework for generating, ranking, and combining multiple types of data for genotype-to-phenotype prediction. We applied EFGPP to migraine prediction using UK Biobank data from 733 individuals. The framework combined genotype-derived features, principal components, clinical and metabolomic covariates, and polygenic risk scores generated from migraine and depression GWAS using PLINK, PRSice-2, AnnoPred, and LDAK-GWAS. The best single data type achieved a test AUC of 0.644, while combining multiple data types improved performance to 0.688 using migraine-focused inputs and 0.663 using cross-trait depression-derived inputs. Genetic features alone did not outperform the covariates-only baseline, but genotype-derived features performed better than PRS alone, and depression-derived PRS showed useful predictive signal. Overall, EFGPP provides a practical proof-of-concept framework for prioritising and integrating heterogeneous genetic data sources for complex phenotype prediction.
Drug resistance caused by mutations is a significant global health concern. One way to better understand this phenomenon is by studying changes in protein-ligand binding affinity upon mutation. While recent advances in protein modelling, such as AlphaFold2 and AlphaFold3, have transformed structural assessments, their utility in predicting mutation-induced binding affinity changes remains underexplored. We evaluated various mutation-based methods and scoring functions using computer-generated protein-ligand complexes. Compared to a baseline using experimental structures, we observed a performance drop ranging from 5% to 30% across different computational models. Specifically, using experimental receptors with docked ligands resulted in a ~5% drop, similar to that observed with AlphaFold3 models (~5%), despite the latter offering lower ligand root mean square deviation. However, using AlphaFold2 receptors with docking led to a greater performance loss (10%-20%), comparable to homology models with high sequence identity. Homology models based on low-identity templates showed over 30% decline. These performance differences were most pronounced for interface mutations and low molecular weight ligands. While AlphaFold models offer accurate protein and interaction predictions, they lack mutation-specific information, such as dynamic changes, highlighting the need for complementary mutation-aware methods for reliable analysis. Our findings provide insights into interpreting mutation effects on ligand binding using predicted structures and can guide more robust assessments of drug resistance mechanisms in silico.
Cancer progression is driven by epigenetic reprogramming, where promoter hypermethylation of tumour-suppressor genes and global hypomethylation reshape gene regulation and cellular phenotypes, promoting oncogenesis and disease advancement. We previously introduced the Methylscape, a cancer-specific DNA methylation landscape characterized by clustered promoter hypermethylation and gene body hypomethylation that enhances DNA's physical affinity for gold surfaces. Here, we demonstrate that Methylscape can be leveraged to monitor cancer progression. In a TGF-β-induced breast cancer epithelial-mesenchymal transition (EMT) model, we observe increased Methylscape enrichment of mesenchymal-state DNA, indicating that this method can sensitively detect subtle epigenetic remodelling linked to tumour progression. Using a gold-based DNA desorption enrichment strategy coupled with methylation sequencing and qPCR, we show that hypermethylated regions are preferentially enriched on gold surface. Finally, we developed a low-cost, disposable screen-printed electrode platform for stage-specific breast cancer monitoring. Together, these findings establish Methylscape as a promising biophysical biomarker for non-invasive, real-time monitoring of cancer progression, advancing its potential for clinical translation.
Metal ions play critical structural, regulatory, and enzymatic roles in proteins, making their binding essential for biological processes. Experimental identification of metal-binding sites is resource-intensive and limited in scalability. Recent advances in protein language models have transformed computational predictions, yet current tools do not address how residue-level metal-binding probabilities change upon mutation. To fill this gap, mCSM-metal leverages embeddings from ESMBind with our graph-based structural signatures to accurately predict the effects of single or multiple point mutations on the binding of seven essential ions (Zn2+, Ca2+, Mg2+, Mn2+, Fe3+, Co2+, Cu2+). Our model achieves accuracies, F1-scores, and Matthews Correlation Coefficient values up to 0.97, 0.97, and 0.95, outperforming other approaches. The webserver provides an interactive platform to assess and visualise local and long-range impacts of mutations on metal-ion binding, offering new avenues for applications in structural biology, disease modelling, and protein engineering. The web application is freely available at: https://biosig.lab.uq.edu.au/mcsm_metal/.
Workflow of proposed hepatotoxicity prediction model Drug-induced liver injury (DILI) continues to be one of the main reasons for late-stage drug attrition, post-marketing withdrawals, and acute liver failure are caused by the DILI phenomenon which is characterized by being multifactorial and idiosyncratic. The availability of in silico hepatotoxicity prediction models has primarily relied on either two-dimensional descriptors or graph topology, thus failing to represent the interplay of molecular geometry, physicochemical determinants, and biologically mediated toxicity pathways in the predictions. This study describes a mechanistically informed multimodal deep-learning framework that combines threedimensional molecular geometry, physicochemical profiling, structural alerts, and biologically derived risk signals for the robust prediction of hepatotoxicity. The proposed structure uses a geometry-aware SchNet encoder that directly learns the spatial and electronic interactions from the atomic coordinates, which is further enhanced by the biophysicochemical expert network that incorporates RDKit descriptors, MACCS fingerprints, and validated biological oracles that capture key hepatotoxicity-relevant mechanisms, such as transporter inhibition and metabolic liability. The heterogeneous representations are combined by a late-integration technique, which enables the retention of domain-specific signals prior to categorization. The model was trained and evaluated using a carefully curated hepatotoxicity dataset prepared from DILIrank 2.0 and TOXRIC. The multimodal framework surpassed traditional machine learning models and classic graph neural networks by its performance on an independent test set attaining an AUC-ROC of 0.80, accuracy of 82.8%, and a Matthews correlation coefficient of 0.665. In addition, the model showed high sensitivity (92.1%) for hepatotoxic compounds and lost little predictive power under scaffold-split evaluation, implying partial generalization beyond the known chemical series. The interpretation analyses verified the model's compliance with established toxicology and showed the contributions derived from known reactive substructures, physicochemical risk factors, and biologically contextualized signals.
Genotype-to-phenotype prediction is a central goal of statistical genetics, yet practical comparisons of prediction workflows remain limited in small, heterogeneous, participant-shared genomic datasets. Here, we benchmarked end-to-end case-control prediction across 80 curated binary phenotypes from openSNP using machine learning, deep learning, and polygenic score workflows. We evaluated 29 machine-learning algorithms, 80 deep-learning model variants, and 3 polygenic score tools across 675 clumping and pruning configurations. No workflow family dominated universally. Polygenic score workflows achieved the highest observed discrimination for 53 phenotypes, whereas machine-learning or deep-learning workflows achieved the highest for 27. However, many apparent phenotype-level wins were modest, with 41.2% of comparisons representing practical ties within five discrimination points. Performance was strongly phenotype-dependent and sensitive to modeling and preprocessing choices. Distinct workflow-specific failure modes were also observed, including unstable behaviour in PRSice and non-informative collapse in lassosum for 13 phenotypes. Higher peak performance was concentrated in smaller phenotypes, reinforcing the need for cautious interpretation in limited-data settings. The cohort was predominantly of European ancestry, restricting generalisability. Together, these results position openSNP as a useful stress-test environment for genomic prediction and support benchmark-guided workflow selection under realistic conditions of data scarcity, phenotype heterogeneity, and ancestry imbalance.
DPD-Cancer is a graph-attention deep learning framework for predicting small-molecule DPD-Cancer is a graph-attention deep learning framework for predicting small-molecule anti-cancer activity across the NCI-60 panel, trained and evaluated under a strict chemistry-aware data-partitioning scheme. On the hold-out test set, the classifier achieved an Area Under the Receiver Operating Characteristic Curve (AUROC) of 0.87 (95
Objective: SNP heritability estimates vary substantially across estimation strategies, yet the downstream consequences for polygenic risk score (PRS) construction remain poorly characterised. We systematically benchmarked heritability estimation configurations and assessed their propagation into downstream PRS performance. Methods: We benchmarked 86 heritability-estimation configurations spanning six tool families (GEMMA, GCTA, LDAK, DPR, LDSC, and SumHer) and ten method groups across 10 UK Biobank phenotypes, yielding 844 configuration-level estimates. Each estimate was propagated into GCTA-SBLUP and LDpred2-lassosum2 PRS frameworks and evaluated across five cross-validation folds using null, PRS-only, and full models. Eleven binary analytical contrasts were tested using Mann-Whitney U tests to identify drivers of heritability variability. Results: Heritability ranged from -0.862 to 2.735 (mean = 0.134, SD = 0.284), with 133 of 844 estimates (15.8 Conclusion: SNP heritability is best interpreted as a configuration-sensitive modelling parameter rather than a universally stable scalar input. Heritability estimates should always be reported alongside their full estimation specification, and downstream PRS performance is comparatively robust to moderate variation in the heritability input.
Protein structure-based comparison provides a framework for uncovering deep evolutionary relationships that can escape conventional sequence-based approaches. Encoding three-dimensional protein structures using a simplified structure-aware alphabet can lead to compact, comparable strings that retain key spatial relationships. Although this enables comparison, structure-aware alignments can experience misaligned regions, particularly when comparing proteins with substantial divergence in fold architecture. To address this, a web-based resource, Structome-AlignViewer, is introduced in this work for evaluating the quality of structure-aware alignments through both spatial mapping of alignment columns to protein structures and quantitative confidence scoring. Confidence is computed from pairwise structural substitutions between adjacent inputs and normalized within each alignment to highlight relatively well-supported columns. To provide broader context, thousands of alignments from established structural classification systems were analyzed, allowing for an empirical comparative statistic to be derived to assess alignment quality. The option to exclude gap-rich regions enables users to refine alignments and focus on conserved structural cores. This approach provides an interpretable method for assessing structural alignment quality and supports more robust comparative and evolutionary analyses. Structome-AlignViewer is freely available at https://biosig.lab.uq.edu.au/structome_alignviewer/.
Polygenic risk score (PRS) tools differ substantially in statistical assumptions, input requirements, and implementation complexity, making direct comparison difficult. We developed a harmonized, implementation-aware benchmarking framework to evaluate 46 PRS tools across seven binary UK Biobank phenotypes and one continuous trait under three model configurations: null, PRS-only, and PRS plus covariates. The framework integrates standardized preprocessing, tool-specific execution, hyperparameter exploration, and unified downstream evaluation using five-fold cross-validation on high-performance computing infrastructure. In addition to predictive performance, we assessed runtime, memory use, input dependencies, and failure modes. A Friedman test across 40 phenotype–fold combinations confirmed significant differences in tool rankings (χ^2 = 102.29, p = 2.57 × 10^-11), with no single method universally optimal. These findings provide a reproducible framework for comparative PRS evaluation and demonstrate that tool performance is shaped not only by statistical methodology but also by phenotype architecture, preprocessing choices, covariate structure, computational demands, software robustness, and practical implementation constraints.
Transformers are rapidly reshaping structural biology. We argue the reason is "Emergent Latent Biology" (ELB): transformers place proteins into high-dimensional representations where hidden biophysical patterns become easier to see. We explore this concept across four key areas: protein folding, variant effects, protein-protein and protein-drug interactions. Highlighting recent gains, we note that traditional, physics-based calculations are still required for the hardest quantitative jobs, like predicting precise binding strength. Furthermore, we draw attention to major pitfalls, arguing progress depends on solving the critical "chemistry gap," modelling chemical modifications, and the "dynamics gap", predicting protein movement, which requires better validation methods and new large-scale experiments.
Isoniazid remains a cornerstone of tuberculosis therapy, yet resistance-associated mutations in its target, InhA, are not fully explained by reduced NADH binding alone. Here, we combine X-ray crystallography, cryo-electron microscopy, calorimetry, thermal stability profiling, and multi-scale molecular dynamics to dissect ten clinically observed InhA missense variants. The structures show that many substitutions preserve the canonical cofactor-bound architecture, but thermodynamic and simulation analyses reveal divergent routes to resistance. Active-site variants remodel water and hydrophobic networks to weaken NADH binding, whereas I95P and several interface-proximal variants populate a coenzyme closed-loop state that occludes the nicotinamide/adduct-binding pocket. A190S and Y259H instead appear to preserve NADH affinity while rigidifying the Phe149–Asp150 region required to accommodate the larger INH–NAD adduct. These findings establish InhA coding-region resistance as a spectrum of local, allosteric, and adduct-selective mechanisms, providing a framework for interpreting clinical variants and designing next-generation InhA inhibitors.
Abstract Genotype-phenotype prediction plays a crucial role in identifying disease-causing single nucleotide polymorphisms and precision medicine. In this manuscript, we benchmark the performance of various machine/deep learning algorithms and polygenic risk score tools on 80 binary phenotypes extracted from the openSNP dataset. After cleaning and extraction, the genotype data for each phenotype is passed to PLINK for quality control, after which it is transformed separately for each of the considered tools/algorithms. To compute polygenic risk scores, we used the quality control measures for the test data and the genome-wide association studies summary statistic file, along with various combinations of clumping and pruning. For the machine learning algorithms, we used p-value thresholding on the training data to select the single nucleotide polymorphisms, and the resulting data was passed to the algorithm. Our results report the average 5-fold Area Under the Curve (AUC) for 29 machine learning algorithms, 80 deep learning algorithms, and 3 polygenic risk scores tools with 675 different clumping and pruning parameters. Machine learning outperformed for 44 phenotypes, while polygenic risk score tools excelled for 36 phenotypes. The results give us valuable insights into which techniques tend to perform better for certain phenotypes compared to more traditional polygenic risk scores tools.
Motivation: GWAS (genome-wide association study) summary statistic files are essential inputs for polygenic risk score (PRS) calculation. However, identifying suitable files across thousands of catalog entries typically requires downloading large datasets and manually inspecting their column structures, a process that is both time-consuming and storage-intensive. Results: We present GWASPoker, a phenotype-driven, GWAS-Catalog-specific pre-download triage tool that scans candidate GWAS files for PRS column availability through partial downloads and header detection, without requiring full-file transfer. Analysing 60,499 records from the GWAS Catalog, 60,281 (99.6 Availability and implementation: GWASPoker is implemented in Python 3 and is freely available at https://github.com/MuhammadMuneeb007/GWASPokerforPRS under the MIT licence. Example outputs and documentation are provided in the repository. The tool was tested on Linux (HPC cluster) with Python 3.8 or later. The LLM-based code-generation step is entirely optional; a rules-based column-mapping template is provided for fully offline use.
Missense variant interpretation remains challenging because pathogenicity depends on heterogeneous evidence from population frequency, evolutionary conservation, transcript context, amino acid substitution severity, prior pathogenicity predictors and protein-language-model-derived features. We present AnnotateMissense, a scalable annotation, benchmarking and genome-wide prediction framework for missense variant interpretation. AnnotateMissense integrates hg38 missense variants derived from dbNSFP v5.1 with ANNOVAR annotations, dbNSFP transcript/protein descriptors, AlphaMissense scores, ESM-derived features, conservation metrics, population-frequency variables, established pathogenicity predictors and engineered amino acid/codon-context features. Using 132,714 ClinVar-labelled missense variants, we benchmarked machine-learning and deep-learning models under controlled feature configurations. The full 303-feature benchmark set achieved the strongest performance with XGBoost, reaching mean MCC = 0.9411 and ROC-AUC = 0.9950 across stratified five-fold cross-validation. Restricted naive and location-oriented feature sets achieved lower best MCC values of 0.4989 and 0.5113, respectively. Circularity-controlled ablations showed that removing prior-predictor, population-frequency and clinically overlapping evidence reduced performance, whereas excluding AlphaMissense and ESM-derived features alone had minimal effect. Temporal ClinVar validation on newly observed pathogenic/benign variants achieved MCC = 0.7613, accuracy = 0.8798 and F1-score = 0.8750. The final model was applied to 90,643,830 hg38 missense variants to generate AnnotateMissense pathogenicity scores and binary prediction labels. Code and outputs are available at https://github.com/MuhammadMuneeb007/CAGI7_Annotate_All_Missense and https://doi.org/10.5281/zenodo.19981867.
Large language models (LLMs) often lack specialized knowledge for complex bioinformatics applications. We present a reproducible pipeline for fine-tuning LLMs on specialized bioinformatics data, demonstrated through two use cases: PRSGPT, focused on polygenic risk score (PRS) tools, and BioStarsGPT, trained on community forum discussions. The nine-step pipeline integrates diverse data sources, structured preprocessing, prompt-based question-answer (QA) generation (via Google Gemini), natural language inference (NLI) for quality control, semantic deduplication, clustering-based data splitting, and parameter-efficient fine-tuning using LoRA. We fine-tuned three LLMs (LLaMA-3.2-3B, Qwen2.5-7B, Gemma) and benchmarked them on over 14 lexical and semantic metrics. Qwen2.5-7B emerged as the best performer, with BLEU-4 and ROUGE-1 improvements of 82% and 70% for PRSGPT and 6% and 18% for BioStarsGPT, respectively. The open-source datasets produced include over 28,000 QA pairs for PRSGPT and 154,282 for BioStarsGPT. Human evaluation of PRSGPT yielded 61.9% accuracy on the PRS tools comparison task, comparable to Google Gemini (61.4%), but with richer methodological detail and accurate citations. BioStarsGPT demonstrated 59% conceptual accuracy across 142 curated bioinformatics questions. Our pipeline enables scalable, domain-specific fine-tuning of LLMs. It enables privacy-preserving, locally deployable bioinformatics assistants, explores their practical applications, and addresses the challenges, limitations, and mitigation strategies associated with their development and use.