
The rapid emergence of drug-resistant pathogens poses a critical threat to global health. With traditional antibiotics losing efficacy, antimicrobial peptides (AMPs) have gained attention for their unique mechanisms and lower resistance potential. We aimed to accelerate AMP discovery by proposing a closed-loop framework that combines AMP-Hunter (a shared-architecture discriminator for AMP classification and MIC prediction that integrates convolutional neural networks with graph neural networks), and AMP-Forge (a generator integrating multiple sequence alignment to select original candidates) and is guided by minimum inhibitory concentration (MIC)for latent space optimization and candidate selection. AMP-Hunter outperformed baseline models in both AMP classification and MIC prediction, achieving 95.82% accuracy and a 95.80% F1 score on the test set for classification, and an R2 of 0.9245 with an MAE of 0.2305 for MIC prediction. Guided by its predictions, AMP-Forge generated peptide sequences with lower MIC values and improved physicochemical properties associated with antimicrobial activity. Molecular dynamics simulations further provided in silico evidence supporting the antimicrobial potential of selected sequences by identifying stable membrane disruption and insertion behaviors consistent with membrane-targeting activity. Thus, the generation-screening-validation workflow enables reliable discovery of potent AMPs, and provides a practical strategy for rational peptide design, rapid prediction, and translational applications.
In taxonomic research, traditional phylogenetic tree- and structure-based analyses of genetic data are increasingly complemented by machine-learning-based identification and representation learning. Although the amount of DNA data needed to train state-of-the-art machine learning models often exceeds what can realistically be collected and sequenced in biological studies, the number of samples can be extended artificially through data augmentation. Genetic data augmentation usually refers to the introduction of random base variations, translocations, and reverse complementing. These augmentations do not take into account the inherent structures of populations and species, potentially blurring the lines between entities within genetic datasets. Here, we propose DNAInterpolator, an approach based on interpolation of DNA sequences within a given dataset that presents a neighbor-guided alternative to random mutations. We tested interpolation as an augmentation technique using four flowering plant datasets and an artificial neural network trained to predict genetic distances between paired samples. To address unequally distributed distances within our training datasets, we examined the effect of balancing the distance distribution by curating interpolated sequences. We found that balancing helps models capture genetic distances across the full distance range by strengthening performance in underrepresented regions of the distribution. Our new approach leverages the potential of taxonomic DNA datasets for modern machine learning applications.
The simulated assembly of molecular building blocks into functional complexes is central to computational biology and materials science. Protein-assembly simulations, driven by short-range nonpolar interactions, can in principle reach their biologically correct structures, but rugged energy landscapes often trap simulations in non-functional local minima. We introduce a long-range topological potential, quantified by weighted total persistence, and combine it with the morphometric approach to solvation free energy. Across four protein systems, this combination increases assembly success rates by up to sixteen-fold and enables assembly in cases that otherwise fail. Unlike previous topology-based approaches, our method uses topological measures as an active energetic bias rather than a descriptive tool. Depending only on atom geometry, the method extends in principle to other self-assembling systems, offering a general strategy for overcoming kinetic barriers in molecular simulations.
Protein complexes are molecular machines that execute essential cellular functions, but their computational identification remains challenging. Existing protein complex identification methods largely rely on PPI network topology, functional annotations, or protein-level biochemical evidence. Although these approaches have recovered many biologically meaningful assemblies, they are often sensitive to incomplete or noisy interactomes and provide limited mechanistic insight into the residue- and interface-level determinants of complex formation. In particular, conventional PPI-based graph representations indicate whether proteins are associated, but usually ignore how protein subunits physically interact through spatially organized residues and structural interfaces. These limitations motivate the development of computational frameworks that connect residue-scale structural cues with interactome-scale organization. Here we present PCIPG, a multi-scale probabilistic graph framework that jointly models residues, proteins, interactions and complexes. PCIPG encodes residue-level physicochemical descriptors on intra-chain contact maps, screens informative residues to construct structure-aware protein representations and propagates these representations over the PPI graph to infer a protein-complex membership matrix. To couple complex membership with sparse interaction evidence, PCIPG reconstructs the network using a zero-inflated Bernoulli-Exponential likelihood, providing a principled learning signal under missing-edge and noise regimes. Across five Saccharomyces cerevisiae benchmarks, PCIPG achieved higher average F1 and Acc than the representative baseline methods included in this study, with average improvements of 11.46% and 3.64%, respectively. On the evaluated human interactomes, PCIPG achieved the highest F1 score among the compared methods on HCT116 and HEK293T, whereas its performance on HuRI was below that of AdaPPI and ClusterONE. Embedding-guided interaction completion improved PCIPG's performance relative to its results on the corresponding original human PPI networks. Beyond complex calling, PCIPG supports core-module mining by recovering known cores and delineating coherent accessory modules within assemblies; several predictions match previously reported functional entities, including TRAPPII- and PCNA-loading-factor-related complexes. At the residue level, residues prioritized by PCIPG show increased overlap with experimentally defined protein-binding interfaces in the evaluated structures. In a computational CFTR case study, the model generated state-dependent interaction predictions that partially overlapped with experimentally profiled wild-type and ΔF508 interaction networks. Together, PCIPG bridges residue-scale structural cues with interactome-scale organization to enable interpretable and scalable protein complex identification. Code and data are available at https://github.com/hyx-1/PCIPG.
Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing at random data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values. The performance of this approach is evaluated through an extensive empirical study using both synthetic and controlled real-world proteomics datasets. Overall, we believe that this Bayesian framework for (multivariate) differential analysis provides a valuable and intuitive counterpart to standard methods at a comparable computational cost.
Microbial symbiosis is widespread among metabolically coupled cells; it presumably gave rise to mitochondria. However, how such symbioses emerge, evolve, and stabilize are unknown, particularly in the prokaryotic domain where endosymbiosis is virtually nonexistent. Yet there is growing evidence suggesting that mitochondria originated from such a metabolically driven prokaryotic partnership rather than phagocytotic predation. While prokaryotes almost ubiquitously engage in metabolic syntrophy, it is unknown whether syntrophy alone can enable stable physical associations that could pave the road toward physical integration. Here, we tested the hypothesis that syntrophy can transition into stable ectosymbiosis, using an ecological mathematical model. Starting from an existing syntrophic partnership between free-living hosts and symbionts, we demonstrate that population-level obligate ectosymbiosis can emerge and stabilize, even in unilateral syntrophy where only the symbiont consumes a host-produced metabolite. A key assumption is that the hosts’ by-product inhibits their growth when it accumulates. By consuming the toxic by-product, the symbiont locally reduces hosts’ self-inhibition at the contact surface, manifesting as a private benefit providing selective advantage. Our results show that due to the direct and indirect benefits, the ectosymbiotic consortium is stable against free-living forms and the consortial cooperation is ecologically selected for. Furthermore, solid metabolic coupling promotes population-level obligacy, ultimately excluding free-living individuals under stricter conditions. Our results support the hypothesis that cooperative, syntrophic microbes (particularly prokaryotes) are capable of forming stable, physical, and species-specific ectosymbiosis through inhibition reduction, providing a plausible first step toward potential, gradual endosymbiotic integration. Our work bridges the gap between models of microbial cooperation between free-living species and models that assume already-concluded, fully integrated endosymbiosis under multilevel selection.
E. coli relies on the heat shock response (HSR) to preserve protein homeostasis under stress, through three feedback modules: feedforward translational control, chaperone-mediated sequestration and targeted degradation. Although previous studies have highlighted how this layered architecture ensures rapid and robust protection compared to simpler designs, not much attention is paid to how these modules interact. Moreover, how do interactions among the three modules balance performance trade-offs, where gains in one module may come at the expense of another, yet together yield an optimal overall response? We address this using a mathematical model that integrates protein folding with σ 32 regulation. We show that the feedback modules both cooperate and compete, giving rise to nonmonotonic dynamics that govern HSR performance. Specifically, increasing feedforward strength does accelerate response, but beyond a threshold, despite increasing chaperone levels, it paradoxically slows recovery. Similarly, while sequestration enhances relative chaperone production and per-chaperone efficiency, when excessive, it traps σ 32 in inactive complexes, prolonging recovery and delaying shutdown. Mapping the parameter space reveals regimes of synergy as well as trade-offs between speed and efficiency, with wild-type parameters lying near the optimal region. These results reveal design principles that produces a robust and efficient heat shock response.
Cryptic species complexes pose fundamental challenges to biologists, as species exhibit minimal morphological differences that require integrating morphology, genetics, and biogeography for identification. Here, we present a deep learning approach to support species identification in the freshwater snail genus Radomaniola (Hydrobiidae), a morphologically cryptic group from the Balkans. Our approach mirrors the integrative workflow of expert taxonomists by combining shell images, morphometric measurements, and collection‑site metadata, with optional phylogenetic information. Despite being trained on fewer than 700 specimens across 20 visually similar species with strongly imbalanced class sizes, the system achieved high identification performance. Careful control of spurious correlations, such as those arising from site‑specific imaging conditions or overly precise geographic metadata, was essential to ensure that the network learned biologically meaningful features. Across all experiments, integrating multiple data types and jointly optimizing meaningful embeddings and classification consistently improved performance over image‑only and classification‑only baselines. On specimens from collection sites seen during training we achieved a macro-averaged F1 score of 0.93. Even though this dropped as low as 0.14 when evaluating on specimens from previously unsampled localities, it could be rapidly recovered by retraining with 2-3 newly labeled specimens. Additionally, model top-3 accuracy stayed consistently above 80% in all settings. These results show that relatively lightweight deep learning models can provide practical decision support in real taxonomic workflows.
Process-based malaria transmission models are important tools for evaluating intervention strategies, quantifying uncertainty, and informing malaria control policy. Individual-based models such as malariasimulation are computationally demanding, which limits their practicality for applications that require large numbers of simulation runs. In this paper we present malariasimple, a simplified, compartmental model implemented as an R package which approximates the epidemiological structure and parameter definitions of malariasimulation while operating at a fraction of the computational cost. Across a range of transmission intensities and intervention scenarios, malariasimple closely reproduces key outputs of malariasimulation while reducing runtimes by up to 99.6%. Its computational efficiency enables full Bayesian parameter inference, allowing estimation of complete posterior distributions. malariasimple provides a fast, flexible, and mechanistically consistent addition to the Imperial College London Malaria Model framework, bridging the gap between computational efficiency and epidemiological realism. The malariasimple R package is freely available for download at https://github.com/mrc-ide/malariasimple.
Musculoskeletal simulations can offer valuable insight into how the properties of our musculoskeletal system influence the biomechanics of our daily movements. One such property is muscle's initial resistance to stretch, also known as short-range stiffness, which is key to stabilizing movements in response to external perturbations. Short-range stiffness is poorly captured by existing musculoskeletal simulations since they employ phenomenological Hill-type models lacking activation-dependent stiffness properties. Existing simulations also do not capture the history-dependent reduction in short-range stiffness after muscle shortening, known as muscle thixotropy. While cross-bridge models can reproduce muscle short-range stiffness, it remains unclear which model properties are necessary to capture its history dependence. Here, we tested the ability of various cross-bridge models to reproduce empirical short-range stiffness and its history-dependent changes across a broad range of behaviorally relevant length changes and activation levels, using an existing dataset on 11 permeabilized rat soleus muscle fibers. We quantified muscle thixotropy using the ratio between the observed short-range stiffnesses after and before shortening. We computed the root-mean-square deviation (σSRS) between the predicted short-range stiffness ratio of various muscle models and the measured stiffness ratio. We found that cross-bridge models captured short-range stiffness changes across conditions with both small and large history-dependent stiffness reductions (σSRS ≤ 0.1), but only when including cooperative activation of both thin and thick myofilaments. In contrast, Hill-type models and a cross-bridge model without cooperative myofilament activation underestimated short-range stiffness and did not capture its change across conditions with large history-dependent stiffness reductions (σSRS > 0.2). Similar results were obtained when using a Gaussian-approximated solution method to simulate the cross-bridge distribution, but at an approximately eightfold lower computational cost. We therefore propose to implement Gaussian-approximated cross-bridge models with cooperative myofilament activation into musculoskeletal simulations to improve the prediction of short-range stiffness during movements.
T4 and T5 neurons are the first direction-selective neurons in the visual pathway. They are the most numerous cell types in the fly brain (~6000 within each optic lobe) and, as a population, their compact dendritic arbours span the entire visual field. They are classified into four subtypes (a, b, c, and d). Each subtype encodes one of four orthogonal motion directions (up, down, forwards, backwards). Crucially, the dendrites of these neurons are oriented inversely to the functional direction of motion which they encode. This dendritic orientation is what ultimately determines their functional directional encoding. The development of these neurons is well characterised up to the point of neuropil innervation. However, the full population of these neurons innervate their target neuropil prior to the emergence of directionality within their dendrites. As it stands, development prior to the emergence of dendritic orientations, and the adult oriented dendrite are both well understood, but the key components relating to the emergence of orientation itself are missing. Recent whole-brain electron microscopy (EM) connectomes of Drosophila melanogaster provide an unprecedented level of resolution and completeness when considering the morphology of neurons. Utilising this, we isolate the dendritic arbour of every T4 and T5 neuron within a female adult Drosophila brain, made available through FAFB-FlyWire. In doing so we are able to rigorously quantify the morphology of these dendrites in order to understand their similarities and differences. In doing so we aim to shed light on the origins of dendritic directionality. We reason that either this emerges through a tightly controlled, subtype specific mechanism, or is the result of a subtype agnostic mechanism and external factors. In the former case, we would expect evidence of this in differences between the morphological structure of individual dendrites between T4 and T5, and their subtypes. Our analysis however reveals a high degree of structural similarity between T4 and T5, and within their subtypes. Particularly, the geometry of branching, section orientation, and tree-graph structure of these dendrites show only minor variability, with no consistent separation between T4 and T5, or their subtypes. These results indicate that, despite forming in different neuropils, and serving distinct motion directions, T4 and T5 dendrites follow closely aligned morphological patterns. This suggests a shared mechanism of directed outgrowth, as opposed to symmetry breaking emerging through neuron type or subtype specific mechanisms.
The increasing interconnectedness of the modern world calls for globally equitable solutions to combat pandemic challenges. However, we have seen a tendency in recent decades for high-income countries to resort to “vaccine nationalism,” hoarding vaccine production to the detriment of lower-income countries. In addition, vaccine nationalism can prove detrimental to hoarding countries in the long term, as inequitable global vaccine distribution during the COVID-19 risked exacerbating the rise of harmful immune-escape variants that largely counteracted the original benefits of vaccine hoarding. Thus, vaccine hoarding may create a problem of time preference for a vaccine-producing country, where countries heavily discounting the future would opt for vaccine hoarding while countries lightly discounting the future would opt for vaccine sharing. Using a novel modeling framework integrating epidemiological, evolutionary, and economic processes, we demonstrate how high temporal discounting, low levels of outgroup prosociality, and high vaccine-distribution costs for low-income countries can promote vaccine-hoarding tendencies. We further show how these factors interact with epidemiological and evolutionary parameters to incentivize vaccine sharing in different ways: in some parameter regimes, vaccine sharing helps by reducing variant infections, while in others, vaccine sharing helps by reducing the probability of initial variant emergence. As a result, the optimal fraction of vaccines a country should share in our model is a bimodal function of the pathogen’s transmissibility. We thus provide a nuanced, model-based exploration of how various factors may contribute to vaccine nationalism’s emergence, emphasizing the need for international organizations to coordinate global vaccination responses to future pandemics.
Men who have sex with men (MSM) remain disproportionately affected by HIV worldwide. This systematic review summarizes the results of mathematical modeling studies that evaluated prospects of HIV elimination among MSM by geographical setting, type of intervention(s), elimination definition, and model characteristics. We searched Embase and PubMed for studies published between July 1, 2016 and September 1, 2025 which used a dynamic mathematical model to assess the impact of interventions on HIV transmission among MSM. Data were extracted on study population, interventions, elimination definitions, model type, model structure, and calibration. Studies were critically appraised for model comprehensiveness in addressing elimination. 135 of the 4,595 records were included. MSM populations in six of the eight Joint United Nations Programme on HIV/AIDS regions were modeled, with 47% of models considering MSM in the USA. Agent-based models (ABMs) were as common as compartmental models overall, with ABMs more frequently used in Western and Central Europe and North America (WCENA), while compartmental models predominated elsewhere. Of the 135 included studies, 41 defined elimination, and they defined it as follows: (i) reduction in HIV incidence/prevalence, or (ii) threshold of HIV incidence/prevalence, or (iii) reproduction number below one. Elimination was achieved in 42 out of 51 modeled scenarios, of which 32 (82.05%) were in WCENA, but the authors of only 28 of these 42 scenarios discussed the real-world elimination feasibility with 10 of these scenarios considered elimination feasible by the original authors. There was also a strong regional divide in the elimination scenarios considered feasible, with 6 (60.00%) in Asia and the Pacific (AP) and 4 (40.00%) in WCENA. Models in which elimination was achieved commonly used combinations of interventions. Modeling efforts to understand HIV elimination prospects outside WCENA should be intensified, and models assessing HIV elimination prospects should account for HIV acquisition outside of the local context. To enhance study comparability and ensure that models contribute effectively to public health policy, an elimination definition based on an HIV incidence threshold would be the most valuable. By identifying gaps in current studies, we recommend novel research directions for modeling to inform a coordinated global response for HIV elimination among MSM.
G protein-coupled receptors (GPCRs) are central membrane receptors and major therapeutic targets. However, predicting GPCR function remains difficult because experimentally annotated receptors are scarce, functional labels follow a long-tailed distribution, and structural information remains underused. Here, we present GPCR-GO, a relation-aware heterogeneous graph attention framework that integrates structural similarity with protein-protein interactions (PPIs). GPCR-GO uses the Dictionary of Protein Secondary Structure (DSSP) to transform three-dimensional protein structures into residue-level structural descriptors, aggregates these descriptors into protein-level structural vectors, and uses the resulting vectors to define structural-similarity edges. The framework builds a heterogeneous graph linking proteins and Gene Ontology (GO) terms through PPI edges, structural-similarity edges, GO hierarchy edges, and reviewed protein-GO annotations. Relation-aware graph attention aggregates complementary biological signals, whereas graph decomposition, hard negative mining, and semi-supervised learning improve learning under sparse supervision and class imbalance. On the held-out GPCR test split, GPCR-GO outperforms existing methods and achieves F-score (Fmax) values of 0.514, 0.767, and 0.631 on biological process (BP), cellular component (CC), and molecular function (MF), respectively. These results show that structure-derived relations complement curated annotation and support accurate GPCR function prediction under limited supervision.
Alternative splicing is a fundamental biological mechanism that increases protein diversity and regulates critical cellular processes across eukaryotes. Dysregulation of splicing is implicated in a wide range of diseases, including cancer, neurological disorders, and autoimmune conditions. Accurate prediction of splicing metrics such as percent spliced in (PSI) is therefore essential for understanding splicing regulation and improving disease characterization. However, existing approaches typically require high sequencing depth and are thus poorly suited for low-coverage settings such as single-cell RNA sequencing, where sparse read counts limit reliable splicing analysis. Here, we present ASPIRE (Accurate Splicing Prediction from Limited RNA Sequencing), a deep learning framework for predicting alternative splicing metrics from low-depth RNA-seq gene expression data. ASPIRE infers PSI values from gene expression profiles with limited read coverage and incorporates an embedded feature selection mechanism that identifies a minimal, informative subset of genes relevant to splicing regulation. This design enables accurate prediction while reducing reliance on extensive sequencing and mitigating noise introduced by irrelevant or weakly informative genes. By focusing on biologically meaningful features, including RNA-binding proteins, ASPIRE maintains strong predictive performance even under conditions typical of single-cell transcriptomics. We demonstrate that ASPIRE accurately predicts PSI values across a range of sequencing depths, including those characteristic of single-cell RNA-seq, and performs comparably to or better than existing methods in both simulated and real datasets. By enabling robust expression-based splicing inference from sparse data, ASPIRE facilitates the study of alternative splicing at cellular resolution and provides a practical framework for investigating splicing regulation in development, disease, and heterogeneous cell populations.
Metal-binding sites (MBSs) are critical determinants of protein stability and biological function, yet methods for comparing their local binding environments lag behind those for whole-structure alignment. Here, we represent MBSs as atomic point clouds surrounding bound metal ligands and align them with a fine-tuned iterative closest point algorithm. Applying this framework to a redundancy-reduced collection of MBSs derived from all metalloproteins in the Protein Data Bank (PDB), we perform pairwise alignments across 23,342 sites to construct a similarity network of metal-binding environments. The resulting network topology recapitulates metal coordination chemistry and enzyme function: links are strongly enriched within metal types and across shared EC subclasses. Conserved metalloenzyme families form cohesive subnetworks; for example, the binuclear ureohydrolase domain appears as two tightly connected components that also capture atypical members such as the dinickel metformin hydrolase. We observe only a moderate global association between protein sequence and MBS geometry, yet many network links connect near-identical binding-site architectures across proteins with low sequence identity, consistent with either divergent evolution with local MBS conservation or candidate cases of molecular convergent evolution. Integrating network proximity with structural evidence of drug binding identifies drugs with enriched connectivity among their targets and predicts 528 drug-off-target combinations across 88 drugs and 151 human proteins, recovering both known off-targets (e.g., ADAM/ADAMTS for matrix metalloproteinase inhibitors) and proposing novel ones. The MBS network thus provides a scalable resource for probing metalloprotein evolution, functional convergence, and the structural basis of drug cross-reactivity.
Mathematical modelling of antibiotic resistance plays an important role in understanding the mechanisms of resistance emergence and spreading, testing the feasibility of new treatment protocols, and antimicrobial stewardship. However, many assumptions underlying some of the most commonly used mathematical models have not been rigorously tested experimentally. We verify whether one of these models - a birth-death-mutation process - is able to quantitatively predict the outcome of laboratory experiments. We grow bacteria in a bioreactor in conditions that closely resemble the assumptions of the model, and compare the model predictions with experimental observables such as the probability and time to resistance evolution, mutant number distribution, and the genetic composition of the evolved populations. We show that the model fails to reproduce some aspects of the experiments (failing differently for different antibiotics) but that simple modifications of the model significantly improve its predictive power. These modifications give insight into the population dynamics of resistant mutants for each antibiotic tested, and highlight the importance of quantitative modelling for accurate prediction of antibiotic resistance evolution.
CLDN18.2 is a promising tumor-specific antigen; however, the development of therapeutic antibodies against it is challenged by the need for simultaneous optimization of affinity and developability. To address this, we present cdrGPT, a deep learning framework based on GPT-2 for de novo generation of complementarity-determining region H3 (CDRH3) sequences. Our approach integrates pre-training on the Observed Antibody Space (OAS) database with structural templating derived from the known antibody zolbetuximab. Generated sequences were iteratively refined through rejection sampling and fine-tuned against a multi-parameter objective function encompassing predicted affinity and MHC class II binding risk. From an initial set of 50,000 sequences, this screening pipeline yielded 313 high-confidence candidates. Subsequent analysis using evolutionary scale modeling 2 (ESM2) embeddings, principal component analysis (PCA), and clustering revealed three structurally distinct clusters, with intra-cluster cosine similarities exceeding 0.99. Validation of seven representative sequences from the dominant cluster using AlphaFold3 confirmed high structural fidelity to the zolbetuximab template, demonstrating a root mean square deviation (RMSD) of 1.331 Å for the CDRH3 loop and positional deviations of less than 0.4 Å for key paratope residues. These results indicate that the designed variants preserve the core binding mode of the parent antibody. This study establishes a feasible pipeline for integrating AI-generated CDRH3 loops into functional antibody scaffolds, providing a foundation for the accelerated development of therapeutics targeting CLDN18.2 and other clinically relevant antigens.
A key computational step in reference-based variant calling is distinguishing true genetic variants from sequencing errors. Advanced tools and workflows have been developed to handle this by computational modelling of technical errors from the sequencing machines. However, these recalibration workflows have largely been evaluated for human data only and its exact applicability for non-human data remains unknown. Here, we conducted a systematic evaluation of variant calling on human, rice, sheep, and chickpea data, and found that existing workflows introduce unexpected statistical bias, thus leading to suboptimal variant calls for non-human data. To address this problem, we present simple guidelines for constructing a "pseudo-"database (pseudoDB) of genetic variants as a scalable and portable solution for recalibration and variant calling. With human data, our pseudoDB-based workflow performs comparably to existing dbSNP-based GATK3 workflows and those using DeepVariant, Strelka2, and FreeBayes. We extend this to other non-human genomes, namely cattle, brown bear, swan goose, African oil palm, Komodo dragon, and stevia, altogether resulting in the identification of up to 242.0% unique genetic variants. The majority of newly identified variants are within the non-coding regions, hinting at the rich diversity of genome regulation in the non-human population. Our pseudoDB-based workflow is agnostic to reference genomes and modular for easy integration with other computational workflows for human and non-human resequencing data.