Motivation Large-scale bioinformatics analyses increasingly require collaboration across multiple cohorts and institutions, yet existing workflows often rely on data co-localization, which is slow, difficult to scale, and raises privacy concerns. We present a privacy-preserving federated analytics framework that enables secure statistical analysis across distributed datasets without transferring raw data, by performing all computations on encrypted data via cryptographic methods.Results We evaluate the framework by validating polygenic risk scores and conducting meta-analyses on two real-world cohorts. The proposed solution achieves over 99.9% accuracy relative to plaintext analyses, while maintaining scalable runtime performance with increasing data size and number of participating sites. These results demonstrate the feasibility of secure federated analytics for practical bioinformatics applications involving sensitive data.
Population-scale proteomics is driving precision medicine by enabling systematic drug target identification and robust biomarker discovery. Comparable, well-powered studies across diverse global populations are essential to elucidate shared and population-specific disease pathways. Mass-spectrometry-based and multiplexed affinity-based assays have emerged as complementary, leading strategies for quantifying proteomic variation in large-scale population-based studies. With the objective of performing a comprehensive comparative evaluation of the latest assays for each of these technologies, we compare three leading affinity-based and mass-spectrometry-based proteomic platforms (SomaScan11K, Olink Explore HT, Orbitrap Astral with Seer Proteograph [MS-Seer]) in a multi-ethnic Asian cohort, to inform biomarker discovery and functional genomic studies of global populations. We find limited overlap of 1,740 proteins out of 12,825 total proteins quantified across the three platforms, with modest correlations (0.34–0.10). SomaScan had lower missingness (< 1
Allergic reactions to fish pose complex food safety challenges, driven by species diversity and underexplored intraspecies variability. This study comparatively examines allergen profiles in Malabar red snappers (Lutjanus malabaricus, n = 39) across body sizes, anatomical regions and production origins using SDS-PAGE, immunoblotting and quantitative mass spectrometry. Protein profiles varied greatly by fish size and muscle region, but not by origin. Smaller fish contained higher levels of major allergen parvalbumin and creatine kinase, while larger fish exhibited elevated levels of heat-labile allergens enolase, aldolase, and glyceraldehyde-3-phosphate dehydrogenase. Parvalbumin levels were highest in head, followed by belly, dorsal, and tail. Greatest variation was observed for the three heat-stable allergens - parvalbumin, tropomyosin, and collagen. Minimal origin-dependent differences affected 2 of 11 registered fish allergens. We established an integrated proteomics workflow for systematic allergenicity assessments to uncover intraspecies variability and provide foundational knowledge for understanding intraspecies variability to improve food safety strategies.
One of the promises of precision medicine is to understand and act on inter-individual genetic differences in drug responses. SNPdrug3D contains the complete genomic landscape of missense single nucleotide variants (SNV) across the human proteome and at a population-wide level that could affect drug binding. Here, we map SNVs in over 80,000 individuals from the Singapore SG10K Health and gnomAD cohorts to identify ~1.17 million variants mapped to residues near ~6000 bound drugs in protein-drug complexes and experimentally verify effects of selected SNVs, including previously uncharacterized variants, on drug binding in relevant proteins ranging from kinases to cytochrome P450s (CYPs). The latter led to a specific predictor for interpreting variants in the CYP family that outperforms existing tools in the prediction of pharmacogenetic effects based on database-annotated (AUROC = 0.9) or assay-based (AUROC = 0.8) test sets. By placing variants and drugs in structural contexts, SNPdrug3D aids drug development by pre-emptively flagging potential resistance sites based on population-specific variability.
Shellfish is a leading cause of life-threatening food anaphylaxis. Protein extraction is a critical step in allergenic protein (allergen) detection and quantification. This study evaluates the impact of protein extraction strategies on allergen detection in the two most widely consumed shellfish species, black tiger prawn (Penaeus monodon) and white leg prawn (Litopenaeus vannamei). Proteins were extracted from raw and boiled muscle tissues in four different buffers: urea-based, sodium dodecyl sulphate (SDS)-based, SDS-based with reducing agent, and phosphate-buffered saline (PBS), the most commonly used buffer in allergy investigations. Protein abundances were analysed using SDS-PAGE and label-free liquid chromatography-tandem mass spectrometry. Allergenicity was predicted using AllerCatPro and subsequently evaluated for antibody-based immunoreactivity. Protein extraction strategies significantly influenced proteome coverages, IgE-antibody binding, allergen diversities and relative abundances. Although PBS-based extracts yielded the highest overall relative allergen abundances (~84-95%), they generated the lowest allergen diversities and reduced relative abundances of key allergens, including tropomyosin and myosin light chain from raw tissues and arginine kinase from boiled tissues. In contrast, denaturing extracts enabled broader allergen representation, with urea-based extracts from boiled tissues generating the most comprehensive allergen repertoire. In addition to well-characterised allergens, up to 72 low-abundant proteins with strong bioinformatic evidence of allergenicity were identified, with their relative recovery varying across extraction strategies. Optimised protein extraction improves allergen detection workflows, supporting the discovery of underreported and low-abundant candidate allergens for future immunological validation, and facilitating more comprehensive risk assessment strategies.
Molecular docking is a vital computational task in drug discovery, wherein the objective is to efficiently identify optimal binding poses between a ligand and a target receptor protein. Due to the combinatorial explosion of possible binding configurations, docking of large and flexible molecules remains a computationally intensive problem, especially at scale. Early studies have revealed that the molecular docking can be re-cast as a maximum vertex-weighted clique problem (MVWCP) problem on a compatibility graph to be solved classically. In this work, we proposed a hybrid quantum-classical approach for molecular docking leveraging the MVWCP formalism with a variational full-basis encoding (FBE) strategy, which enables efficient encoding of classical binary variables with Bloch sphere vectors. We further prove that a global minimizer of the FBE objective can always be chosen to be a pure product state, thereby providing a rigorous justification for its optimization using a unitary variational circuit. The molecular docking problem is first mapped to a cost Hamiltonian that is minimized within a variational framework, optimized via a randomized imaginary time evolution (ITE)-inspired warm start, and gradient-based techniques. Finally, we also executed the circuit on an IBM quantum computer, underlying the feasibility and of quantum-assisted optimization for structure-based drug design and point towards the broader utility of advanced encoding techniques in quantum optimization for computational biology.
Novel clade 2.3.2.1e A(H5N1) virus was detected in cerebrospinal fluid but not in respiratory, rectal swab, or blood samples of an 8-year-old boy presenting with meningoencephalitis without respiratory symptoms. Cerebrospinal fluid A(H5N1) hemagglutinin-specific antibody levels were higher than those of sera. Clinicians should be aware of emerging clade 2.3.2.1e A(H5N1)-associated meningoencephalitis.
Shellfish allergy is a major cause of food-induced severe adverse reactions worldwide, with shrimp representing one of the most implicated triggers. Detection and monitoring of allergenic proteins in shrimp-containing foods rely largely on immunoassay-based detection, typically targeting tropomyosin without species-level specificity. However, allergen composition may vary across species and processing conditions, limiting the accuracy of existing approaches. Here, we applied quantitative proteomics to characterise allergen profiles in five commonly consumed shrimp species under raw and heat-treated conditions. Protein extracts (n = 3 per species/treatment) were analysed by liquid chromatography-tandem mass spectrometry (LC-MS/MS) following in-solution digestion. Identified proteins were mapped against curated allergen databases and species-specific transcriptomes to assess presence and relative abundance. Complementary SDS-PAGE and immunoblot analyses using allergen-specific antibodies were performed to evaluate qualitative detection across species and treatments. Proteomic analysis revealed marked interspecies variability in allergen presence and abundance. Several allergens exhibited differential stability following heat treatment, reflecting distinct biophysical properties influencing persistence in heat-processed foods. Selected putative allergens identified by transcriptome analysis, including aldolase and enolase, were confirmed at the protein level, while others were not detected. Moreover, weak concordance between transcriptomic and proteomic abundance was observed for multiple allergens, including arginine kinase and myosin light chain. Immunoblotting for tropomyosin and myosin light chain demonstrated inconsistent detection across species and treatments, highlighting limitations of antibody-based detection systems. These findings establish that shrimp allergen composition is species- and heat processing-dependent, with implications for risk assessment frameworks. Quantitative proteomics provides a robust platform for comprehensive allergen profiling and supports improved detection strategies for aquatic food safety.
Seafood allergy is complex due to extensive species diversity, posing major challenges in food safety assessments, clinical diagnosis and dietary management. However, the absence of established workflows to resolve allergenomes limits correlations between allergen abundance, clinical sensitisation, and consumer risk. Mass spectrometry (MS)-based proteomics overcomes limitations of conventional immunoassay allergen detection by enabling unbiased protein identification and quantification, including allergen isoforms and low-abundance proteins within complex matrices. An integrated workflow combining immunological analyses, liquid chromatography-MS/MS proteomics, and bioinformatics was developed to characterise allergenomes across eight commonly consumed Asia-Pacific fish species. Comprehensive allergen profiles were established using in silico allergenicity predictions with AllerCatPro, combined with immunological validations using allergen-specific antibody and pooled patient sera. Across eight species, 529-1012 protein groups were identified, including all 11 fish muscle allergens, with pronounced interspecies differences in allergen composition and isoform distribution. The major fish pan-allergen parvalbumin was the most abundant allergen of most species and varied in abundance by up to 7-fold. Mackerel displayed a distinct low-parvalbumin profile with enriched metabolic allergens. Tissue heating induced a consistent shift toward enrichment of heat-stable and tissue-retained allergens, particularly parvalbumin, tropomyosin, and collagen. In contrast, heating of raw extracts generated more variable and species-specific retention of selected proteins, including heat-labile metabolic enzymes. In silico analysis predicted 14 proteins with strong allergenicity evidence for further validation. Fish allergenomes are species-specific and processing-dependent, positioning quantitative proteomics as a powerful platform for improved molecular risk assessment, and the development of representative diagnostic and food safety reference materials.
OBJECTIVE:The COVID-19 pandemic response relied heavily on statistical and machine learning models to predict key outcomes such as case prevalence and fatality rates. These predictions were instrumental in enabling timely public health interventions that helped break transmission cycles. In this study, we aimed to assess the effectiveness of multimodal data in forecasting SARS-CoV-2 case surges across different phases of the pandemic, characterized by varying levels of data availability. MATERIALS AND METHODS:Most existing models are grounded in traditional epidemiological data. The potential of alternative datasets, such as those derived from genomic information and human behavior, remains underexplored. In the current study, we investigated the usefulness of diverse modalities of feature sets in predicting case surges using machine learning models. RESULTS AND DISCUSSION:Our results highlight the relative effectiveness of biological (e.g., mutations), public health (e.g., case counts, policy interventions) and human behavioral features (e.g., mobility and social media conversations) in predicting country-level case surges. Importantly, we uncover considerable heterogeneity in predictive performance across countries and feature modalities, suggesting that surge prediction models based on alternative data may need to be tailored to specific national contexts and pandemic phases. CONCLUSION:Overall, our work highlights the value of integrating alternative data sources into existing disease surveillance frameworks to enhance the prediction of pandemic dynamics.
Biocatalysis provides a sustainable approach for highly selective chemical transformations, however the complex, non-additive nature of enzyme fitness landscapes remains a significant barrier hindering the predictive design of optimized enzymes. Here, we document a curated dataset of kinetic colorimetric activity measurements of galactose oxidase variants for biocatalytic oxidation of a secondary alcohol, 1-phenylbutan-1-ol. Using the engineered galactose oxidase variant, GOh1052, as the backbone, site saturation mutagenesis libraries covering 352 residue positions were constructed, yielding 6,686 unique single mutants. These mutants were subsequently screened via a colorimetric assay to capture their kinetic activity profiles over a 16-hour reaction period. This dataset provides a data foundation for the development of machine learning models for applications such as enzyme-substrate activity prediction or biocatalytic process optimization.
We compared three leading affinity-based and mass-spectrometry-based proteomic platforms (SomaScan11K, Olink Explore HT, Orbitrap Astral with Seer Proteograph [MS-Seer]) in a multi-ethnic Asian cohort, to inform biomarker discovery and functional genomic studies of global populations. We found limited overlap (1,740 proteins) across the three platforms, with modest correlations (0.34-0.10). SomaScan had lower missingness (<1%) and CV (<10%), compared to Olink (51%, 23%) and MS-Seer (12%, 19%). The new assays in Olink Explore HT (absent in Explore3072) primarily drove its higher missingness and CV. The number of phenotypic associations varied by trait, while the number of genetic associations ( cis- pQTLs at P<5e-08) were similar for Olink and SomaScan. Protein levels differed between ethnicities, with SomaScan identifying more ethnicity-differentiated proteins than Olink (FDR<0.05). Finally, SomaScan ANML normalization attenuated biologically relevant associations in our study. These findings underscore the importance of platform evaluation and data normalization strategies for application in large-scale, diverse population cohorts. ### Competing Interest Statement L.D.W works for Alnylam Pharmaceuticals and holds stocks as part of employment. O.B works for Bayer AG and does not hold stocks as part of employment. Z.D works for Boehringer Ingelheim Pharma GmbH & Co. KG and does not hold stocks as part of employment. J.F works for Novo Nordisk A/S and holds minor share portions as part of employment. The rest of the authors declare no competing interests. NMRC Singapore, NMRC/StaR/0028/2017, MOH-000271-00, NMRC/PRECISE/2020
Zoonotic influenza viruses, including highly pathogenic avian influenza and swine-origin variants, continue to cause sporadic human infections with, in some cases, high case fatality rates and potential for sustained human-to-human transmission. The COVID-19 pandemic underscored both the possibilities of rapid vaccine innovation and the persistent challenges in equitable access and public trust. This paper synthesizes the vaccine-related priorities from the 2024 update of the World Health Organization Public Health Research Agenda for Influenza, integrating evidence from systematic literature reviews commissioned, expert consultations, and analysis of lessons learned from recent health emergencies, to outline a research and policy roadmap for zoonotic and pandemic influenza vaccine preparedness. Key research priorities identified include development of broadly protective animal and human vaccines; improved understanding of correlates of protection; rapid and scalable manufacturing platforms; predictive modelling for strain selection; and targeted communication strategies to strengthen uptake. Experts have considered that implementing these priorities will require One Health integration, sustained investment, harmonized regulatory frameworks, and proactive community engagement to ensure that advances in vaccine science translate into timely, equitable public health protection.
Malassezia are commensal lipid-dependent yeasts and opportunistic pathogens that cause superficial mycoses and systemic infection. Azole antifungals target cell wall ergosterol synthesis and are the first line of antifungal treatment. ERG11 gene mutations and overexpression are major mechanisms conferring azole resistance and resulting in antifungal therapy failure. Malassezia restricta is found ubiquitously on healthy and diseased skin, with azole-resistant isolates described. Malassezia arunalokei is a relatively new, closely related common skin species. Ketoconazole and itraconazole were the most effective at inhibiting both species. Isolates of M. restricta and M. arunalokei from healthy skin of Singapore subjects were cultured, evaluated, and generally susceptible to common over-the-counter azoles, including clotrimazole, except for select less-susceptible strains. Some less-susceptible strains have novel or reported non-synonymous mutations in the ERG11 gene, such as R88C. The QK178RQ ERG11 sequence variation was observed to be associated with differences in M. restricta and M. arunalokei as independent species. In the absence of identified ERG11 mutations, strains with elevated MICs were observed to have elevated ERG11 expression and drug efflux pump expression/activity. We conclude that antifungal susceptibility is determined by a combination of intrinsic (e.g., mutations, gene expression, efflux pump activity) and extrinsic (e.g., skin condition, prior antifungal exposure) factors and that the skin microbiome serves as a reference for the emergence of new mutations and strain phenotypes. IMPORTANCE:Malassezia over colonization is associated with conditions such as dandruff and seborrheic dermatitis, which give rise to unpleasant itching and swelling on the skin. Azole antifungals such as ketoconazole, clotrimazole, and miconazole are the primary treatments of choice available as over-the-counter creams or shampoos. However, the emergence of antifungal resistance leads to a loss of treatment efficacy and persistent fungal infection. To understand the mechanisms underlying antifungal resistance, we profiled the susceptibility profiles of commensal Malassezia isolates from the skin and identified novel ERG11 mutations. Our results indicate that antifungal susceptibility is determined by a combination of factors (mutations, efflux pump activity, gene expression, copy number) and suggest that the healthy skin microbiome serves as a reference for the emergence of new mutations and strain phenotypes.
The COVID-19 pandemic has prompted an unprecedented global response. In particular, extraordinary efforts have been dedicated toward monitoring and predicting variant emergence due to its huge impact, particularly for vaccine escape. Broadly, we classify such methods into two categories: forward mutation prediction, where phenotypes are first observed and the responsible genotypes traced, and reverse mutation prediction, which starts with selected pathogen genetic profiles and characterizes their associated phenotypes. Reverse mutation prediction strategies have advantages in being able to sample a more complete evolutionary space since sequences that do not yet exist can be sampled. The rapid improvement in the maturity and scale of reverse mutation prediction strategies, such as deep mutational scanning, has led to significant amounts of data for machine learning, with concomitant improvement in the prediction results from computational tools. Such integrated prediction approaches are generalizable and offer significant opportunities for anticipating viral evolution and for pandemic preparedness.
The growth of industrial biocatalysis for sustainable chemical manufacturing has been limited by the narrow range of chemistries associated with natural enzymes and experiment-intensive regimes of enzyme engineering. Consequently, there has been deep interest to expand enzyme substrate scopes for broader synthetic utility, and to streamline the enzyme engineering process. In the field of alcohol oxidation, galactose oxidase (GOase) is one of the most established enzymes capable of this important chemical transformation under benign conditions. However, the applicability of GOase towards more complex molecules such as those frequently found in the pharmaceutical, or agrochemical industries remains restricted. Here, by employing a combined approach of directed evolution and predictive modelling, we have identified new GOases with significantly expanded substrate specificity toward both bulky benzylic and unactivated secondary alcohols, showing activity enhancements of up to 2,400-fold compared to the reported benchmark M3-5 mutant. Beneficial mutations conveying relaxed substrate enantioselectivity biases (R/S ratios down to 1.05) and higher thermostabilities (up to 20-fold versus benchmark) have also been identified. We have developed predictive models based on computational tools YASARA, FoldX, SCWRL and Glide that are well correlated with features related to enzyme structure, selectivity, protein stability and catalytic activity. The generated enzyme activity models based on Glide-MM/GBSA (r = -0.85) and YASARA (r = -0.89) have successfully predicted the activity trend of a family of related substrates based on the 1-phenyl-1-alkyl alcohol scaffold with varying alkyl chain lengths. It is envisioned that these in silico models can serve as valuable tools to explore desirable enzyme characteristics, establish enzyme substrate scopes, and accelerate biocatalyst development, thus promoting it as a competitive and competent solution for sustainable chemical manufacturing.
Advances in sequencing technology have enabled whole genome sequencing of microorganisms, allowing their rapid identification and characterisation and measure of relatedness to the highest resolution possible. This facilitates the tracking of pathogen evolution and spread, as well as the identification of variants of interest or concern. Pathogen genomics came into the spotlight with the COVID-19 pandemic. The unprecedented, close to real-time collection of SARS-CoV-2 whole genome sequences spurred rapid diagnostic development and informed outbreak response and surveillance strategies. This chapter provides an overview of the technical processes of sequencing and discusses the value and limitations of pathogen genomics in surveillance and outbreak response.
IntroductionFish is a major food allergy trigger with a complex variety of allergenic protein isoforms and vast species diversity exhibiting variable allergenicity. This is the first study to systematically compile fish isoallergen and variant entries associated with ingestion-related allergic reactions.MethodsEntries were compiled from four major allergen databases: World Health Organization and International Union of Immunological Societies (WHO/IUIS), AllergenOnline, Comprehensive Protein Allergen Resource (COMPARE), and Allergome, including evidence from in vitro IgE-binding assays and complete amino acid sequences. Challenges in predicting the allergenicity of fish isoallergens and variants were evaluated, and the sensitivity of five widely used in silico tools (AllerCatPro 2.0, AlgPred 2.0, pLM4Alg, AllergenFP v.1.0, and AllerTop v.2.0) was assessed. Epitope mapping and phylogenetic analyses were performed for the major fish allergen parvalbumin, incorporating experimentally validated B-cell epitope data from the Immune Epitope Database (IEDB) and evolutionary relationships.ResultsA comprehensive dataset of 79 unique fish isoallergen and variant entries from 34 fish species was identified, with 25 entries common across all four databases. AllerCatPro 2.0 achieved the highest sensitivity (97.5%). A phylogenetic tree was constructed, integrating epitope data to optimize protein family-specific thresholds for differentiating allergenic from less/non-allergenic parvalbumins. A threshold of ≥4 IEDB-mapped epitopes allowing up to two mismatches captured 52 out of 54 parvalbumin sequences (96%) in the dataset, effectively distinguishing between parvalbumin classes.DiscussionThis study enhances understanding of fish allergy by systematically compiling fish isoallergens and variants and integrating B-cell epitope data. The optimized thresholds improve the performance of allergenicity prediction tools and can be applied to other protein families in future studies.
The demand for botanicals and natural substances in consumer products has increased in recent years. These substances usually contain proteins and these, in turn, can pose a risk for immunoglobulin E (IgE)-mediated sensitization and allergy. However, no method has yet been accepted or validated for assessment of potential allergenic hazards in such materials. In the studies here, a dual proteomic-bioinformatic approach is proposed to evaluate holistically allergenic hazards in complex mixtures of plants, insects, or animal proteins. Twelve commercial preparations of source materials (plant products, dust mite extract, and preparations of animal dander) known to contain allergenic proteins were analyzed by label-free proteomic analyses to identify and semi-quantify proteins. These were then evaluated by bioinformatics using AllerCatPro 2.0 (https://allercatpro.bii.a-star.edu.sg/) to predict no, weak, or strong evidence for allergenicity and similarity to source-specific allergens. In total, 4,586 protein sequences were identified in the 12 source materials combined. Of these, 1,665 sequences were predicted with weak or strong evidence for allergenic potential. This first-tier approach provided top-level information about the occurrence and abundance of proteins and potential allergens. With regards to source-specific allergens, 129 allergens were identified. The sum of the relative abundance of these allergens ranged from 0.8% (lamb's quarters) to 63% (olive pollen). It is proposed here that this dual proteomic-bioinformatic approach has the potential to provide detailed information on the presence and relative abundance of allergens, and can play an important role in identifying potential allergenic hazards in complex protein mixtures for the purposes of safety assessments.