
Tea (Camellia sinensis L.), a major global economic crop in Asia, poses challenges for genetic identification because its highly heterozygous, repetitive genome reduces the efficacy of conventional single-nucleotide polymorphism (SNP) and microsatellite markers, and interspecific hybridization further complicates the situation. To address these issues, CamK-DB was developed as a reference-free Camellia fingerprinting database built on MIKE MinHash sketches. We curated 418 candidate resequencing datasets, and built a database using standardized 5× genome-coverage fingerprints. Each accession is stored as a MIKE. jac fingerprint generated with k = 21 and recommended sketch/pre_cnt = 2000. CamK-DB provides a command-line interface for data management and a custom C++ query engine that computes top-10 matches using Jaccard similarity, complemented by a QT-based graphical interface for interactive analysis. This resource offers a robust and scalable framework for precise and routine germplasm identification, genomic phylogenetic inference, and strategic breeding program design. CamK-DB (database and code) is publicly available at https://github.com/sc-zhang/CamK-DB. CamK-DB binaries are provided for Windows 10/11 and Linux (x86_64, glibc ≥ 2.27).
Single-cell long-read transcriptomics (scLR-seq) extends single-cell analysis beyond gene abundance by resolving full-length transcript structures in individual cells. It can directly interrogate isoform usage, alternative splicing, and transcription start and end site selection, thereby revealing regulatory variation that is often obscured by short-read measurements. In this review, we examine the experimental and computational foundations of scLR-seq, including platform selection, library design, cell barcode and unique molecular identifier recovery, transcript discovery, and isoform quantification. We discuss how these choices influence the reliability of downstream biological interpretation, and summarize emerging insights into isoform usage, alternative splicing, transcription start and end site selection, allele-specific expression, fusion transcripts, transposable element-derived transcripts, and RNA modifications. Finally, we highlight applications of scLR-seq in diverse biological systems, such as the immune system, neural development, and tumor microenvironments, and consider future opportunities and challenges in integrating multi-omics data to decode cellular programs and disease evolution.
BACKGROUND:The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. FINDINGS:To address this, we developed a public and sustainable solution leveraging ontology-based APIs. All available ICTV Master Species List (MSL) releases, from MSL1 to MSL41, were transformed into a unified, semantically structured ontology comprising more than 195,000 current and historical entities and deployed through the Ontology Lookup Service (OLS). The ontology is automatically rebuilt and republished whenever a new MSL release becomes available. Complementary ICTV-NCBI mappings and helper libraries support integration into downstream systems. CONCLUSIONS:Together, these resources enable, for the first time, public programmatic retrieval of current and historical ICTV taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints, including resolution of former taxonomic terms to their current accepted taxon or taxa and retrieval of taxon histories across releases. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness.
BACKGROUND:Protein nanopores are essential molecular gateways in biology and have inspired transformative technologies in biosensing and single-molecule sequencing. However, the discovery and engineering of novel nanopore scaffolds remains limited due to the scarcity of experimentally resolved pore structures. RESULTS:Here, we present NanoporeDB, an open-access structural resource comprising about 7,000 high-confidence multimeric models across 4 representative pore types. Using a structure- and sequence-guided mining strategy, we identified candidate nanopores from large protein datasets, including the AlphaFold Protein Structure Database, UniRef90, and MGnify90, and generated high-confidence multimeric models using AlphaFold-Multimer and AlphaFold3. Collectively, these models represent a >170-fold expansion of the structurally annotated nanopore repertoire. Each model is further annotated with predicted membrane embedding, pore geometry, and constriction profiles, enabling structure-informed functional inference. NanoporeDB features an interactive web interface with 3D visualization and quantitative metrics. CONCLUSIONS:NanoporeDB provides the first comprehensive structural resource of multimeric protein nanopores with explicit membrane and pore annotations. This resource provides a structural gateway for advancing nanopore-based molecular sensing, precision diagnostics, and synthetic biology. NanoporeDB is publicly available at https://db.genomics.cn/nanopore.
Recent advances in multi-omics technologies have catalyzed the construction of comprehensive brain cell atlases, providing essential data foundations for artificial intelligence (AI)-driven analyses in precision neurology. This review systematically examines how the integration of AI with single-cell multi-omics and spatial multi-omics advances the resolution in deciphering brain cellular architecture across health and disease states. Through systematic evaluation of multi-omics datasets from neurodegenerative, psychiatric, and neurodevelopmental disorders, we demonstrate how AI facilitates disease subtype stratification, biomarker discovery, and therapeutic target identification. We critically address translational challenges, including data standardization, model interpretability, and regulatory frameworks for clinical implementation. Notably, the establishment of the International Consortium for Primate Brain Mapping in 2025 exemplifies ongoing global collaborative efforts toward systematic multi-omics atlas construction across species and disease states. This synthesis underscores a paradigm shift toward AI-enabled, mechanism-driven analyses, ultimately positioning precision neurology as a realizable framework for individualized diagnosis and targeted interventions in complex brain disorders. .
BACKGROUND:Spatial transcriptomics (ST) enables a high-resolution interrogation of molecular characteristics within specific spatial contexts and tissue morphology. Despite its potential, visualization of ST data is a challenging task due to the complexities in handling, sharing, and visualizing large image datasets together with molecular information. RESULTS:We introduce ScopeViewer, a browser-based software designed to overcome these challenges. ScopeViewer offers the following functionalities: (1) it visualizes large image data and associated annotations at various zoom levels, allowing for intricate exploration of the data; (2) it enables dual interactive viewing of the original images along with their annotations, providing a comprehensive understanding of the context; (3) it displays spatial molecular features with optimized bandwidth, ensuring a smooth user experience; and (4) it bolsters data security by circumventing data transfers. CONCLUSIONS AND DISCUSSIONS:ScopeViewer offers the research community a convenient, powerful, and secure software for high-resolution images, including pathology images and ST. It serves as an open-source platform for imaging-based research. Future enhancements and new features will be shared on GitHub by the creators and are open for contributions from other researchers. ScopeViewer is freely available on the website at https://cdc.biohpc.swmed.edu/scopeviewer.
BACKGROUND:The rapid advancement in single-cell, spatial omics, imaging, and genomic technologies requires robust analytical and visualisation platforms capable of managing complex biological data. Tools such as Multi-Dimensional Viewer (MDV) offer comprehensive interfaces for data exploration but often require advanced computational expertise and manual configuration to generate visualisation outputs, limiting accessibility for many users. RESULTS:We present ChatMDV, a natural language interface integrated with MDV that enables users to generate high-quality, interactive visualisations and analyses through natural language commands. ChatMDV employs a retrieval-augmented generation pipeline in combination with large language models to translate user queries into executable, reproducible Python code and interactive output. This conversational layer facilitates both exploratory and targeted analyses in diverse biological domains. We demonstrate ChatMDV's capabilities using 3 datasets of increasing complexity: the Peripheral Blood Mononuclear Cells 3K single-cell RNA-sequencing (scRNA-seq) dataset, the lung cancer atlas scRNA-seq dataset included in the Human Cell Atlas, and the longitudinal TAURUS study scRNA-seq dataset. Across all use cases, ChatMDV produced high-quality, reproducible visualisations from simple natural language queries, achieving a high semantic success rate between 79% and 97% when visualising the datasets. CONCLUSIONS:By bridging the gap between natural language processing and bioinformatics visualisation, ChatMDV reduces technical barriers, enhances reproducibility, and supports more inclusive scientific inquiry. Its modular design and adherence to Findability, Accessibility, Interoperability, and Reuse (FAIR) principles make it a scalable and adaptable framework for accelerating biological data analysis.
BACKGROUND:Spectral libraries are essential for mass spectrometry-based metabolomics, enabling accurate metabolite annotation. Collision-induced dissociation (CID) dominates existing public libraries, but is rarely sufficient for structural elucidation. Electron-activated dissociation (EAD) provides complementary, radical-driven fragmentation, but remains sparsely represented. The lack of datasets spanning multiple dissociation mechanisms, energies, and ionization modes limits both analytical workflows and the development of robust machine learning models. FINDINGS:We present MultiMS2, a curated metabolomics spectral library comprising 43,728 MS/MS spectra from 2,899 unique compounds. Spectra were acquired using both CID and EAD at three energies each, in positive and negative ionization modes. The dataset substantially expands publicly available EAD coverage while preserving matched acquisition conditions across energies and dissociation types. CONCLUSIONS:By systematically combining CID and EAD across multiple energies and polarities, MultiMS2 provides a unique resource for metabolite annotation, benchmarking, and machine learning. The library supports energy-aware and dissociation-aware analysis, enabling methodological innovation and improved generalization in computational metabolomics.
BACKGROUND:We present the design and implementation of a data curation framework to generate a large-scale clinical brain imaging dataset suitable for artificial intelligence (AI) enabled image analysis. FINDINGS:The dataset is accessible through the Brain Health Data (BHD) initiative, which includes ~417,341 magnetic resonance imaging (MRI) and 846,077 computerized tomography head studies, linked electronic health records, and associated free-text imaging reports from clinical practice between 2010 and 2018 in Scotland, exceeding 185 TB in size. The data curation framework was developed during the SCottish AI in Neuroimaging to Predict Dementia and Neurodegenerative Disease (SCANDAN) study, which used a subset of 41,966 MRI series from the BHD for dementia prediction.We describe the processing of the BHD metadata and our multilabel classification output. We discuss the strengths of the BHD, including clinical relevance thanks to its unprecedented scale, population-wide representativeness of a national free-at-the-point-of-delivery healthcare, long-term follow-up to neurodegenerative disease, and real-world variability. We describe the challenges and lessons learnt in developing a framework to curate data, including the time needed to obtain permissions, the need for easily accessible, secure, responsive and affordable computational environments, the variability of clinical data, and the challenge of extracting linked clinical data and images at scale. CONCLUSION:This resource will be crucial for clinical research, fostering the development of personalized medicine approaches, and fast-tracking the implementation of AI models in clinical workflows. We encourage the use of the BHD data through a streamlined application to the Public Benefit and Privacy Panel for Health and Care via the electronic Data Research and Innovation Service of Public Health Scotland (eDRIS).
BACKGROUND:Identifying de novo mutations (DNM) is an important component of both genetic research studies and clinical diagnostic workflows, but is complicated by distinguishing true mutations from sequencing errors. Likelihood-based error models are more accurate than inferring mutations from genotypes alone, but the resulting callsets still have high false positive rates. RESULTS:We identify that the main source of false positive DNMs comes from the use of genotype likelihoods in an otherwise robust mutational model. To address this issue, we propose two alternative methods that build on an existing DNM calling approach, DeNovoGear, but with higher accuracy and no decrease in sensitivity.Furthermore, we developed a method that collects allele-specific frequency profiles in the sequenced cohort from across many unrelated samples and identifies sites that either demonstrate high rates of sequencing and mapping errors, or are unlikely to be clinically significant due to their high recurrence rate.
Traditional Kaplan-Meier curves capture aggregate survival trends within broad patient subgroups but overlook the heterogeneity of individual patients. In contrast, single-patient survival risk models bridge this gap by incorporating each patient's unique clinical, genomic, and demographic characteristics, generating personalized survival curves. These individualized visualizations enhance patient-clinician communication by translating complex statistics into intuitive, time-based visuals that are easier to interpret. However, the complexity, high dimensionality, and heterogeneity of multiomics data present significant challenges for analysis, interpretation, and model development. To address these challenges, we introduce the Cancer Patient Survival Model (CPSM), an R package designed to deliver individualized survival and risk predictions through a fully integrated, reproducible computational pipeline. CPSM includes 10 core functions organized into 4 key steps: (i) data preprocessing and normalization, (ii) feature selection, (iii) survival risk group prediction modeling, and (iv) visualization and nomogram construction. We demonstrate the utility of CPSM using publicly available datasets from The Cancer Genome Atlas for 4 cancer types: glioblastoma multiforme (GBM), acute myeloid leukemia (LAML), pancreatic adenocarcinoma (PAAD), and breast invasive cancer (BRCA). CPSM efficiently handles high-dimensional datasets with over 60,000 RNA transcripts and diverse clinical variables, enabling robust and interpretable individualized survival predictions under varying data conditions. Model performance was evaluated using repeated cross-validation with uncertainty quantification, ensuring robust and reliable estimates in high-dimensional, small-sample settings. In summary, CPSM provides an efficient, user-friendly, end-to-end solution for integrating patient data and generating personalized survival and risk predictions. Its integrated visual tools enhance interpretability and support more informed clinical decision-making. The package is freely available on Bioconductor (https://bioconductor.org/packages/devel/bioc/html/CPSM.html) and GitHub (https://github.com/hks5august/CPSM).
BACKGROUND:African swine fever (ASF) remains a persistent threat to global pig production, with no licensed vaccines or effective treatments available. Observations of surviving individuals within low-virulence infected herds suggest that host genetic resistance plays a crucial role. RESULTS:Here, we present a multi-dimensional integrative analysis to uncover host genomic variants associated with ASF resistance. Combining genome-wide association studies (GWAS), genetic differentiation, and functional genomic approaches, including TWAS, SMR, colocalization, and Bayesian network GWAS, we prioritized 135 high-priority candidate resistance genes from an initial gene set of 1,102 candidates. These prioritized genes are enriched in immune-related pathways, such as chemokine signaling and IL-15-mediated activation. Heritability enrichment and transcriptomic analyses further revealed tissue- and cell-type-specific expression patterns, particularly in peripheral immune organs and pulmonary alveolar macrophages. Dynamic infection-responsive genes, including CXCL10, CXCL11, and IL15, exhibited robust antiviral signatures, which highlighted Mac_CD163 as key cellular mediators in the immune response to ASF. Moreover, multiple genes (such as SOS1, FCGR2B, FCGR3) converged on the PI3K-AKT and Fcγ receptor signaling axes pathways, underscoring their functional importance. Finally, we developed a polygenic resistance score using 40 prioritized independent SNPs, which effectively discriminates phenotypic outcomes and showed a positive correlation with health traits such as platelet distribution width. CONCLUSIONS:These findings provided a genomic foundation for the precision breeding of ASF-resistant pigs and inform host-targeted disease control strategies.
This study presents the first publicly accessible electroencephalography (EEG) dataset explicitly targeting sit-to-stand and stand-to-sit transitions during both motor execution (ME) and motor imagery (MI) tasks. Twenty-two healthy participants performed sitting and standing transitions under well-controlled experimental conditions while 60-channel EEG, electrooculography (EOG), and electromyography (EMG) signals were synchronously recorded. The dataset enables the exploration of neural activation patterns associated with lower-limb movements and supports the development of EEG-based brain-computer interface (BCI) algorithms for mobility assistance and rehabilitation. To validate the dataset, benchmark classification was conducted on three baseline deep learning methods-CTNet, EEGNet, and TCANet. Given the high inter-subject variability inherent to EEG, leave-one-subject-out cross-validation is used to ensure no subject bias during evaluation. Results demonstrated consistent decoding performance with mean accuracies of approximately 81% for ME and 73% for MI, indicating the reliability and usability of the dataset. Additionally, analyses of movement-related cortical potentials (MRCPs) and event-related desynchronization/synchronization (ERD/ERS) patterns revealed distinct neural signatures across the transition phases. This dataset provides a comprehensive foundation for studying lower-limb motor control, neural dynamics, and the advancement of MI-based BCIs for rehabilitation and assistive technologies.
The Carnegie stages represent a critical window in human development, during which organ primordia emerge, tissue identities diversify, and many congenital disorders are thought to originate. However, this period has remained difficult to study at the whole-embryo scale. Recent advances in spatial and single-cell genomics are beginning to close this gap. This article discusses the significance of the spatiotemporal transcriptomic atlas of post-gastrulation human embryos spanning Carnegie stages 12-23, which provides one of the most comprehensive dynamic molecular maps of this developmental window to date. Beyond its value as a reference resource, this atlas offers biological insights into early cardiac patterning, brain regionalization, the timing of inhibitory and excitatory neurogenesis, tissue-specific susceptibility to prenatal infection, and spatially resolved allelic imbalance. Equally importantly, it illustrates a broader shift in the field from fragmented organ-specific datasets toward integrated developmental frameworks that can connect embryology, human genetics, and disease mechanisms. Future progress will depend not only on generating more data but also on harmonizing multimodal datasets across sources, stages, and platforms, an effort that will increasingly rely on advances in computational biology and artificial intelligence.
BACKGROUND:Understanding how organisms reconstruct complex tissue architectures following injury requires precise mapping of gene expression and cellular responses across space and time. Although planarians serve as a classic model for whole-body regeneration, capturing the continuous spatiotemporal dynamics of positional information and cell fate decisions at the organismal scale remains a significant challenge. RESULTS:Using high-definition spatial transcriptomics, we generated a 4-dimensional atlas encompassing over 3.5 million cells from whole animals across 8 distinct regeneration timepoints. This comprehensive dataset enabled the definition of 36 spatial domains and the tracing of body axis restoration, revealing that positional control genes recover through self-organizing dynamics analogous to an underdamped control system. We identified an injury-induced spatial domain termed the anterior regenerative zone. This unique region is characterized by the convergence of epidermal, muscular, and neural lineages enriched with positional signals. Furthermore, we demonstrated that the transcriptional co-factor Mediator 8 is a critical regulator of this zone. Depletion of Mediator 8 impairs the formation of the anterior regenerative zone, disrupts polarity establishment, and prevents successful blastema formation. CONCLUSIONS:Our study provides a holistic molecular and cellular reconstruction of whole-body regeneration, directly linking dynamic gene expression gradients to morphological restoration. The discovery of the Mediator 8-regulated anterior regenerative zone highlights the importance of spatial domains in coordinating tissue repair. The resulting interactive atlas serves as a foundational resource for deciphering the logic of spatiotemporal patterning in regeneration.
BACKGROUND:Antimicrobial resistance genes (ARGs) and virulence factors (VFs) are central contributors to the global health crisis surrounding drug-resistant infections. FINDINGS:We introduce PathoFact 2.0, an enhanced pipeline for improved ARG, VF, toxin, and biosynthetic gene clusters (BGCs) prediction. Key improvements include an updated machine learning (ML) model for VF identification, expanded hidden Markov model profiles for VFs and toxin-associated proteins, a new ML model for toxin and toxin-associated proteins identification, and the integration of antiSMASH 7.0 for predicting BGCs. CONCLUSIONS:Our upgrades make PathoFact 2.0 a more powerful and user-friendly platform for predicting microbiome-based pathogenicity and resistance, providing a crucial tool for better understanding and addressing the challenges posed by antimicrobial resistance and infectious diseases.PathoFact 2.0 is available at https://gitlab.com/uniluxembourg/lcsb/systems-ecology/pathofact2. It is compatible with Linux operating systems.
Background Many research domains are producing large, multi-scale, multi-modal datasets at growing rates with mixed variable types (continuous, discrete, censored). Identifying possible cause-effect associations in such datasets is essential for predicting outcomes and proposing possible interventions. Probabilistic graphical models (PGMs) have emerged as a robust, interpretable way to analyze such datasets, but current graph learning algorithms cannot incorporate time-to-event (censored) variables, which are important in many systems (e.g., patient survival). Instead, regression models are typically used for survival analysis of single censored variables, but these cannot assess cause-effect interactions.Results Here, we present a new mathematical framework to incorporate multiple censored variables into mixed graphical models. A novel efficient algorithm, CausalCoxMGM, is implemented, which is extensively evaluated on synthetic and real-life high-dimensional biomedical datasets (cardiovascular disease, breast cancer). CausalCoxMGM was able to recover effectors of censored variables, supported by literature, and provided new mechanistic insights on the differences between ER+ and ER- breast cancers.Conclusions CausalCoxMGM is a flexible computational framework for learning potential cause-effect relations from observational data of mixed data types, including multiple censored variables. The resulting graphs are interpretable and can be used to generate testable hypotheses or build efficient predictors of any outcome.
Asian genomic datasets possess unparalleled potential to advance global understanding of human genetic diversity. Encompassing the world's largest population pool with diverse ethnicities, these datasets capture comprehensive genomic variations shaped by heterogeneous socioeconomic conditions, climate exposures, and clinical environments. However, current national genome initiatives across Asia demonstrate substantial disunity, stemming from limited cross-border communication and collaborative infrastructure, thereby diminishing their collective impact on biomedical research and precision medicine development. The MedHackathon Asia 2025 catalyzed crucial dialogues toward establishing a regional community dedicated to three pillars: harmonized biobank collaboration, standardized genomic data protocols, and cooperative governance frameworks. This multidisciplinary convening brought together researchers, clinicians, bioinformaticians, and national precision medicine program leaders from across Asia to share best practices, identify implementation challenges, and formulate foundational strategies for sustained cooperation. This community review synthesizes critical outcomes from these deliberations, emphasizing the imperative for continuous regional collaboration while advocating for the development of sustainable architectures enabling: (1) equitable biobank resource sharing, (2) genomic data standardization, and (3) ethical governance models. Through consolidation and expansion of this emerging network, Asian nations are expected to lead transformative contributions to global genomic science while ensuring appropriate representation in biomedical innovation. Such coordinated efforts promise to accelerate healthcare advancements with equitable benefits extending throughout the region and worldwide.