
Butterflies have been a model system for studying the evolution of colour. This is partly due to their complex patterns that reflect human-visible (VIS) and ultraviolet (UV) light, which are perceived by conspecifics and predators. Many studies have sourced data from publicly available images, but most of these images only consider the visible spectrum of light. Including the UV spectrum is crucial for fully understanding the evolution of butterfly morphology and behavioural ecology. Here we provide standardized images (VIS and UV) of over 4 000 individuals from 16 communities of Australian butterflies. These communities represent different climates and urbanization levels spanning over 2 500 km. The dataset contains at least one individual of 125 different species from five families, constituting over one quarter of Australian butterfly diversity. In addition to photographs, we provide spectral measurements of butterfly wings for at least one individual of each species and sex, and Cytochrome Oxidase subunit 1 (CO1) sequences of 1 635 individuals. All these data are accessible in Zenodo and an associated R package simplifies the download of subsets of the database. This database will be of use to evolutionary biologists and ecologists interested in a broad range of topics related to phenotypic variation.
Querying the RDF Portal knowledge graph maintained by DBCLS-which aggregates ~60 life-science databases-requires proficiency in both SPARQL and database-specific RDF schemas, placing this resource beyond the reach of most researchers. Large Language Models (LLMs) can, in principle, translate natural-language questions into executable SPARQL, but without schema-level context, they frequently fabricate non-existent predicates or fail to resolve entity names to database-specific identifiers. We present TogoMCP, a system that recasts the LLM as a protocol-driven inference engine orchestrating specialized tools via the Model Context Protocol (MCP). Two mechanisms are essential to its design: (i) the MIE (Metadata-Interoperability-Exchange) file, a concise YAML document that dynamically supplies the LLM with each target database's structural and semantic context at query time; and (ii) a two-stage workflow separating entity resolution via external REST APIs from schema-guided SPARQL generation. On a benchmark of 50 biologically grounded questions spanning five types and 23 databases, TogoMCP achieved a large improvement over an unaided baseline (Cohen's $d = 1.82$, Wilcoxon $P \lt .001$), with win rates exceeding 80% for question types with precise, verifiable answers. An ablation study shows that all component configurations deliver significant improvements, with MIE schema files providing the largest marginal contribution on mean per-question score ($\Delta = +0.50$ relative to a no-MIE condition, two-sided Wilcoxon $P = .067$; 90% bootstrap CI $[+0.04,\,\,+0.94]$ excludes zero); a one-line instruction to load the relevant MIE file recovers the same mean improvement as a full procedural protocol, while the protocol additionally reduces downside risk (loss rate 1.6% vs. 4.8%, Fisher $P = .036$). These results suggest a general design principle: concise, dynamically delivered schema context is more valuable than complex orchestration logic for mean score performance, while procedural guidance plays a complementary role in narrowing variance. Database URL: https://togomcp.rdfportal.org/.
The family Cyprinidae represents the most taxonomically diverse group of freshwater fishes, encompassing over 3 000 species of ecological and economic importance in aquaculture, conservation, and ecological monitoring. Simple sequence repeats (SSRs), also known as microsatellites, are highly informative molecular markers widely used for genetic diversity analysis, population structure assessment, and marker-assisted breeding. However, comprehensive genome-wide SSR resources for cyprinids remain limited. Existing databases, such as FishMicrosat, provide restricted taxonomic coverage and lack standardized chromosome-level datasets suitable for comparative genomic analyses. To address this gap, we developed CypriSSR, a genome-wide SSR database encompassing 11 representative cyprinid species. Chromosome-level genome assemblies were retrieved from NCBI. SSR loci were identified using MISA, and primer pairs were designed using Primer3. In total, over 7.8 million SSR loci were identified. Mononucleotide repeats were the most abundant class (39-53%), followed by dinucleotide and trinucleotide motifs. SSRs were predominantly distributed in genic regions (54%-71% across species), suggesting potential functional roles. Each database entry includes repeat type, genomic coordinates, primer sequences, melting temperatures, and predicted PCR product sizes. The CypriSSR web interface enables flexible querying of SSR markers based on species, chromosome, motif class, and genomic location, and supports sequence similarity searches through an integrated BLAST module along with data export options. CypriSSR provides a comprehensive and standardized multi-species microsatellite resource for cyprinid genomes and supports applications in population genetics, molecular breeding, and conservation genomics. Database URL: http://46.202.167.198/fishssr/.
Fast-neutron mutagenesis creates diverse genome-wide mutations, providing a powerful tool for crop functional genomics. Here, we present an expanded genomic and phenotypic analysis of 3268 fast-neutron (FN)-induced mutant rice lines (Oryza sativa L. cv. Kitaake). All FN lines were whole-genome sequenced, and mutations were identified by alignment in the Nipponbare and KitaakeX reference genomes. We cataloged over 428,000 mutations affecting 78.49% of Nipponbare genes and 70.38% of KitaakeX genes. In silico expression analysis indicates that 575 non-mutated Nipponbare genes are highly expressed and likely essential for viability. Each mutant carries, on average, 68.5 mutations in the Nipponbare alignments or 63.2 mutations for KitaakeX alignments, distributed randomly across all 12 chromosomes with no evident hotspots. FN lines have approximately 8.5% fewer mutations when using the KitaakeX alignment, underscoring the unique contributions of each reference genome and the importance of utilizing both for comprehensive mutation discovery. The majority of mutations are small deletions and single-base substitutions, with deletions predominating in their effect on genes. We found that 74.4% of all transcription factor Nipponbare genes were mutated at least once. Phenotypic characterization of over 2700 lines revealed a broad spectrum of variation in core agronomic traits (heading date, tiller number, plant height, panicle weight, seed yield components) and other morphological variants of interest. The integration of genomic and phenotypic data through the KitBase platform enabled the identification of candidate genes for several traits of interest. The KitBase website (https://kitbase.ucdavis.edu) has been updated to provide open access to all mutation data and seed stocks, as well as an intuitive query interface, facilitating forward and reverse genetic analyses in rice. This expanded resource enriches the rice functional genomics toolkit and highlights the value of coupling high-density mutation mapping with phenotypic data for rapid gene discovery and crop improvement.
Animal venomics is a growing field of research with evolutionary and biotechnological significance. Yet, fundamental questions regarding the origin, diversification, and bioactivity of venoms remain unresolved. Here, we analysed venom tissue-related data curated in Tox-Prot, currently the most comprehensive database for animal venoms, across three snapshots spanning two decades (2005, 2015, 2025). We assessed the taxonomic landscape related to Tox-Prot entries, sequence length distribution, protein family abundances, and habitat-specific venom patterns. Our results consistently show that snakes, spiders, cone snails, and scorpions, along with their associated protein families such as Phospholipase A2, Snake Three Finger Toxin and Long 4(C-C) scorpion toxin family, dominate across Tox-Prot. Nevertheless, the taxonomic and protein family diversity has been steadily increasing, with 503 new species and 188 new protein families added by 2025 compared to the 2005 dataset. Marine species account for 16%-20% of total species, of which 63%-85% are Neogastropoda reflecting limited marine species diversity coverage; likewise, terrestrial taxa are disproportionately represented by Squamata (39%) and Hymenoptera (20%) relative to their natural diversities of 2.3% and 50.26%, respectively. At the molecular level, half of all entries correspond to mature peptide sequences of 26-75 amino acids, featuring three to four disulfide bridges and C-terminal amidations as the most frequently recorded post-translational modifications. Further, protein language model embeddings infer taxonomical diverse peptide and delimited large enzyme clusters. Our study maps two decades of venom data diversification, reflecting both the field's rapid expansion and the need for robust integrated datasets to propel and disseminate that knowledge.
Food allergies represent a significant and growing global challenge. However, researchers and risk assessors still lack a dedicated and comprehensive resource that centralizes food allergen information. As a result, obtaining the necessary structural, biochemical, and immunological data often requires time-consuming searches across multiple heterogeneous databases, hindering efficient analysis and slowing scientific progress. The Food Allergen Database (FAD) merges biochemical, structural, immunological, and functional properties of all known food allergens and isoallergens (1168 entries). It is an essential resource that facilitates the extraction of all information about food allergens.
Cancer metastasis involves complex molecular mechanisms that cannot be fully explained by individual gene expression profiles. Previous studies have shown that correlations between RNA and miRNA expression can capture metastatic behaviour more effectively than expression of individual genes. However, no publicly available databases provide systematic analysis of RNA-miRNA correlations specific to cancer metastasis. We developed an efficient computational method to identify differential correlations between miRNAs and RNAs that are specific to individual tumour samples. Using data from The Cancer Genome Atlas (TCGA), we computed differential correlations for tumour samples across 9 cancer types and 21 metastatic sites, encompassing ~200 million RNA-miRNA pairs. Statistical analysis identified RNA-miRNA pairs with site-specific correlations using Mann-Whitney U-tests. MetaCancerDB contains RNA-miRNA correlation networks for 9 primary cancer types and 21 metastatic sites. Site-specific correlations showed distinct patterns, with lung metastasis displaying the most conserved correlations across cancer types. Survival analysis revealed that specific RNA-miRNA pairs are prognostic for patient outcomes in a metastatic site-dependent manner. MetaCancerDB provides a comprehensive resource for exploring RNA-miRNA correlations in cancer metastasis. The database enables researchers to identify molecular signatures specific to metastatic sites and can serve as a foundation for developing predictive biomarkers. MetaCancerDB is freely available for academic purposes.
Determining the functional consequence of missense mutations acquired in the development of cancer is critical to the understanding of the evolution and the therapeutic vulnerabilities of an individual tumour. Several million missense mutations associated with cancer have been reported across different databases with little functional annotation accompanying each mutation. We have designed the MOKCa-3D database, (https://bioinformaticslab.sussex.ac.uk/MOKCa-3D/) to enable the contextualization and interpretation of cancer somatic missense mutations, including the structural impact of the mutation on the 3D structure, and whether the mutation results in a gain or loss of the protein's function. For each protein, a sequence feature viewer enables interactive visualization of the amino acid sequence, missense mutations, post-translational modification sites, protein domains, active sites, binding sites, protein-protein interaction sites, and mutational frequency. The mutation-level page concisely presents functional insights for each individual mutation, and an interactive MOL* viewer highlights mutated residue on an AlphaFold protein structural model. The SAAP structural impact analysis pipeline was used to identify the structural impact of the mutation. MOKCa-3D concisely presents functional insights and structural impacts of cancer somatic missense mutations enabling users to interpret their functional consequences. It is freely accessible and easy to navigate, making it usable by the widest range of researchers.
The safety and modernization of traditional Chinese medicine (TCM) are significant concerns for human beings. Recently, the adverse effects caused by the use of certain TCMs have been frequently reported. Although TCMs may cause toxic reactions, they also play critical roles in treating multiple complex diseases. Therefore, toxicity research is urgently needed for the safe usage of TCMs. However, existing databases for TCMs primarily focus on the pharmacological effects of TCMs, with limited attention to the toxicity. They neither distinguish the toxic effects of formulas, herbs, and ingredients, nor classify and summarize targets for specific toxic manifestations, or assemble evidence from previous studies. We developed TCMToxDB, a comprehensive database that focuses on the toxicity and safe usage of TCMs. TCMToxDB systematically integrates and analyses the research results of toxic TCMs, offering users diverse information acquisition and analysis services. In addition, it assembles five canonical herb-target and ingredient-target interaction prediction algorithms with different advantages, which support the prediction of toxic targets of herbs and ingredients to empower the toxicity research of TCMs and to meet users' personal needs. TCMToxDB is accessible at https://www.sdu-idea.cn/TCMToxDB.
Scientific articles have been searched for experimentally established critical pharmacokinetic parameter-renal clearance. The main source of the documents was PubMed database, considered one of the most important and comprehensive repositories of biomedical literature. After manual data collection and thorough quality check database presenting human renal clearance values of exogenous substances was developed. After collecting the data, preliminary processing and simple analysis were carried out. The final database contains over 1700 experimental observations from 761 scientific articles and covers over 500 unique substances studied. Database URL: doi: https://10.17632/3427x3wzzc.2
Emerging studies highlight the importance of protein isoforms, which often exhibit distinct functional roles and contribute to physiological diversity, disease mechanisms, and phenotypic variation, despite originating from the same gene. However, comprehensive isoform-level resources that characterize protein isoforms remain limited. IsoProDB is an integrative and unified one-stop database that aligns protein isoforms from RefSeq and UniProtKB, enabling cross-sequence visualization for protein isoform analysis in humans. It integrates features such as domain architecture, intrinsically disordered regions, sequence variants, transmembrane topology, and 52 distinct post-translational modifications (PTMs) mapped to protein isoforms from multiple resources. Currently, IsoProDB enables users to perform gene wise comparative analyses across 110 149 protein isoforms derived from 20 536 protein-coding genes for all integrated features, supported by effective visualizations. This provides insights into conserved and nonconserved PTM sites, domains, isoform-specific membrane localization, the impact of variants on protein function, and disease relevance across protein isoforms. With specific isoforms emerging as markers and theragnostic targets for various disorders, IsoProDB is integrated with multiple global resources for easy navigation and exploration of multiomics information on isoforms.
Accurate, consistent and comprehensive metadata are essential for the reuse of functional genomics data deposited in repositories such as the Gene Expression Omnibus (GEO), however, achieving this often requires careful manual curation, which is time-consuming, costly and prone to errors. In this paper, we evaluate the performance of Large Language Models (LLMs), focusing on OpenAI's GPT-4o, as an assistive tool for entity-to-ontology annotation of two commonly encountered descriptors in transcriptomic experiments, mouse strains and cell lines. Using over 9 000 manually curated experiments from the Gemma database and over 5 000 associated journal articles, we assess the model's ability to identify relevant free-text entries and map them to appropriate ontology terms. Using zero-shot prompting and retrieval-augmented generation (RAG) to incorporate domain-specific ontology knowledge, GPT-4o correctly annotated 77% of mouse strain and 59% of cell line experiments, and uncovered manual curation errors in Gemma for over 200 experiments (2% of total). GPT-4o substantially outperformed non-LLM alternatives, and was statistically indistinguishable from the highest-performing 2026 frontier models. Model errors often arose from typographical mistakes or inconsistent naming in the GEO record or publication, and resembled those made by human curators. Along with annotations, our approach requests that the model output supporting context and verbatim quotes from the sources. These were typically accurate and enabled rapid curator verification. We further found that for the difficult cell line task, an ensemble of LLMs can boost precision at the cost of recall. These findings suggest that while LLMs are not ready to fully replace manual curators, they can effectively support them. A human-in-the-loop workflow, in which LLM's annotations are provided to human curators for validation, should improve the efficiency and quality of large-scale biomedical metadata curation.
Projects such as the European 1+ Million Genomes initiative and the European Genomic Data Infrastructure project are paving the way towards the age of genomic medicine. To address the challenge of balancing genomic data privacy with biomedical research, the proposed solution is to enable discovery of private datasets through public metadata. Yet enabling data discovery based on genomic variants present in a dataset-which is the goal of Beacon-raises the risk of re-identifying data subjects. We have implemented a Portuguese Beacon endpoint within the scope of the European Genomic Data Infrastructure project, which features a re-identification prevention algorithm to ensure the privacy of data subjects-the first Beacon endpoint to do so. We assessed the impact of the algorithm on data discovery, which varies with the size of the Beacon dataset.
The laboratory rat, Rattus norvegicus, is an important model of many human diseases, and experimental findings in the rat have relevance to both human physiology and disease. The Rat Genome Database (RGD, https://rgd.mcw.edu/) is a model organism database that provides access to a wide variety of curated rat data including disease associations, phenotypes, pathways, molecular functions, biological processes, cellular components, and chemical interactions for genes, quantitative trait loci (QTL), and strains. Because the laboratory rat is used often as a model of human physiology and disease, RGD has incorporated data from the NHGRI-EBI (National Human Genome Research Institute-European Bioinformatics Institute) catalog of human genome-wide association studies (GWAS) (https://www.ebi.ac.uk/gwas/). This provides the basis for easy integration of RGD data for rat and other species with human SNP (single nucleotide polymorphism)-phenotype associations. When data from different sources use different vocabularies, there must be some standard way to compare the data. Since the human GWAS data are annotated with Experimental Factor Ontology (EFO) terms, RGD needed to map those EFO terms to various ontologies used at RGD for annotating rat genomic and strain data, so the human data can be translationally associated with the wealth of preclinical data including rat genes, rat QTL, and rat strains. To bridge the ontology/vocabulary gap, the curators at RGD have mapped all the EFO terms (http://www.ebi.ac.uk/efo) used to annotate the human GWAS SNPs in the NHGRI-EBI Catalog of human GWAS Catalog Data (https://www.ebi.ac.uk/gwas/) to multiple ontologies used at RGD. RGD has used the mappings to translate human GWAS disease/phenotype EFO annotations into RGD Disease Ontology annotations, human phenotype ontology annotations, clinical measurement ontology annotations, and vertebrate trait ontology annotations. Also, RGD has made the ontological mappings available via Simple Standard for Sharing Ontological Mappings (SSSOM) files on the RGD download site (https://download.rgd.mcw.edu/ontology/mappings/).
Biological databases play a crucial role in life sciences research by organizing vast amounts of data, enabling efficient access and analysis. Numerous databases have been published across various research areas, yet there remains a need for updated platforms in the field of molecular docking and molecular dynamics simulation research. To address this gap, we have developed an extensive and user-friendly platform focused on compiling the binding energies of compounds associated with a wide range of biological activities. The database offers free access to data on 1321 compounds, including abstracts, references, isomeric SMILES, and 22 molecular properties. Researchers can also securely store their docking and screening data. To demonstrate its capabilities, molecular docking was performed on the top 10 compounds with the best binding energies against human metapneumovirus (HMPV) using AutoDock Vina and the crystal structure (PDB ID: 8FPJ). MK-3207 and Etoposide exhibited docking scores of -10.3 and -9.6, respectively. The top two compounds were further selected for MD simulations, confirming stable binding interactions with the viral protein. Additional compounds, including Teniposide, UK432097, 85019940, Setileuton, Orvepitan, Cep-11981, Tadalafil, and VS-5584, were also analyzed, providing further insights into their binding mechanisms and potential therapeutic relevance. The database is developed using PHP, HTML, CSS, JavaScript, and Python and is freely accessible at https://www.pbed.habdsk.org/.
Tumour heterogeneity often leads to substantial differences in responses to same drug treatment. The presence of pre-existing or acquired drug-resistant cell subpopulations within a tumour survive and proliferate, ultimately resulting in tumour relapse and metastasis. The drug resistance is the leading cause of failure in clinical tumour therapy. Therefore, accurate identification of drug-resistant tumour cell subpopulations could greatly facilitate the precision medicine and novel drug development. However, the scarcity of single-cell drug response data significantly hinders the exploration of tumour cell resistance mechanisms and the development of computational predictive methods. In this paper, we propose scDrugAtlas, a comprehensive database devoted to integrating the drug response data at single-cell level. We manually compiled more than 100 datasets containing single-cell drug responses from various public resources. The current version comprises large-scale single-cell transcriptional profiles and drug response labels from 1023 samples, across 77 unique drugs and 31 major cancer types. Particularly, we assigned a confidence level to each response label based on the tissue source (primary or relapse/metastasis), drug exposure time, and drug-induced cell phenotype. We believe scDrugAtlas could greatly facilitate the Bioinformatics community for developing computational models and biologists for identifying drug-resistant tumour cells and underlying molecular mechanism.Database URL: http://drug.hliulab.tech/scDrugAtlas/.
BACKGROUND:Skin diseases are among the most prevalent conditions worldwide, posing significant threats to human health by causing physical discomfort, psychological distress, and reduced quality of life. With the rapid advancement of high-throughput technologies, a substantial number of transcriptomic datasets, including single-cell RNA sequencing (scRNA-seq), spatial transcriptomics, and bulk RNA-seq, have been generated in the field of dermatology over the past decade. However, the lack of effective integration and standardized analysis pipelines limits the full utilization of these valuable resources in skin disease research. OBJECTIVES:To address this gap, we aimed to construct a comprehensive, integrative, and user-friendly atlas that enables systematic exploration of skin transcriptomic data across multiple diseases and modalities. METHODS:We developed the Human Skin Atlas (huSA) ('https://humanskinatlas.com/index.html'), a publicly accessible database that incorporates data from 17 skin diseases and 63 independent datasets, including 1 434 scRNA-seq, 63 spatial transcriptomics, and 1 502 bulk RNA-seq samples. The database provides standardized cell-type annotations, differential gene expression analysis, cell-cell interaction mapping, pathway and metabolic module enrichment, transcription factor regulatory inference, and differentiation state assessment for scRNA-seq data. Data from identical skin diseases were further integrated to enhance biological signal detection. For visualization, we embedded the 'cell × gene' and 'Cirrocumulus' platforms, offering interactive and customizable gene expression visualizations at both single-cell and spatial levels with user-defined parameters. RESULTS:The huSA enables both individual dataset analysis and cross-dataset integration, providing robust, consistent, and scalable insights into skin disease biology. Demonstration analyses confirmed that results derived from either single datasets or aggregated multi-dataset integrations exhibited high reliability and biological relevance. The platform successfully supports diverse research needs, including cell-type-specific expression profiling, regulatory network construction, and spatial transcriptomic exploration. CONCLUSIONS:The Human Skin Atlas (huSA) represents a state-of-the-art integrative resource for the skin research community. By offering multiscale analyses and interactive visualization tools, the huSA accelerates the discovery of molecular mechanisms underlying skin diseases and facilitates translational research efforts aimed at improving skin health.
Sickle cell disease (SCD) is one of the most prevalent monogenic disorders worldwide, with the highest burden in Africa, where ~75% of the 7.74 million global cases occur. Scientific progress in understanding its epidemiology, clinical heterogeneity, and treatment outcomes has been constrained by heterogeneous, non-standardized, and non-interoperable datasets that limit data integration and cross-country analyses. To address this, the Sickle Africa Data Coordinating Centre (SADaCC) was established as the data science hub of the SickleInAfrica consortium to support the development and expansion of Pan-African SCD registry. SADaCC now coordinates one of the largest patient-consented SCD datasets globally, with data from over 40 000 persons living with SCD in seven countries (Ghana, Mali, Nigeria, Tanzania, Uganda, Zambia, and Zimbabwe) within the Sickle Pan-African Research Consortium (SPARCo), as well as genomic data from SADaCC satellite sites in Cameroon, South Africa, and Malawi. The registry is built on FAIR-compliant architecture, the Sickle Cell Disease Ontology, and powered by a suite of digital platforms such as REDCap, NextCloud, RStudio, GitHub, Docker, and Jupyter. In partnership with SPARCo, SADaCC is also piloting a biobank that will link biospecimens with data in the registry to advance multi-omics research. Beyond infrastructure, SADaCC leads training and/or research in big data analytics, genomics, bioethics, implementation science, qualitative research, and psychosocial studies. Ethical, legal, and social considerations are embedded across all operations with emphasis on equitable intra-African collaboration and patient involvement in research. Looking ahead, SADaCC will integrate real-time data streams, AI-driven analytics, and multi-omics data to drive big data and genetic medicine research for SCD in Africa.
Motivation: Epilepsy is a diverse group of neurological disorders affecting over 50 million people worldwide. While common epilepsy types are well studied, rare epilepsies-often severe and genetically complex-pose significant challenges in diagnosis, research, and treatment. Accurate and interoperable etiology and disease classifications are critical for improving data sharing, supporting clinical decision-making, and advancing rare disease research. Results: To enhance the accuracy of epilepsy-related disease concept representation within the Mondo Disease Ontology (Mondo), we conducted a series of expert-driven workshops in collaboration with the team from the Rare Disease Cures Accelerator-Data and Analytics Platform (RDCA-DAP). Specialists in epileptology, genetics, neurodevelopment, biomedical ontology, and patient community advocates systematically reviewed and revised the epilepsy hierarchy in Mondo, aligning it with the International League Against Epilepsy (ILAE) classification system. These updates include reclassification of epilepsy subtypes, including syndromes, age-related epilepsies, and developmental epileptic encephalopathies, resulting in a more granular, standardized, and clinically relevant structure. Mondo now offers an enhanced framework for integrating epilepsy data across resources, enabling improved interoperability and facilitating rare disease research and data curation, with continued efforts underway to further refine and expand this integration.