
This multi-hop biomedical QA system combines dense retrieval with a bounded agentic workflow for the MedHopQA (BioCreative IX) task, which requires short, entity-centric answers from a Wikipedia-derived corpus evaluated via exact match with normalization. Our approach replaces prior prompt decomposition with a controlled state machine that routes questions into four strategies (direct, definition, intersection, and multi-hop). The pipeline performs dense retrieval with cross-encoder reranking, executes up to three hops of intermediate entity extraction when needed, and applies answer validation with a bounded query-repair loop. On the MedHopQA test set (N = 1000), MedHopper achieves 0.55 Exact Match (MedHopQA evaluation metric) under official CodaBench evaluation. Ablation confirms capped multi-hop execution, reranking, and query repair as primary contributors; validation modules provide smaller consistent gains. Analysis reveals persistent challenges under exact-match scoring, including answer-type inconsistency (chromosome vs cytoband; yes/no vs symptom) and surface variation across plausible lexicalizations. All experiments were run on a single consumer GPU (RTX 5080) with no API dependencies, producing deterministic outputs at zero marginal cost per query.
Understanding the regulatory impact of non-coding genetic variants remains a major challenge in human genetics. Here, we present scRiskDB, a comprehensive and user-friendly database that maps genetic risk variants to their downstream regulatory elements, target genes, and relevant cell types at single-cell resolution. By integrating genome-wide association studies (GWAS) with single-cell datasets across 45 tissues and developmental stages, scRiskDB implements a variant-to-function framework that systematically outlines potential regulatory cascades from single nucleotide variants to cell-specific risk mechanisms. This multi-layered design allows users to explore trait-associated regulatory architectures across cell types and developmental stages. The platform provides interactive, multi-level visualizations and curated results, facilitating hypothesis generation and mechanistic insights into disease aetiology.
Butterflies have been a model system for studying the evolution of colour. This is partly due to their complex patterns that reflect human-visible (VIS) and ultraviolet (UV) light, which are perceived by conspecifics and predators. Many studies have sourced data from publicly available images, but most of these images only consider the visible spectrum of light. Including the UV spectrum is crucial for fully understanding the evolution of butterfly morphology and behavioural ecology. Here we provide standardized images (VIS and UV) of over 4 000 individuals from 16 communities of Australian butterflies. These communities represent different climates and urbanization levels spanning over 2 500 km. The dataset contains at least one individual of 125 different species from five families, constituting over one quarter of Australian butterfly diversity. In addition to photographs, we provide spectral measurements of butterfly wings for at least one individual of each species and sex, and Cytochrome Oxidase subunit 1 (CO1) sequences of 1 635 individuals. All these data are accessible in Zenodo and an associated R package simplifies the download of subsets of the database. This database will be of use to evolutionary biologists and ecologists interested in a broad range of topics related to phenotypic variation.
Querying the RDF Portal knowledge graph maintained by DBCLS-which aggregates ~60 life-science databases-requires proficiency in both SPARQL and database-specific RDF schemas, placing this resource beyond the reach of most researchers. Large Language Models (LLMs) can, in principle, translate natural-language questions into executable SPARQL, but without schema-level context, they frequently fabricate non-existent predicates or fail to resolve entity names to database-specific identifiers. We present TogoMCP, a system that recasts the LLM as a protocol-driven inference engine orchestrating specialized tools via the Model Context Protocol (MCP). Two mechanisms are essential to its design: (i) the MIE (Metadata-Interoperability-Exchange) file, a concise YAML document that dynamically supplies the LLM with each target database's structural and semantic context at query time; and (ii) a two-stage workflow separating entity resolution via external REST APIs from schema-guided SPARQL generation. On a benchmark of 50 biologically grounded questions spanning five types and 23 databases, TogoMCP achieved a large improvement over an unaided baseline (Cohen's $d = 1.82$, Wilcoxon $P \lt .001$), with win rates exceeding 80% for question types with precise, verifiable answers. An ablation study shows that all component configurations deliver significant improvements, with MIE schema files providing the largest marginal contribution on mean per-question score ($\Delta = +0.50$ relative to a no-MIE condition, two-sided Wilcoxon $P = .067$; 90% bootstrap CI $[+0.04,\,\,+0.94]$ excludes zero); a one-line instruction to load the relevant MIE file recovers the same mean improvement as a full procedural protocol, while the protocol additionally reduces downside risk (loss rate 1.6% vs. 4.8%, Fisher $P = .036$). These results suggest a general design principle: concise, dynamically delivered schema context is more valuable than complex orchestration logic for mean score performance, while procedural guidance plays a complementary role in narrowing variance. Database URL: https://togomcp.rdfportal.org/.
The family Cyprinidae represents the most taxonomically diverse group of freshwater fishes, encompassing over 3 000 species of ecological and economic importance in aquaculture, conservation, and ecological monitoring. Simple sequence repeats (SSRs), also known as microsatellites, are highly informative molecular markers widely used for genetic diversity analysis, population structure assessment, and marker-assisted breeding. However, comprehensive genome-wide SSR resources for cyprinids remain limited. Existing databases, such as FishMicrosat, provide restricted taxonomic coverage and lack standardized chromosome-level datasets suitable for comparative genomic analyses. To address this gap, we developed CypriSSR, a genome-wide SSR database encompassing 11 representative cyprinid species. Chromosome-level genome assemblies were retrieved from NCBI. SSR loci were identified using MISA, and primer pairs were designed using Primer3. In total, over 7.8 million SSR loci were identified. Mononucleotide repeats were the most abundant class (39-53%), followed by dinucleotide and trinucleotide motifs. SSRs were predominantly distributed in genic regions (54%-71% across species), suggesting potential functional roles. Each database entry includes repeat type, genomic coordinates, primer sequences, melting temperatures, and predicted PCR product sizes. The CypriSSR web interface enables flexible querying of SSR markers based on species, chromosome, motif class, and genomic location, and supports sequence similarity searches through an integrated BLAST module along with data export options. CypriSSR provides a comprehensive and standardized multi-species microsatellite resource for cyprinid genomes and supports applications in population genetics, molecular breeding, and conservation genomics. Database URL: http://46.202.167.198/fishssr/.
Fast-neutron mutagenesis creates diverse genome-wide mutations, providing a powerful tool for crop functional genomics. Here, we present an expanded genomic and phenotypic analysis of 3268 fast-neutron (FN)-induced mutant rice lines (Oryza sativa L. cv. Kitaake). All FN lines were whole-genome sequenced, and mutations were identified by alignment in the Nipponbare and KitaakeX reference genomes. We cataloged over 428,000 mutations affecting 78.49% of Nipponbare genes and 70.38% of KitaakeX genes. In silico expression analysis indicates that 575 non-mutated Nipponbare genes are highly expressed and likely essential for viability. Each mutant carries, on average, 68.5 mutations in the Nipponbare alignments or 63.2 mutations for KitaakeX alignments, distributed randomly across all 12 chromosomes with no evident hotspots. FN lines have approximately 8.5% fewer mutations when using the KitaakeX alignment, underscoring the unique contributions of each reference genome and the importance of utilizing both for comprehensive mutation discovery. The majority of mutations are small deletions and single-base substitutions, with deletions predominating in their effect on genes. We found that 74.4% of all transcription factor Nipponbare genes were mutated at least once. Phenotypic characterization of over 2700 lines revealed a broad spectrum of variation in core agronomic traits (heading date, tiller number, plant height, panicle weight, seed yield components) and other morphological variants of interest. The integration of genomic and phenotypic data through the KitBase platform enabled the identification of candidate genes for several traits of interest. The KitBase website (https://kitbase.ucdavis.edu) has been updated to provide open access to all mutation data and seed stocks, as well as an intuitive query interface, facilitating forward and reverse genetic analyses in rice. This expanded resource enriches the rice functional genomics toolkit and highlights the value of coupling high-density mutation mapping with phenotypic data for rapid gene discovery and crop improvement.
Gut dysbiosis is widely recognized as a contributor to autoimmune diseases, as it can lead to the expression of microbial antigens that disrupt immune regulation through specific molecular mechanisms. However, existing resources do not systematically link gut microbial antigen sequences to the specific autoimmune mechanisms through which they act. Here, we present GUTAID (Gut Microbes in Autoimmune Disorders), a literature-curated database of gut microbial antigens annotated with experimentally supported autoimmune mechanisms. Peer-reviewed studies published from October 1970 to September 2024 were manually screened, yielding 73 potential antigens that operate through nine molecular mechanisms, including protein citrullination, epitope spreading, molecular mimicry, and immune modulation, amongst others. The corresponding protein sequences were retrieved from UniProtKB, and redundancy was removed with MMseqs2. For the database implementation, data were delivered through a lightweight LAMP (Linux-Apache-MySQL/MariaDB-PHP) stack with server-side HTML/Bootstrap rendering, MySQL indexing, and HTTPS-secured downloads. Users can browse, keyword-search, or bulk-download sequence archives via a five-tab interface (Home, Downloads, Search, Team, and About). GUTAID thus enables mechanism-oriented exploration of gut microbial antigens and supports downstream biomarker and therapeutic discovery in autoimmune research. Database URL: https://gutaid.mgdiscoverylab.com/.
Animal venomics is a growing field of research with evolutionary and biotechnological significance. Yet, fundamental questions regarding the origin, diversification, and bioactivity of venoms remain unresolved. Here, we analysed venom tissue-related data curated in Tox-Prot, currently the most comprehensive database for animal venoms, across three snapshots spanning two decades (2005, 2015, 2025). We assessed the taxonomic landscape related to Tox-Prot entries, sequence length distribution, protein family abundances, and habitat-specific venom patterns. Our results consistently show that snakes, spiders, cone snails, and scorpions, along with their associated protein families such as Phospholipase A2, Snake Three Finger Toxin and Long 4(C-C) scorpion toxin family, dominate across Tox-Prot. Nevertheless, the taxonomic and protein family diversity has been steadily increasing, with 503 new species and 188 new protein families added by 2025 compared to the 2005 dataset. Marine species account for 16%-20% of total species, of which 63%-85% are Neogastropoda reflecting limited marine species diversity coverage; likewise, terrestrial taxa are disproportionately represented by Squamata (39%) and Hymenoptera (20%) relative to their natural diversities of 2.3% and 50.26%, respectively. At the molecular level, half of all entries correspond to mature peptide sequences of 26-75 amino acids, featuring three to four disulfide bridges and C-terminal amidations as the most frequently recorded post-translational modifications. Further, protein language model embeddings infer taxonomical diverse peptide and delimited large enzyme clusters. Our study maps two decades of venom data diversification, reflecting both the field's rapid expansion and the need for robust integrated datasets to propel and disseminate that knowledge.
Food allergies represent a significant and growing global challenge. However, researchers and risk assessors still lack a dedicated and comprehensive resource that centralizes food allergen information. As a result, obtaining the necessary structural, biochemical, and immunological data often requires time-consuming searches across multiple heterogeneous databases, hindering efficient analysis and slowing scientific progress. The Food Allergen Database (FAD) merges biochemical, structural, immunological, and functional properties of all known food allergens and isoallergens (1168 entries). It is an essential resource that facilitates the extraction of all information about food allergens.
The nucleolus is a well-characterized sub-nuclear compartment primarily responsible for ribosomal RNA (rRNA) synthesis and ribosome biogenesis. In addition to these canonical functions, it plays roles in a variety of other cellular processes. Despite its importance, the contributions of the nucleolus to fungal development and pathogenicity remain poorly understood, especially in filamentous fungi. The structure and function of the nucleolus are regulated by numerous proteins that either reside within it or shuttle between the nucleolus and nucleoplasm, often guided by short amino acid motifs known as nucleolar localization signals (NoLSs). However, comprehensive resources cataloguing nucleolar proteins and their localization signals in fungi are currently lacking, hindering systematic investigations of nucleolar function across species. To address this gap, we developed FuNGI (Fungal Nucleolar Genomic Inventory), a web-based interactive database for the exploration of fungal proteins containing predicted nucleolar localization signals. The current version of FuNGI includes proteins containing predicted nucleolar localization signals and their associated NoLSs across 769 fungal proteomes spanning eight phyla. The database offers a user-friendly interface that enables browsing, retrieval, and comparison of NoLS-containing proteins across multiple species. Each entry integrates sequence-based predictions and functional annotations to support comparative and functional analyses. To our knowledge, FuNGI is the first comprehensive and interactive database dedicated to fungal proteins containing predicted nucleolar localization signals. By enabling systematic and cross-species analyses, FuNGI provides a valuable resource for advancing our understanding of fungal nucleoli and their roles in fungal biology and pathogenicity.
Cancer metastasis involves complex molecular mechanisms that cannot be fully explained by individual gene expression profiles. Previous studies have shown that correlations between RNA and miRNA expression can capture metastatic behaviour more effectively than expression of individual genes. However, no publicly available databases provide systematic analysis of RNA-miRNA correlations specific to cancer metastasis. We developed an efficient computational method to identify differential correlations between miRNAs and RNAs that are specific to individual tumour samples. Using data from The Cancer Genome Atlas (TCGA), we computed differential correlations for tumour samples across 9 cancer types and 21 metastatic sites, encompassing ~200 million RNA-miRNA pairs. Statistical analysis identified RNA-miRNA pairs with site-specific correlations using Mann-Whitney U-tests. MetaCancerDB contains RNA-miRNA correlation networks for 9 primary cancer types and 21 metastatic sites. Site-specific correlations showed distinct patterns, with lung metastasis displaying the most conserved correlations across cancer types. Survival analysis revealed that specific RNA-miRNA pairs are prognostic for patient outcomes in a metastatic site-dependent manner. MetaCancerDB provides a comprehensive resource for exploring RNA-miRNA correlations in cancer metastasis. The database enables researchers to identify molecular signatures specific to metastatic sites and can serve as a foundation for developing predictive biomarkers. MetaCancerDB is freely available for academic purposes.
The reactive thiol group of cysteine (Cys) acts as a nucleophile and undergoes many cysteine post-translational modifications (Cys-PTMs). Cys-PTMs, called protein redox switch, contribute to various cellular and physiological processes, including reactive oxygen species (ROS)-induced signalling, ROS mitigation, and scavenging. Consolidation of Cys-PTMs into a database would facilitate the mechanistic elucidation of biological processes and therapeutic applications. The existing databases store information on cysteine motifs, oxidation states, a few of the Cys-PTMs, etc., specific to species or kingdoms, and lack general applicability. There was no mention of the impact of the protein microenvironments and cellular localizations on the Cys-PTMs. The current study reports a database containing 7 Cys-PTMs (disulphide, S-nitrosylation, S-palmitoylation, S-glutathionylation, S-sulphenylation, metal-binding, and thioether), 11 features, 33 06 395 UniProt IDs, and 1 14 56 639 cysteine residues, across the taxonomy, encompassing cellular organelles, enzyme classes, sequence motifs, protein structures, and microenvironments. The maximum number of cysteine residues is reported here compared to 16 contemporary cysteine databases. Twenty-one types of metal-binding cysteines and thioether modifications are reported for the first time. Enzyme classes, cellular localization, taxonomic preferences, and microenvironment around Cys-PTMs were systematically analysed and curated, indicating the pathogenic involvement of those Cys-PTMs. The database has a web access (https://cysdbase.bits-hyderabad.ac.in/) and a programmatic access via GitHub link (https://github.com/devhimd19/CysDBase). The query inputs to the repositories are UniProt ID, biological pathway, location, or genus name. Query outputs are 11 biological features, namely, protein name, Cys-PTMs, cysteine residue number, cysteine sequence motif, cell organelle, biological pathway, protein microenvironment (buried fraction and relative hydrophobicity [rHpy]), EC number and enzyme class, secondary structure, organism, and PubMed ID.
Determining the functional consequence of missense mutations acquired in the development of cancer is critical to the understanding of the evolution and the therapeutic vulnerabilities of an individual tumour. Several million missense mutations associated with cancer have been reported across different databases with little functional annotation accompanying each mutation. We have designed the MOKCa-3D database, (https://bioinformaticslab.sussex.ac.uk/MOKCa-3D/) to enable the contextualization and interpretation of cancer somatic missense mutations, including the structural impact of the mutation on the 3D structure, and whether the mutation results in a gain or loss of the protein's function. For each protein, a sequence feature viewer enables interactive visualization of the amino acid sequence, missense mutations, post-translational modification sites, protein domains, active sites, binding sites, protein-protein interaction sites, and mutational frequency. The mutation-level page concisely presents functional insights for each individual mutation, and an interactive MOL* viewer highlights mutated residue on an AlphaFold protein structural model. The SAAP structural impact analysis pipeline was used to identify the structural impact of the mutation. MOKCa-3D concisely presents functional insights and structural impacts of cancer somatic missense mutations enabling users to interpret their functional consequences. It is freely accessible and easy to navigate, making it usable by the widest range of researchers.
Microbial communities associated with desert plants play a pivotal role in enhancing host survival under extreme environmental stressors, including drought, salinity, and nutrient limitation. The Desert Plant Endophyte Microbial Collection is one of the largest curated repositories of 2500 cultivable endophytic bacteria isolated from 23 native desert plant species across Saudi Arabia, Jordan, and Pakistan. Representing a broad spectrum of arid microhabitats from inland deserts and mountain wadis to coastal mangroves and date palm oases, the collection supports integrative studies on microbial ecology and plant-microbe interactions in water-limited ecosystems. A central component of this initiative is the Desert Plant Endophyte Genome Database, which currently hosts whole-genome sequences of 534 endophytic bacterial isolates annotated with extensive ecological metadata, assembly statistics, functional traits, and host associations. The database interface provides tools for genome exploration, metadata filtering, and functional gene mining, enabling users to identify taxa and traits of agronomic interest, particularly for applications in sustainable agriculture and sustainable desert revegetation. By combining genomic, ecological, and functional data, the Desert Plant Endophyte Genome Database serves as a foundational platform for the development of targeted microbial inoculants and fosters data-driven research into desert microbiomes and plant resilience mechanisms.
The safety and modernization of traditional Chinese medicine (TCM) are significant concerns for human beings. Recently, the adverse effects caused by the use of certain TCMs have been frequently reported. Although TCMs may cause toxic reactions, they also play critical roles in treating multiple complex diseases. Therefore, toxicity research is urgently needed for the safe usage of TCMs. However, existing databases for TCMs primarily focus on the pharmacological effects of TCMs, with limited attention to the toxicity. They neither distinguish the toxic effects of formulas, herbs, and ingredients, nor classify and summarize targets for specific toxic manifestations, or assemble evidence from previous studies. We developed TCMToxDB, a comprehensive database that focuses on the toxicity and safe usage of TCMs. TCMToxDB systematically integrates and analyses the research results of toxic TCMs, offering users diverse information acquisition and analysis services. In addition, it assembles five canonical herb-target and ingredient-target interaction prediction algorithms with different advantages, which support the prediction of toxic targets of herbs and ingredients to empower the toxicity research of TCMs and to meet users' personal needs. TCMToxDB is accessible at https://www.sdu-idea.cn/TCMToxDB.
Scientific articles have been searched for experimentally established critical pharmacokinetic parameter-renal clearance. The main source of the documents was PubMed database, considered one of the most important and comprehensive repositories of biomedical literature. After manual data collection and thorough quality check database presenting human renal clearance values of exogenous substances was developed. After collecting the data, preliminary processing and simple analysis were carried out. The final database contains over 1700 experimental observations from 761 scientific articles and covers over 500 unique substances studied. Database URL: doi: https://10.17632/3427x3wzzc.2
The rapid progress in tumour genome sequencing has created a need for bioinformatics tools to interpret the clinical significance of detected variants. VarStack² integrates information from several publicly available resources, including the Catalogue of Somatic Mutations in Cancer (COSMIC), ClinVar, cBioPortal, UCSC Genome Browser, and ClinicalTrials.gov, CIViC and presents it through a user-friendly interface. VarStack2 simplifies the process of retrieving data, saving users significant time compared to manually navigating each database individually. Users can input a variant by specifying a gene symbol, amino acid change, and coding sequence change, with the option to search tumour-specific studies in cBioPortal alongside their primary query. Results are organized into separate sections and can be exported in CSV format for further analysis. Additionally, VarStack2 offers a smart search feature that suggests variants for the gene of interest based on its database search results. These features make VarStack2 a useful tool for scientists and clinicians by enhancing the variant interpretation process and integrating somatic variant information into workflows. VarStack² is freely available at http://varstack.brown.edu/.
Metabolic dysfunction-associated steatohepatitis (MASH), formerly called NASH, is a progressive form of fatty liver disease closely linked to metabolic dysregulation. It has become a leading cause of liver failure and liver transplantation worldwide. Despite its growing clinical significance, there is currently no centralized, publicly accessible, cross-cohort and cross-species transcriptomic database specifically dedicated to MASH. To address this gap, we developed MASH-GA (MASH-Genomic Atlas; https://mashga.com), the first and only database in the fatty liver research field that systematically integrates transcriptomic data across human cohorts and mouse models. MASH-GA contains 45 human transcriptomic datasets (2740 samples) from 17 countries and 126 mouse datasets (1950 samples) covering major MASH models. The platform provides a range of unique features, including integrated weighted gene coexpression network analysis (WGCNA), cross-model classification and comparison of mouse datasets, multigene correlation analysis and pairwise coexpression exploration, and interactive visualizations (e.g. boxplots and correlation plots). Human and mouse data are presented in separate, intuitively navigable modules. As the only standardized, multicohort, cross-species MASH transcriptomic database currently available, MASH-GA enables reproducible data exploration, informed model selection, and the identification of regulatory modules. It is therefore a valuable resource for researchers in hepatology, metabolic disease, and systems biology. Database URL: https://mashga.com.
Dysregulated cell-cell communication (CCC) is increasingly recognized as a driver of brain disease pathology, contributing to neuroinflammation, synaptic dysfunction, and neurodegeneration. Nevertheless, existing resources remain limited in brain specificity, regional coverage, and functional annotation. To address this gap, we develop the Brain Disease Cell-cell communication Database (BDCD), the first comprehensive resource focused on CCC networks across major brain diseases. BDCD integrates 38 manually curated datasets, comprising 8 519 425 single cells from single-cell RNA-seq studies and 140 744 spots from spatial transcriptomic maps, spanning 14 brain regions and 13 canonical cell types covering Alzheimer's disease, Parkinson's disease, schizophrenia, bipolar disorder, and multiple sclerosis. BDCD reconstructs more than 495 000 ligand-receptor interaction events and links them to structural features, genetic associations, pathways, and therapeutic modulators, including 6100 single nucleotide polymorphisms from genome-wide association studies, 3350 drugs, and 72 477 allosteric modulators. This comprehensive atlas enables cross-disease comparison and supports dynamic hypothesis generation, providing a foundation for mechanistic insights and therapeutic discovery. BDCD is publicly available at https://bioinfo.uth.edu/bdcd/. Database URL: https://bioinfo.uth.edu/bdcd/.