
This protocol details Dataset Readiness Assessment for Training (DRAFT), a systematic method for determining if a high-dimensional biological dataset is suitable for developing reliable and equitable machine learning models. Standard model validation often fails to detect when performance is driven by spurious correlations within the dataset, leading to irreproducible findings. DRAFT provides a step-by-step methodology to perform a multiaxis assessment of a dataset's integrity before extensive modeling begins. The assessment comprises three basic protocols to probe the dataset's characteristics: (1) evaluating its potential for generalization by testing baseline model performance and stability under rigorous cross-validation; (2) auditing for inherent biases by analyzing model performance and calibration across demographic subgroups; and (3) measuring its capacity for scientific utility by assessing the stability of predictive features. The Cancer Genome Atlas Lung Adenocarcinoma (TCGA-LUAD) dataset serves as a primary case study to demonstrate DRAFT's ability to reveal a dataset's potential to produce fragile, inequitable, and scientifically uninformative models, even when preliminary modeling achieves high performance metrics. This framework provides a necessary template for vetting datasets, ensuring that subsequent computational models built upon them are robust, equitable, and capable of generating true scientific insight. © 2025 Wiley Periodicals LLC. Support Protocol 1: Environment and software installation via Conda Support Protocol 2: TCGA-LUAD case study: Data acquisition and preparation Basic Protocol 1: Generalization audit: Assessing for overfitting and illusory performance Basic Protocol 2: Equity audit: Assessing for hidden stratification and algorithmic bias Basic Protocol 3: Stability audit: Assessing for spurious discoveries and scientific utility.
JBrowse 2 is an open-source genome browser that provides unique features for visualizing syntenic relationships between multiple genomes. This article describes a protocol for setting up synteny views in JBrowse 2, using an assembly-to-assembly whole-genome alignment example. We detail data preparation steps, including the generation and formatting of whole-genome alignment data into formats compatible with JBrowse 2's synteny visualization capabilities, and show the GUI-driven process for setting up interactive synteny views and generating publication-quality figures. This protocol establishes methods for using JBrowse 2 to explore conserved sequences across multiple genomes. © 2025 The Author(s). Current Protocols published by Wiley Periodicals LLC. Basic Protocol: Using JBrowse 2 to show genomic synteny.
BioModelAnalyzer (BMA) is an open-source graphical tool for the development of executable models of protein and gene networks within cells. Based upon the Qualitative Networks formalism, the user can rapidly construct large networks, either manually or by connecting motifs selected from a built-in library. After the appropriate functions for each variable are defined, the user has access to three analysis engines to test the model. In addition to standard simulation tools, BMA includes an interface to the stability-testing algorithm and to a graphical Linear Temporal Logic (LTL) editor and analysis tool. Alongside this, we have developed a novel ChatBot to aid users constructing LTL queries and to explain the interface and run through tutorials. Here we present worked examples of model construction and testing via the interface. As an initial example, we discuss fate decisions in Dictyostelium discoidum and cAMP signaling. We go on to describe the workflow leading to the construction of a published model of the germline of C. elegans. Finally, we demonstrate how to construct simple models from the built-in network motif library. © 2020 by John Wiley & Sons, Inc. Basic Protocol 1: Modeling the signaling network of Dictyostelium discoidum Basic Protocol 2: Modeling the germline progression of Caenorhabditis elegans Basic Protocol 3: Constructing a model of the cell cycle using motifs.
MetaBridge is a web-based tool designed to facilitate the integration of metabolomics with other "omics" data types such as transcriptomics and proteomics. It uses data from the MetaCyc metabolic pathway database and the Kyoto Encyclopedia of Genes and Genomes (KEGG) to map metabolite compounds to directly interacting upstream or downstream enzymes in enzymatic reactions and metabolic pathways. The resulting list of enzymes can then be integrated with transcriptomics or proteomics data via protein-protein interaction networks to perform integrative multi-omics analyses. MetaBridge was developed to be intuitive and easy to use, requiring little to no prior computational experience. The protocols described here detail all steps involved in the use of MetaBridge, from preparing input data and performing metabolite mapping to utilizing the results to build a protein-protein interaction network. © 2020 by John Wiley & Sons, Inc. Basic Protocol 1: Mapping metabolite data using MetaCyc identifiers Basic Protocol 2: Mapping metabolite data using KEGG identifiers Support Protocol 1: Converting compound names to HMDB IDs Support Protocol 2: Submitting mapped genes produced by MetaBridge for protein-protein interaction (PPI) network construction.
Non-coding RNAs are essential for all life and carry out a wide range of functions. Information about these molecules is distributed across dozens of specialized resources. RNAcentral is a database of non-coding RNA sequences that provides a unified access point to non-coding RNA annotations from >40 member databases and helps provide insight into the function of these RNAs. This article describes different ways of accessing the data, including searching the website and retrieving the data programmatically over web APIs and a public database. We also demonstrate an example Galaxy workflow for using RNAcentral for RNA-seq differential expression analysis. RNAcentral is available at https://rnacentral.org. © 2020 The Authors. Basic Protocol 1: Viewing RNAcentral sequence reports Basic Protocol 2: Using RNAcentral text search to explore ncRNA sequences Basic Protocol 3: Using RNAcentral sequence search Basic Protocol 4: Using RNAcentral FTP archive Support Protocol 1: Using web APIs for programmatic data access Support Protocol 2: Using public Postgres database to export large datasets Support Protocol 3: Analyze non-coding RNA in RNA-seq datasets using RNAcentral and Galaxy.
Pharos is an integrated web-based informatics platform for the analysis of data aggregated by the Illuminating the Druggable Genome (IDG) Knowledge Management Center, an NIH Common Fund initiative. The current version of Pharos (as of October 2019) spans 20,244 proteins in the human proteome, 19,880 disease and phenotype associations, and 226,829 ChEMBL compounds. This resource not only collates and analyzes data from over 60 high-quality resources to generate these types, but also uses text indexing to find less apparent connections between targets, and has recently begun to collaborate with institutions that generate data and resources. Proteins are ranked according to a knowledge-based classification system, which can help researchers to identify less studied "dark" targets that could be potentially further illuminated. This is an important process for both drug discovery and target validation, as more knowledge can accelerate target identification, and previously understudied proteins can serve as novel targets in drug discovery. Two basic protocols illustrate the levels of detail available for targets and several methods of finding targets of interest. An Alternate Protocol illustrates the difference in available knowledge between less and more studied targets. © 2020 by John Wiley & Sons, Inc. Basic Protocol 1: Search for a target and view details Alternate Protocol: Search for dark target and view details Basic Protocol 2: Filter a target list to get refined results.
The MPI Bioinformatics Toolkit (https://toolkit.tuebingen.mpg.de) provides interactive access to a wide range of the best-performing bioinformatics tools and databases, including the state-of-the-art protein sequence comparison methods HHblits and HHpred. The Toolkit currently includes 35 external and in-house tools, covering functionalities such as sequence similarity searching, prediction of sequence features, and sequence classification. Due to this breadth of functionality, the tight interconnection of its constituent tools, and its ease of use, the Toolkit has become an important resource for biomedical research and for teaching protein sequence analysis to students in the life sciences. In this article, we provide detailed information on utilizing the three most widely accessed tools within the Toolkit: HHpred for the detection of homologs, HHpred in conjunction with MODELLER for structure prediction and homology modeling, and CLANS for the visualization of relationships in large sequence datasets. © 2020 The Authors. Basic Protocol 1: Sequence similarity searching using HHpred Alternate Protocol: Pairwise sequence comparison using HHpred Support Protocol: Building a custom multiple sequence alignment using PSI-BLAST and forwarding it as input to HHpred Basic Protocol 2: Calculation of homology models using HHpred and MODELLER Basic Protocol 3: Cluster analysis using CLANS.
DisProt is the major repository of manually curated data for intrinsically disordered proteins collected from the literature. Although lacking a stable tertiary structure under physiological conditions, intrinsically disordered proteins carry out a plethora of biological functions, some of them directly arising from their flexible nature. A growing number of scientific studies have been published during the last few decades in an effort to shed light on their unstructured state, their binding modes, and their functions. DisProt makes use of a team of expert biocurators to provide up-to-date annotations of intrinsically disordered proteins from the literature, making them available to the scientific community. Here we present a comprehensive description on how to use DisProt in different contexts and provide a detailed explanation of how to explore and interpret manually curated annotations of intrinsically disordered proteins. We describe how to search DisProt annotations, using both the web interface and the API for programmatic access. Finally, we explain how to visualize and interpret a DisProt entry, p53, a widely studied protein characterized by the presence of unstructured N-terminal and C-terminal regions. © 2020 Wiley Periodicals LLC. Basic Protocol 1: Performing a search in DisProt Support Protocol 1: Downloading options Support Protocol 2: Programmatic access with DisProt REST API Basic Protocol 2: Visualizing and interpreting DisProt entries: the p53 use case Basic Protocol 3: Providing feedback and submitting new intrinsic disorder-related data.
The BioGateway App is a plugin for the Cytoscape network editor, allowing users to interactively build biological networks by querying the Biogateway Resource Description Framework (RDF) triple store. BioGateway contains information from several curated resources including UniProtKB, IntAct, Gene Ontology Annotations, various datasets containing transcription-factor regulatory relations to specific target genes, and more. The BioGateway App facilitates the step-by-step creation of complex SPARQL queries through an intuitive Graphical User Interface, allowing users to build and explore biological interaction networks to assess, among other things, gene regulatory relationships, gene ontology annotations, and protein-protein interactions. As the BioGateway information content is most abundant for human proteins and genes, this article describes the utility of the tool through a series of use cases on these human data, starting from the most basic levels and then detailing applications that address some of the rich complexity of the integrated data. Network refinement and display can be further optimized via the selection and filtering possibilities that the Cytoscape framework provides. The use cases also provide examples to explore network information in other species, as they become supported by BioGateway. © 2020 The Authors. Basic Protocol 1: Introducing a node from the canvas Basic Protocol 2: Introducing a node from the query builder Basic Protocol 3: Exploring molecular relationships between diseases Basic Protocol 4: Find proteins with protein kinase activity involved in a disease and explore the context around them Basic Protocol 5: Exploring the potential downstream effects after targeted inhibition of proteins Support Protocol: Installation of the BioGateway plugin through the Cytoscape App Manager and from source.
The Perseus software provides a comprehensive framework for the statistical analysis of large-scale quantitative proteomics data, also in combination with other omics dimensions. Rapid developments in proteomics technology and the ever-growing diversity of biological studies increasingly require the flexibility to incorporate computational methods designed by the user. Here, we present the new functionality of Perseus to integrate self-made plugins written in C#, R, or Python. The user-written codes will be fully integrated into the Perseus data analysis workflow as custom activities. This also makes language-specific R and Python libraries from CRAN (cran.r-project.org), Bioconductor (bioconductor.org), PyPI (pypi.org), and Anaconda (anaconda.org) accessible in Perseus. The different available approaches are explained in detail in this article. To facilitate the distribution of user-developed plugins among users, we have created a plugin repository for community sharing and filled it with the examples provided in this article and a collection of already existing and more extensive plugins. © 2020 The Authors. Basic Protocol 1: Basic steps for R plugins Support Protocol 1: R plugins with additional arguments Basic Protocol 2: Basic steps for python plugins Support Protocol 2: Python plugins with additional arguments Basic Protocol 3: Basic steps and construction of C# plugins Basic Protocol 4: Basic steps of construction and connection for R plugins with C# interface Support Protocol 4: Advanced example of R Plugin with C# interface: UMAP Basic Protocol 5: Basic steps of construction and connection for python plugins with C# interface Support Protocol 5: Advanced example of python plugin with C# interface: UMAP Support Protocol 6: A basic workflow for the analysis of label-free quantification proteomics data using perseus.
Ten of thousands of open reading frames (ORFs) are hidden within genomes. These alternative ORFs, or small ORFs, have eluded annotations because they are either small or within unsuspected locations. They are found in untranslated regions or overlap a known coding sequence in messenger RNA and anywhere in a "non-coding" RNA. Serendipitous discoveries have highlighted these ORFs' importance in biological functions and pathways. With their discovery came the need for deeper ORF annotation and large-scale mining of public repositories to gather supporting experimental evidence. OpenProt, accessible at https://openprot.org/, is the first proteogenomic resource enforcing a polycistronic model of annotation across an exhaustive transcriptome for 10 species. Moreover, OpenProt reports experimental evidence cumulated across a re-analysis of 114 mass spectrometry and 87 ribosome profiling datasets. The multi-omics OpenProt resource also includes the identification of predicted functional domains and evaluation of conservation for all predicted ORFs. The OpenProt web server provides two query interfaces and one genome browser. The query interfaces allow for exploration of the coding potential of genes or transcripts of interest as well as custom downloads of all information contained in OpenProt. © 2020 The Authors. Basic Protocol 1: Using the Search interface Basic Protocol 2: Using the Downloads interface.
SPAdes-St. Petersburg genome Assembler-was originally developed for de novo assembly of genome sequencing data produced for cultivated microbial isolates and for single-cell genomic DNA sequencing. With time, the functionality of SPAdes was extended to enable assembly of IonTorrent data, as well as hybrid assembly from short and long reads (PacBio and Oxford Nanopore). In this article we present protocols for five different assembly pipelines that comprise the SPAdes package and that are used for assembly of metagenomes and transcriptomes as well as assembly of putative plasmids and biosynthetic gene clusters from whole-genome sequencing and metagenomic datasets. In addition, we present guidelines for understanding results with use cases for each pipeline, and several additional support protocols that help in using SPAdes properly. © 2020 Wiley Periodicals LLC. Basic Protocol 1: Assembling isolate bacterial datasets Basic Protocol 2: Assembling metagenomic datasets Basic Protocol 3: Assembling sets of putative plasmids Basic Protocol 4: Assembling transcriptomes Basic Protocol 5: Assembling putative biosynthetic gene clusters Support Protocol 1: Installing SPAdes Support Protocol 2: Providing input via command line Support Protocol 3: Providing input data via YAML format Support Protocol 4: Restarting previous run Support Protocol 5: Determining strand-specificity of RNA-seq data.
The iSwathX web application processes and normalizes mass spectrometry−based proteomics spectral libraries generated in the data‐dependent acquisition (DDA) approach. These libraries are stored in various proteomics repositories such as PeptideAtlas and NIST, or are user‐generated and provide reference data for data‐independent acquisition (DIA) targeted data extraction and analysis. iSwathX 2.0 can efficiently normalize DDA data from different instruments, gathered at different instances, and make it compatible with specific DIA experiments. Novel functions for parallel processing of DDA libraries and DIA report files, along with various data visualizations, are available in iSwathX 2.0. The step‐by‐step protocols provided here describe how the libraries are uploaded, processed, visualized, and downloaded using various modules of the application. They also provide detailed guidelines on the use of DIA report files for data analysis and visualization. © 2020 Wiley Periodicals LLC.
QIIME 2 is a completely re-engineered microbiome bioinformatics platform based on the popular QIIME platform, which it has replaced. QIIME 2 facilitates comprehensive and fully reproducible microbiome data science, improving accessibility to diverse users by adding multiple user interfaces. QIIME 2 can be combined with Qiita, an open-source web-based platform, to re-use available data for meta-analysis. The following basic protocol describes how to install QIIME 2 on a single computer and analyze microbiome sequence data, from processing of raw DNA sequence reads through generating publishable interactive figures. These interactive figures allow readers of a study to interact with data with the same ease as its authors, advancing microbiome science transparency and reproducibility. We also show how plug-ins developed by the community to add analysis capabilities can be installed and used with QIIME 2, enhancing various aspects of microbiome analyses-e.g., improving taxonomic classification accuracy. Finally, we illustrate how users can perform meta-analyses combining different datasets using readily available public data through Qiita. In this tutorial, we analyze a subset of the Early Childhood Antibiotics and the Microbiome (ECAM) study, which tracked the microbiome composition and development of 43 infants in the United States from birth to 2 years of age, identifying microbiome associations with antibiotic exposure, delivery mode, and diet. For more information about QIIME 2, see https://qiime2.org. To troubleshoot or ask questions about QIIME 2 and microbiome analysis, join the active community at https://forum.qiime2.org. © 2020 The Authors. Basic Protocol: Using QIIME 2 with microbiome data Support Protocol: Further microbiome analyses.
IUPred2A is a combined prediction tool designed to discover intrinsically disordered or conditionally disordered proteins and protein regions. Intrinsically disordered regions exist without a well-defined three-dimensional structure in isolation but carry out important biological functions. Over the years, various prediction methods have been developed to characterize disordered regions. The existence of disordered segments can also be dependent on different factors such as binding partners or environmental traits like pH or redox potential, and recognizing such regions represents additional computational challenges. In this article, we present detailed instructions on how to use IUPred2A, one of the most widely used tools for the prediction of disordered regions/proteins or conditionally disordered segments, and provide examples of how the predictions can be interpreted in different contexts. © 2020 The Authors. Basic Protocol 1: Analyzing disorder propensity with IUPred2A online Basic Protocol 2: Analyzing disordered binding regions using ANCHOR2 Support Protocol 1: Interpretation of the results Basic Protocol 3: Analyzing redox-sensitive disordered regions Support Protocol 2: Download options Support Protocol 3: REST API for programmatic purposes Basic Protocol 4: Using IUPred2A locally.
Ggtree is an R/Bioconductor package for visualizing tree-like structures and associated data. After 5 years of continual development, ggtree has been evolved as a package suite that contains treeio for tree data input and output, tidytree for tree data manipulation, and ggtree for tree data visualization. Ggtree was originally designed to work with phylogenetic trees, and has been expanded to support other tree-like structures, which extends the application of ggtree to present tree data in other disciplines. This article contains five basic protocols describing how to visualize trees using the grammar of graphics syntax, how to visualize hierarchical clustering results with associated data, how to estimate bootstrap values and visualize the values on the tree, how to estimate continuous and discrete ancestral traits and visualize ancestral states on the tree, and how to visualize a multiple sequence alignment with a phylogenetic tree. The ggtree package is freely available at https://www.bioconductor.org/packages/ggtree. © 2020 by John Wiley & Sons, Inc. Basic Protocol 1: Using grammar of graphics for visualizing trees Basic Protocol 2: Visualizing hierarchical clustering using ggtree Basic Protocol 3: Visualizing bootstrap values as symbolic points Basic Protocol 4: Visualizing ancestral status Basic Protocol 5: Visualizing a multiple sequence alignment with a phylogenetic tree.
Visualizing protein data remains a challenging and stimulating task. Useful and intuitive visualization tools may help advance biomolecular and medical research; unintuitive tools may bar important breakthroughs. This protocol describes two use cases for the CellMap (http://cellmap.protein.properties) web tool. The tool allows researchers to visualize human protein-protein interaction data constrained by protein subcellular localizations. In the simplest form, proteins are visualized on cell images that also show protein-protein interactions (PPIs) through lines (edges) connecting the proteins across the compartments. At a glance, this simultaneously highlights spatial constraints that proteins are subject to in their physical environment and visualizes PPIs against these localizations. Visualizing two realities helps in decluttering the protein interaction visualization from "hairball" phenomena that arise when single proteins or groups thereof interact with hundreds of partners. © 2019 The Authors. Basic Protocol 1: Visualizing proteins and their interactions on cell images Basic Protocol 2: Displaying all interaction partners for a protein.
We present Mothulity-a novel interface for Mothur, a well-established tool for 16S/ITS biodiversity analysis. Although Mothur is a well-documented and virtually complete software suite, its proper execution might be challenging for first-time users, and editing the Mothur batch scripts is time consuming even for experienced users. Mothur produces little to no graphical output, leaving the generation of plots to the user. Mothulity minimizes the chance of human error through a minimalistic yet powerful interface, with most of the analysis parameters predefined or adjusted automatically. Time spent on running the analysis is drastically reduced, since Mothulity produces an HTML report with publication-quality figures. Finally, Mothulity can be conveniently used with the SLURM workload manager, and is thereby suitable for a range of computing facilities. © 2020 by John Wiley & Sons, Inc. Basic Protocol 1: Standard operational procedure (SOP) Basic Protocol 2: Generating report on pre-processed data.
The Molecular INTeractions Database (MINT) is a public database designed to store information about protein interactions. Protein interactions are extracted from scientific literature and annotated in the database by expert curators. Currently (October 2019), MINT contains information on more than 26,000 proteins and more than 131,600 interactions in over 30 model organisms. This article provides protocols for searching MINT over the Internet, using the new MINT Web Page. © 2020 by John Wiley & Sons, Inc. Basic Protocol 1 : Searching MINT over the internet Alternate Protocol : MINT visualizer Basic Protocol 2 : Submitting interaction data
MathIOmica is a package for bioinformatics, written in the Wolfram language, that provides multiple utilities to facilitate the analysis of longitudinal data generated from omics experiments, including transcriptomics, proteomics, and metabolomics data, as well as any generalized time series. MathIOmica uses Mathematica's notebook interface, wherein users can import longitudinal datasets, carry out quality control and normalization, generate time series, and classify temporal trends. MathIOmica provides spectral methods based on periodograms and autocorrelations for automatically detecting classes of temporal behavior and allowing the user to visualize collective temporal behavior, and also assess biological significance through Gene Ontology and pathway enrichment analyses. MathIOmica's time-series classification methods address common issues including missing data and uneven sampling in measurements. As such, the software is ideally suited for the analysis of experimental data from individualized profiling of subjects, can facilitate analysis of data from the emerging field of individualized health monitoring, and can detect temporal trends that may be associated with adverse health events. In this article, we import a transcriptomics (RNA-sequencing) dataset collected over multiple timepoints and generate time series for each transcript represented in the data. We classify the time series to identify classes of significant temporal trends (using autocorrelations). We assess statistical significance cutoffs in the classification by generating null distributions using randomly resampled time series. We then visualize the significant trends in heatmaps and assess biological significance using enrichment analyses. Finally, we visualize pathway results for statistically significant pathways of interest. © 2019 by John Wiley & Sons, Inc. Basic Protocol: Time series analysis of transcriptomics expression dataset.