Although cell type annotation has become an integral part of single-cell analysis workflows, the assessment of computational annotations remains challenging. Many annotation tools transfer labels from an annotated reference dataset to a new query dataset of interest, but blindly transferring labels from one dataset to another has its own set of challenges. Often enough there is no perfect alignment between datasets, especially when transferring annotations from a healthy reference atlas for the discovery of disease states. We present scDiagnostics, a new open-source software package that facilitates the detection of complex or ambiguous annotation cases that may otherwise go unnoticed, thus addressing a critical unmet need in current single-cell analysis workflows. scDiagnostics is equipped with novel diagnostic methods that are compatible with all major cell type annotation tools. We demonstrate that scDiagnostics reliably detects complex or conflicting annotations using both carefully designed simulated datasets and diverse real-world single-cell datasets. Our evaluation demonstrates that scDiagnostics reliably identifies misleading annotations that systematically distort downstream analysis and interpretation and that would otherwise remain undetected.
The National Health and Nutrition Examination Survey (NHANES) provides extensive public data on demographics, health, and nutrition, collected in 2-year cycles since 1999. Although invaluable for epidemiological and health-related research, the complexity of NHANES data, involving numerous files and disjoint metadata, makes accessing, managing, and analysing these datasets challenging. This paper presents a reproducible computational environment built upon Docker containers, PostgreSQL databases, and R/RStudio, designed to streamline NHANES data management, facilitate rigorous quality control, and simplify analyses across multiple survey cycles. We introduce specialized tools, such as the enhanced nhanesA R package and the phonto R package, to provide fast access to data, to help manage metadata, and to handle complexities arising from questionnaire design and cross-cycle data inconsistencies. Furthermore, we describe the Epiconnector platform, established to foster collaborative sharing of code, analytical scripts, and best practices, which taken together, can significantly enhance the reproducibility, extensibility, and robustness of scientific research using NHANES data.
Summary:AlphaMissense is an AI model from Google DeepMind that predicts the pathogenicity of every possible missense mutation in the human proteome. We present AlphaMissenseR, an R/Bioconductor package that facilitates performant and reproducible access to these predictions and that provides functionality for analysis, visualization, validation, and benchmarking. AlphaMissenseR integrates with Bioconductor facilities for genomic region analysis, and provides multi-level visualization and interactive exploration of variant pathogenicity in a genome browser and on 3D protein structures. In addition, AlphaMissenseR integrates with major clinical and experimental variant databases for contrasting predicted and clinically derived pathogenicity scores, and for systematic benchmarking of existing and new variant effect prediction methods across a large collection of deep mutational scanning assays. Availability and implementation:AlphaMissense data resources are distributed under the CC-BY 4.0 license and the AlphaMissenseR package is available from Bioconductor (https://bioconductor.org/packages/AlphaMissenseR) under the Artistic 2.0 license.
The National Health and Nutrition Examination Survey provides comprehensive data on demographics, sociology, health and nutrition. Conducted in 2-year cycles since 1999, most of its data are publicly accessible, making it pivotal for research areas like studying social determinants of health or tracking trends in health metrics such as obesity or diabetes. Assembling the data and analyzing it presents a number of technical and analytic challenges. This paper introduces the nhanesA R package, which is designed to assist researchers in data retrieval and analysis and to enable the sharing and extension of prior research efforts. We believe that fostering community-driven activity in data reproducibility and sharing of analytic methods will greatly benefit the scientific community and propel scientific advancements.Database URL: https://github.com/cjendres1/nhanes
In many fields, research progress may be hindered by indefiniteness of language used to describe experimental conditions and outcomes. Harmonization of data resources generated by independent groups is important for integrative analysis. Adoption of formal ontologies and vocabularies for experiment annotation should help with harmonization tasks, but the use of ontologies also suffers from a lack of definiteness. In this study we explore how natural language characterization of human diseases coupled with ontologic mapping of study outcome terminology can be used to integrate information from multiple studies of genetic origins of disease risk. Open source tools and workflows are presented. This work exposes areas for improvement in tooling for data harmonization, which is a fundamental requirement for efficient research progress. ### Competing Interest Statement The authors have declared no competing interest.
Clinical and biological characteristics of patients evaluated by gene expression profiling; probe sets differentially expressed in molecular groups
A key challenge in the study of rare disease genetics is assembling large case cohorts for well-powered studies. We demonstrate the use of self-reported diagnosis data to study rare diseases at scale. We performed genome-wide association studies (GWAS) for 33 rare diseases using self-reported diagnosis phenotypes and re-discovered 29 known associations to validate our approach. In addition, we performed the first GWAS for Duane retraction syndrome, vestibular schwannoma and spontaneous pneumothorax, and report novel genome-wide significant associations for these diseases. We replicated these novel associations in non-European populations within the 23andMe, Inc. cohort as well as in the UK Biobank cohort. We also show that mixed model analyses including all ethnicities and related samples increase the power for finding associations in rare diseases. Our results, based on analysis of 19,084 rare disease cases for 33 diseases from 7 populations, show that large-scale online collection of self-reported data is a viable method for discovery and replication of genetic associations for rare diseases. This approach, which is complementary to sequencing-based approaches, will enable the discovery of more novel genetic associations for increasingly rare diseases across multiple ancestries and shed more light on the genetic architecture of rare diseases.
We propose computational methods and statistical paradigms to explore the relationships between phenotypic data and cellular organizational units, such as multi-protein complexes or pathways. Indeed, while proteins are often the primary unit used by cells to carry out the many different functions that the cell requires for life, they seldom accomplish important tasks alone, but rather assemble into organizational units. Recent studies suggest that some control of phenotype can be usefully attributed to multi-protein complexes rather than genes (Deutschbauer et al., 2005; Spirin et al., 2006) and hence may help provide elucidation of the underlying roles or mechanisms that directly control changes in phenotype.
Science depends on collaboration, result reproduction, and the development of supporting software tools. Each of these requires careful management of software versions. We present a unified model for installing, managing, and publishing software contexts in R. It introduces the package manifest as a central data structure for representing version specific, decentralized package cohorts. The manifest points to package sources on arbitrary hosts and in various forms, including tarballs and directories under version control. We provide a high-level interface for creating and switching between side-by-side package libraries derived from manifests. Finally, we extend package installation to support the retrieval of exact package versions as indicated by manifests, and to maintain provenance for installed packages. The provenance information enables the user to publish libraries or sessions as manifests, hence completing the loop between publication and deployment. We have implemented this model across two software packages, switchr and GRANbase, and have released the source code under the Artistic 2.0 license.
Motivation Variant calling is the complex task of separating real polymorphisms from errors. The appropriate strategy will depend on characteristics of the sample, the sequencing methodology and on the questions of interest. Results We present VariantTools, an extensible framework for developing and testing variant callers. There are facilities for reproducibly tallying, filtering, flagging and annotating variants. The tools are extensible, modular and flexible, so that they are tunable to particular use cases, and they interoperate with existing analysis software so that they can be embedded in established work flows. Availability and implementation VariantTools is available from http://www.bioconductor.org/. Contact michafla@gene.com Supplementary information Supplementary data are available at Bioinformatics online.
Analysis of splice variants from short read RNA-seq data remains a challenging problem. Here we present a novel method for the genome-guided prediction and quantification of splice events from RNA-seq data, which enables the analysis of unannotated and complex splice events. Splice junctions and exons are predicted from reads mapped to a reference genome and are assembled into a genome-wide splice graph. Splice events are identified recursively from the graph and are quantified locally based on reads extending across the start or end of each splice variant. We assess prediction accuracy based on simulated and real RNA-seq data, and illustrate how different read aligners (GSNAP, HISAT2, STAR, TopHat2) affect prediction results. We validate our approach for quantification based on simulated data, and compare local estimates of relative splice variant usage with those from other methods (MISO, Cufflinks) based on simulated and real RNA-seq data. In a proof-of-concept study of splice variants in 16 normal human tissues (Illumina Body Map 2.0) we identify 249 internal exons that belong to known genes but are not related to annotated exons. Using independent RNA samples from 14 matched normal human tissues, we validate 9/9 of these exons by RT-PCR and 216/249 by paired-end RNA-seq (2 x 250 bp). These results indicate that de novo prediction of splice variants remains beneficial even in well-studied systems. An implementation of our method is freely available as an R/Bioconductor package [Formula: see text].
The Nrf2 pathway is frequently activated in human cancers through mutations in Nrf2 or its negative regulator KEAP1. Using a cell-line-derived gene signature for Nrf2 pathway activation, we found that some tumors show high Nrf2 activity in the absence of known mutations in the pathway. An analysis of splice variants in oncogenes revealed that such tumors express abnormal transcript variants from the NFE2L2 gene (encoding Nrf2) that lack exon 2, or exons 2 and 3, and encode Nrf2 protein isoforms missing the KEAP1 interaction domain. The Nrf2 alterations result in the loss of interaction with KEAP1, Nrf2 stabilization, induction of a Nrf2 transcriptional response, and Nrf2 pathway dependence. In all analyzed cases, transcript variants were the result of heterozygous genomic microdeletions. Thus, we identify an alternative mechanism for Nrf2 pathway activation in human tumors and elucidate its functional consequences.
TheGOstats package has extensive facilities for testing the association of Gene Ontology (GO) The Gene Ontology Consortium (2000) terms to genes in a gene list. You can test for both over and under representation of GO terms using either the standard Hypergeometric test or a conditional Hypergeometric test that uses the relationships among the GO terms for conditioning (similar to that presented in Alexa et al. (2006)). In this vignette we describe the preprocessing required to construct inputs for the main testing function, hyperGTest, the algorithms used, and the structure of the return value. We use a microarray data set (Chiaretti et al., 2004) from a clinical trial in acute lymphoblastic leukemia (ALL) to work an example analysis. In the ALL data, we focus on the patients with B-cell derived ALL, and in particular on comparing the group with ALL1/AF4 to those with no observed cytogenetic abnormalities. To get started, load the packages needed for this analysis:
Background RNA-editing is a tightly regulated, and essential cellular process for a properly functioning brain. Dysfunction of A-to-I RNA editing can have catastrophic effects, particularly in the central nervous system. Thus, understanding how the process of RNA-editing is regulated has important implications for human health. However, at present, very little is known about the regulation of editing across tissues, and individuals. Results Here we present an analysis of RNA-editing patterns from 9 different tissues harvested from a single mouse. For comparison, we also analyzed data for 5 of these tissues harvested from 15 additional animals. We find that tissue specificity of editing largely reflects differential expression of substrate transcripts across tissues. We identified a surprising enrichment of editing in intronic regions of brain transcripts, that could account for previously reported higher levels of editing in brain. There exists a small but remarkable amount of editing which is tissue-specific, despite comparable expression levels of the edit site across multiple tissues. Expression levels of editing enzymes and their isoforms can explain some, but not all of this variation. Conclusions Together, these data suggest a complex regulation of the RNA-editing process beyond transcript expression levels.