Multiple myeloma is a treatable, but currently incurable, hematological malignancy of plasma cells characterized by diverse and complex tumor genetics for which precision medicine approaches to treatment are lacking. The Multiple Myeloma Research Foundation’s Relating Clinical Outcomes in Multiple Myeloma to Personal Assessment of Genetic Profile study ( NCT01454297 ) is a longitudinal, observational clinical study of newly diagnosed patients with multiple myeloma (n = 1,143) where tumor samples are characterized using whole-genome sequencing, whole-exome sequencing and RNA sequencing at diagnosis and progression, and clinical data are collected every 3 months. Analyses of the baseline cohort identified genes that are the target of recurrent gain-of-function and loss-of-function events. Consensus clustering identified 8 and 12 unique copy number and expression subtypes of myeloma, respectively, identifying high-risk genetic subtypes and elucidating many of the molecular underpinnings of these unique biological groups. Analysis of serial samples showed that 25.5% of patients transition to a high-risk expression subtype at progression. We observed robust expression of immunotherapy targets in this subtype, suggesting a potential therapeutic option. Longitudinal genomic and transcriptomic profiling of 1,143 patients with multiple myeloma by the Relating Clinical Outcomes in Multiple Myeloma to Personal Assessment of Genetic Profile study yields an improved copy number and gene expression subtype scheme, most notably a high-risk proliferative subtype associated with complete loss of RB1 or MAX.
Multiple myeloma is a treatable, but currently incurable, hematological malignancy of plasma cells characterized by diverse and complex tumor genetics for which precision medicine approaches to treatment are lacking. The MMRF CoMMpass study is a longitudinal, observational clinical study of newly diagnosed multiple myeloma patients where tumor samples are characterized using whole genome, exome, and RNA sequencing at diagnosis and progression, and clinical data is collected every three months. Analyses of the baseline cohort identified genes that are the target of recurrent gain- and loss-of-function events. Consensus clustering identified 8 and 12 unique copy number and expression subtypes of myeloma, respectively, identifying high-risk genetic subtypes and elucidating many of the molecular underpinnings of these unique biological groups. Analysis of serial samples showed 25.5% of patients transition to a high-risk expression subtype at progression. We observed robust expression of immunotherapy targets in this subtype, suggesting a potential therapeutic option.
With increasing utilization of comprehensive genomic data to guide clinical care, anticipated to become the standard of care in many clinical settings, the practice of diagnostic medicine is undergoing a notable shift. However, the move from single-gene or panel-based genetic testing to exome and genome sequencing has not been matched by the development of tools to enable diagnosticians to interpret increasingly complex or uncertain genomic findings. Here, we present gene.iobio, a real-time, intuitive and interactive web application for clinically-driven variant interrogation and prioritization. We show gene.iobio is a novel and effective approach that significantly improves upon and reimagines existing methods. In a radical departure from existing methods that present variants and genomic data in text and table formats, gene.iobio provides an interactive, intuitive and visually-driven analysis environment. We demonstrate that adoption of gene.iobio in clinical and research settings empowers clinical care providers to interact directly with patient genomic data both for establishing clinical diagnoses and informing patient care, using sophisticated genomic analyses that previously were only accessible via complex command line tools.
When ordering genetic testing or triaging candidate variants in exome and genome sequencing studies, it is critical to generate and test a comprehensive list of candidate genes that succinctly describe the complete and objective phenotypic features of disease. Significant efforts have been made to curate gene:disease associations both in academic research and commercial genetic testing laboratory settings. However, many of these valuable resources exist as islands and must be used independently, generating static, single-resource gene:disease association lists. Here we describe genepanel.iobio (https://genepanel.iobio.io) an easy to use, free and open-source web tool for generating disease- and phenotype-associated gene lists from multiple gene:disease association resources, including the NCBI Genetic Testing Registry (GTR), Phenolyzer, and the Human Phenotype Ontology (HPO). We demonstrate the utility of genepanel.iobio by applying it to complex, rare and undiagnosed disease cases that had reached a diagnostic conclusion. We find that genepanel.iobio is able to correctly prioritize the gene containing the diagnostic variant in roughly half of these challenging cases. Importantly, each component resource contributed diagnostic value, showing the benefits of this aggregate approach. We expect genepanel.iobio will improve the ease and diagnostic value of generating gene:disease association lists for genetic test ordering and whole genome or exome sequencing variant prioritization.
Additional file 1: Table S1. Summary of the diagnostic analysis for 16 previously diagnosed clinical cases, Table S2. Phenotype descriptions that were given to each analyst, whether each analyst was able to correctly identify the diagnostic gene using genepanel.iobio and a likely rationale for the analysts’.
DNA damage repair (DDR) pathways modulate cancer risk, progression, and therapeutic response. We systematically analyzed somatic alterations to provide a comprehensive view of DDR deficiency across 33 cancer types. Mutations with accompanying loss of heterozygosity were observed in over 1/3 of DDR genes, including TP53 and BRCA1/2. Other prevalent alterations included epigenetic silencing of the direct repair genes EXO5, MGMT, and ALKBH3 in similar to 20% of samples. Homologous recombination deficiency (HRD) was present at varying frequency in many cancer types, most notably ovarian cancer. However, in contrast to ovarian cancer, HRD was associated with worse outcomes in several other cancers. Protein structurebased analyses allowed us to predict functional consequences of rare, recurrent DDR mutations. A new machine-learning-based classifier developed from gene expression data allowed us to identify alterations that phenocopy deleterious TP53 mutations. These frequent DDR gene alterations in many human cancers have functional consequences that may determine cancer progression and guide therapy.
Early infantile epileptic encephalopathy (EIEE) is a devastating epilepsy syndrome with onset in the first months of life. Although mutations in more than 50 different genes are known to cause EIEE, current diagnostic yields with gene panel tests or whole-exome sequencing are below 60%. We applied whole-genome analysis (WGA) consisting of whole-genome sequencing and comprehensive variant discovery approaches to a cohort of 14 EIEE subjects for whom prior genetic tests had not yielded a diagnosis. We identified both de novo point and INDEL mutations and de novo structural rearrangements in known EIEE genes, as well as mutations in genes not previously associated with EIEE. The detection of a pathogenic or likely pathogenic mutation in all 14 subjects demonstrates the utility of WGA to reduce the time and costs of clinical diagnosis of EIEE. While exome sequencing may have detected 12 of the 14 causal mutations, 3 of the 12 patients received non-diagnostic exome panel tests prior to genome sequencing. Thus, given the continued decline of sequencing costs, our results support the use of WGA with comprehensive variant discovery as an efficient strategy for the clinical diagnosis of EIEE and other genetic conditions.
Home- and community-based services (HCBS) are critical to helping older adults age in place (AIP). Few studies have examined the association between older adults’ satisfaction with AIP and HCBS utilization, particularly in the context of the Andersen behavioral model of health services use. The purpose of the study is to examine how predisposing, enabling, and need characteristics combine with HCBS utilization to predict satisfaction with AIP. A survey was administered to 229 adults aged ≥65 years. The survey assessed individual predisposing (age, gender, race, marital status, education), enabling (income, activity satisfaction, informal caregiving availability, HCBS costliness), need (health status change), and contextual enabling (neighborhood resources, neighborhood accessibility, HCBS availability) characteristics. Use of nine HCBS was assessed (i.e., nursing, physical/occupational therapy, personal care, counseling, home-delivered meals, prescription delivery, home maintenance, transportation, and yard maintenance). Among those in need of HCBS (n=166), 29% were satisfied with AIP, and 71% were not satisfied; 20% did not use HCBS and 80% used at least 1 service. Satisfaction with AIP was significantly correlated (p<.05) with HCBS use (rho=-.16), activity satisfaction (rho=.26), HCBS costliness (rho=-.43), health status change (rho=.20), and neighborhood accessibility (rho=.19). In a multivariate logistic regression model, satisfaction with AIP increased with health status improvement [Exp(B)=3.45, p<.05] and neighborhood accessibility [Exp(B)=1.49, p<.05], but decreased with HCBS costliness [Exp(B)=0.67, p<.05]. HCBS use decreased satisfaction with AIP [Exp(B)=0.18, p<.05], possibly as a result of the overall decline in ability and activities signified by the need for and use of HCBS. Further research is needed.
INTRODUCTION:Computational analysis of genome or exome sequences may improve inherited disease diagnosis, but is costly and time-consuming.METHODS:We describe the use of iobio, a web-based tool suite for intuitive, real-time genome diagnostic analyses.RESULTS:We used iobio to identify the disease-causing variant in a patient with early infantile epileptic encephalopathy with prior nondiagnostic genetic testing.CONCLUSIONS:Iobio tools can be used by clinicians to rapidly identify disease-causing variants from genomic patient sequencing data.
OBJECTIVES/SPECIFIC AIMS: The objective of the study was 2-fold; to identify potentially deleterious alleles in a child with Treacher Collins syndrome, and; to demonstrate the value of the iobio analysis platform for intuitively and rapidly analyzing genomic data. METHODS/STUDY POPULATION: We used the iobio suite of web-based applications to analyze quality metrics for the sequencing data and called variants for the proband and his parents. We then visually interrogated variants in genes potentially associated with the syndrome in real-time, using the intuitive gene.iobio application. We sought high impact variants that demonstrated a predicted impact on the protein function, and were simultaneously at low allele frequency in the general human population. Variants were also compared against the ClinVar database of known mutations to identify variants that have already been associated with this, or related syndromes in the literature or clinical studies. Finally, the gene.iobio tool allows users to interrogate the primary sequencing data to ensure that no variants had been missed by the primary variant calling pipeline. This analysis pipeline was performed using intuitive web-based apps in real time, and consequently represents a system that is available to users that traditionally are excluded from these analyses. RESULTS/ANTICIPATED RESULTS: The iobio suite was used to rapidly assess data quality and interrogate genetic variants for a child with Treacher Collins syndrome. A compound heterozygote consisting of 2 missense alleles in the TCOF1 gene was identified as a compelling pathogenic allele, necessitating further functional investigation. The study helped validate the use of the intuitive iobio tools in such analyses, strengthening the case for greater involvement of medical professionals in data analysis. DISCUSSION/SIGNIFICANCE OF IMPACT: The performed analyses demonstrated that the whole genome sequencing data for the family being studied was of a very high quality, although 1 gene demonstrated a local region of almost zero coverage. This ensured that study conclusions can be presented with confidence. A variant associated with Treacher Collins syndrome 1 in ClinVar was uncovered in the TCOF1 gene, however, given it’s benign rating, this variant was not considered further. The most interesting candidate was a compound heterozygote, consisting of 2 missense mutations, also in the TCOF1 gene. These mutations occurred with allele frequencies of 22% and 8% in the general population, and additional molecular and functional studies are currently being pursued.
Background High-throughput sequencing enables unbiased profiling of microbial communities, universal pathogen detection, and host response to infectious diseases. However, computation times and algorithmic inaccuracies have hindered adoption. Results We present Taxonomer, an ultrafast, web-tool for comprehensive metagenomics data analysis and interactive results visualization. Taxonomer is unique in providing integrated nucleotide and protein-based classification and simultaneous host messenger RNA (mRNA) transcript profiling. Using real-world case-studies, we show that Taxonomer detects previously unrecognized infections and reveals antiviral host mRNA expression profiles. To facilitate data-sharing across geographic distances in outbreak settings, Taxonomer is publicly available through a web-based user interface. Conclusions Taxonomer enables rapid, accurate, and interactive analyses of metagenomics data on personal computers and mobile devices.
Fluorescent in situ hybridization (FISH) is commonly used in the multiple myeloma field to subtype and risk-stratify patients. There are many benefits to FISH based assays, which are widely used around the world and represent true single cell assays. However, there are significant discrepancies in the specific assays, utilization of reflex testing strategies, and enumeration requirements between clinical centers. By comparison next-generation sequencing tests can be designed to simultaneously detect the copy number abnormalities and translocations detected by clinical FISH along with gene mutations that cannot be detected by FISH. As part of the MMRF CoMMpass Study we have compared the results attained using clinical FISH assays compared to sequencing based FISH (Seq-FISH) results.
correspondence(v) the histogram of mapping quality values and fraction of properly mapped read pairs (to identify poor mapping results).Collecting such vital alignment statistics using current tools requires placing the BAM file on a Unix machine and then installing and running Unix programs such as SAMTools 1 or BamTools 2 on the entire BAM file.This process may take hours to complete, e.g., the 18-gigabyte BAM file in our tests took 8 hours to process (Supplementary Table 1).In contrast, our approach is to collect a random sample of the read alignments (Supplementary Fig. 1) to accurately estimate the same alignment statistics in seconds (Supplementary Fig. 2).Notably, sampling takes place where the BAM file is stored (i.e., on cloud storage or a user' s hard drive), and only the sampled data-a tiny fraction of the entire BAM file-are ever transmitted.The alignments are then streamed to data analysis web services that produce appropriate alignment statistics in seconds before transmitting these to bam.iobio for visualization (for implementation details, licensing and deployment considerations, see Supplementary Note; for system compatibility, see Supplementary Table 2).We can now analyze the same 18-gigabyte alignment file in <10 seconds.Realtime visualization allows the user to experience how the statistical distributions progressively converge and become stable as sampled alignment data are collected.The user can further explore the data interactively by selecting other chromosomes or chromosomal subregions, using the main read coverage panel for navigation.This web app puts forward an interactive and intuitive genomic data analysis paradigm that is not achievable with existing systems, enabling users to analyze both local and remotely stored data, without tool installation or transmitting large data sets, and immediately see informative results.We are developing other real-time analysis applications: for example, to analyze multiple alignment files simultaneously using our sampling approach, and for interactive, complete analysis of genomic data in smaller genomic windows, such as in the region of a gene (demos at http://iobio.io/).We are also creating software libraries for third-party developers to build similar interactive web apps.Although large, whole-genome computation will remain essential for many tasks, we expect that web-based, visually driven, real-time tools will offer a powerful new analysis modality for bioinformatics experts and bench scientists alike.
MOTIVATIONA common question arises at the beginning of every experiment where RNA-Seq is used to detect differential gene expression between two conditions: How many reads should we sequence?RESULTSScotty is an interactive web-based application that assists biologists to design an experiment with an appropriate sample size and read depth to satisfy the user-defined experimental objectives. This design can be based on data available from either pilot samples or publicly available datasets.AVAILABILITYScotty can be freely accessed on the web at http://euler.bc.edu/marthlab/scotty/scotty.php
SUMMARY:Biogem provides a software development environment for the Ruby programming language, which encourages community-based software development for bioinformatics while lowering the barrier to entry and encouraging best practices. Biogem, with its targeted modular and decentralized approach, software generator, tools and tight web integration, is an improved general model for scaling up collaborative open source software development in bioinformatics.AVAILABILITY:Biogem and modules are free and are OSS. Biogem runs on all systems that support recent versions of Ruby, including Linux, Mac OS X and Windows. Further information at http://www.biogems.info. A tutorial is available at http://www.biogems.info/howto.htmlCONTACT:bonnal@ingm.org.
MOTIVATIONHigh-throughput biological research requires simultaneous visualization as well as analysis of genomic data, e.g. read alignments, variant calls and genomic annotations. Traditionally, such integrative analysis required desktop applications operating on locally stored data. Many current terabyte-size datasets generated by large public consortia projects, however, are already only feasibly stored at specialist genome analysis centers. As even small laboratories can afford very large datasets, local storage and analysis are becoming increasingly limiting, and it is likely that most such datasets will soon be stored remotely, e.g. in the cloud. These developments will require web-based tools that enable users to access, analyze and view vast remotely stored data with a level of sophistication and interactivity that approximates desktop applications. As rapidly dropping cost enables researchers to collect data intended to answer questions in very specialized contexts, developers must also provide software libraries that empower users to implement customized data analyses and data views for their particular application. Such specialized, yet lightweight, applications would empower scientists to better answer specific biological questions than possible with general-purpose genome browsers currently available.RESULTSUsing recent advances in core web technologies (HTML5), we developed Scribl, a flexible genomic visualization library specifically targeting coordinate-based data such as genomic features, DNA sequence and genetic variants. Scribl simplifies the development of sophisticated web-based graphical tools that approach the dynamism and interactivity of desktop applications.AVAILABILITY AND IMPLEMENTATIONSoftware is freely available online at http://chmille4.github.com/Scribl/ and is implemented in JavaScript with all modern browsers supported.
Background Phyloinformatic analyses involve large amounts of data and metadata of complex structure. Collecting, processing, analyzing, visualizing and summarizing these data and metadata should be done in steps that can be automated and reproduced. This requires flexible, modular toolkits that can represent, manipulate and persist phylogenetic data and metadata as objects with programmable interfaces. Results This paper presents Bio::Phylo, a Perl5 toolkit for phyloinformatic analysis. It implements classes and methods that are compatible with the well-known BioPerl toolkit, but is independent from it (making it easy to install) and features a richer API and a data model that is better able to manage the complex relationships between different fundamental data and metadata objects in phylogenetics. It supports commonly used file formats for phylogenetic data including the novel NeXML standard, which allows rich annotations of phylogenetic data to be stored and shared. Bio::Phylo can interact with BioPerl, thereby giving access to the file formats that BioPerl supports. Many methods for data simulation, transformation and manipulation, the analysis of tree shape, and tree visualization are provided. Conclusions Bio::Phylo is composed of 59 richly documented Perl5 modules. It has been deployed successfully on a variety of computer architectures (including various Linux distributions, Mac OS X versions, Windows, Cygwin and UNIX-like systems). It is available as open source (GPL) software from http://search.cpan.org/dist/Bio-Phylo