Abstract Background: Rhabdomyosarcoma (RMS) is a highly malignant pediatric soft-tissue sarcoma where molecular subtyping, particularly PAX3/7::FOXO1 fusion status, drives prognosis and treatment. However, histology-based diagnostic approaches remain limited by subjectivity and the scarcity of comprehensive molecular annotations. To overcome these challenges, we improved our previously reported convolutional neural network learning models that predict PAX3/7::FOXO1 fusion status from whole-slide images (WSIs) while additionally trained models to infer gene expression profiles from histology, thereby linking morphology to transcriptomic signatures. Methods: A total of 826 independent WSIs from three sources [Children’s Oncology Group (COG) biobanking protocols = 322, Kids First (KIDS) = 252, Childhood Cancer Data Initiative/Molecular Characterization Initiative (CCDI/MCI) = 252] were used to train and evaluate an Attention-Based Multiple Instance Learning (ABMIL) model using UNI2-h foundation features for fusion classification. For gene expression prediction, 135 RMS WSIs paired with bulk RNA-seq data were used to fine-tune a SEQUOIA transformer model, which was trained on TCGA UCEC/COAD datasets. Model performance was evaluated using the Matthews Correlation Coefficient (MCC), AUC, and Pearson's r correlation, with biological validation through pathway enrichment analysis. Results: The fusion detection model achieved robust and generalizable performance across independent test cohorts (MCC ≥ 0.80, AUC ≥ 0.94), with multi-institutional training improving external generalization (MCC up to 0.84). The gene expression model reliably predicted bulk transcriptomic profiles from WSIs (mean r > 0.6, p < 0.05), identifying biologically meaningful pathways including cell cycle and muscle development. Together, these models demonstrate the feasibility of integrating morphological imaging data to gain molecular insights that would not be possible with histology alone. Conclusions: This work presents the first large-scale validated deep learning framework for simultaneous molecular subtyping and transcriptomic inference in RMS. By combining digital pathology with molecular prediction, our approach offers a scalable, tissue-sparing, and generalizable tool for advancing precision oncology in RMS. Citation Format: Dorsa Ziaei, Hyun Jung, Philip J. Lupo, Pagna Sok, Jack F. Shern, Corinne M. Linardic, Syed Abbas Bukhari, Hsein-Chao Chou, Jun S. Wei, Curtis Lisle, Uma Mudunuri, Javed Khan. PAX3/7::FOXO1 fusion detection and transcriptomic prediction from whole-slide images of rhabdomyosarcoma using attention-based deep learning frameworks: A multi-institutional study [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 2758.
BackgroundApproximately 4-8% of the world suffers from a rare disease. Rare diseases are often difficult to diagnose, and many do not have approved therapies. Genetic sequencing has the potential to shorten the current diagnostic process, increase mechanistic understanding, and facilitate research on therapeutic approaches but is limited by the difficulty of novel variant pathogenicity interpretation and the communication of known causative variants. It is unknown how many published rare disease variants are currently accessible in the public domain.ResultsThis study investigated the translation of knowledge of variants reported in published manuscripts to publicly accessible variant databases. Variants, symptoms, biochemical assay results, and protein function from literature on the SLC6A8 gene associated with X-linked Creatine Transporter Deficiency (CTD) were curated and reported as a highly annotated dataset of variants with clinical context and functional details. Variants were harmonized, their availability in existing variant databases was analyzed and pathogenicity assignments were compared with impact algorithm predictions. 24% of the pathogenic variants found in PubMed articles were not captured in any database used in this analysis while only 65% of the published variants received an accurate pathogenicity prediction from at least one impact prediction algorithm.ConclusionsDespite being published in the literature, pathogenicity data on patient variants may remain inaccessible for genetic diagnosis, therapeutic target identification, mechanistic understanding, or hypothesis generation. Clinical and functional details presented in the literature are important to make pathogenicity assessments. Impact predictions remain imperfect but are improving, especially for single nucleotide exonic variants, however such predictions are less accurate or unavailable for intronic and multi-nucleotide variants. Developing text mining workflows that use natural language processing for identifying diseases, genes and variants, along with impact prediction algorithms and integrating with details on clinical phenotypes and functional assessments might be a promising approach to scale literature mining of variants and assigning correct pathogenicity. The curated variants list created by this effort includes context details to improve any such efforts on variant curation for rare diseases.
ABSTRACTMicroRNAs (miRNAs) function as master regulators of gene expression in many physiological and pathological conditions including cancer. Sequence variants or isoforms (isomiRs) can account for between 40 to 60% of total miRNA counts, yet despite this overwhelming abundance, their function continues to be debated. Recent studies demonstrate that certain isomiRs can regulate unique sets of target mRNAs by altering their seed sequence or stabilizing 3’ pairing, while others are decay intermediates indicating an active miRNA turnover. Given their short sequence length and high heterogeneity, mapping isomiRs can be challenging; without adequate depth and data aggregation, low frequency events are often disregarded. To address these challenges, we present the Tumor IsomiR Encyclopedia (TIE): a dynamic database of isomiRs from over 10,000 adult and pediatric tumor samples in The Cancer Genome Atlas (TCGA) and The Therapeutically Applicable Research to Generate Effective Treatments (TARGET) projects. A key novelty of TIE is its ability to annotate heterogeneous isomiR sequences and aggregate the variants obtained across all samples and datasets. The database provides annotation of templated and non-templated nucleotides as well as other advanced analysis. All data can be browsed online or downloaded as simple spreadsheets. Here we show analysis of isomiRs of miR-21 and miR-30a to demonstrate the utility of TIE. TIE search engine and data are hosted at https://isomir.ccr.cancer.gov/.
SUMMARY:The Annotation, Visualization and Impact Analysis (AVIA) is a web application combining multiple features to annotate and visualize genomic variant data. Users can investigate functional significance of their genetic alterations across samples, genes and pathways. Version 3.0 of AVIA offers filtering options through interactive charts and by linking disease relevant data sources. Newly incorporated services include gene, variant and sample level reporting, literature and functional correlations among impacted genes, comparative analysis across samples and against data sources such as TCGA and ClinVar, and cohort building. Sample and data management is now feasible through the application, which allows greater flexibility with sharing, reannotating and organizing data. Most importantly, AVIA's utility stems from its convenience for allowing users to upload and explore results without any a priori knowledge or the need to install, update and maintain software or databases. Together, these enhancements strengthen AVIA as a comprehensive, user-driven variant analysis portal. AVAILABILITYAND IMPLEMENTATION:AVIA is accessible online at https://avia-abcc.ncifcrf.gov.
We perform an immunogenomics analysis utilizing whole-transcriptome sequencing of 657 pediatric extracranial solid cancer samples representing 14 diagnoses, and additionally utilize transcriptomes of 131 pediatric cancer cell lines and 147 normal tissue samples for comparison. We describe patterns of infiltrating immune cells, T cell receptor (TCR) clonal expansion, and translationally relevant immune checkpoints. We find that tumor-infiltrating lymphocytes and TCR counts vary widely across cancer types and within each diagnosis, and notably are significantly predictive of survival in osteosarcoma patients. We identify potential cancer-specific immunotherapeutic targets for adoptive cell therapies including cell-surface proteins, tumor germline antigens, and lineage-specific transcription factors. Using an orthogonal immunopeptidomics approach, we find several potential immunotherapeutic targets in osteosarcoma and Ewing sarcoma and validated PRAME as a bona fide multi-pediatric cancer target. Importantly, this work provides a critical framework for immune targeting of extracranial solid tumors using parallel immuno-transcriptomic and -peptidomic approaches.
Although individually uncommon, collectively, rare diseases affect 6 8% of the world’s population. Diagnosing rare diseases through phenotypes is a difficult task because symptoms overlap based solely on phenotype there may be multiple causes for a single disease and a single cause may be associated with multiple diseases. Genomic investigations are leading to precise molecular level characterization allowing for a systematic discovery of therapies that either target a specific disease or, more broadly, multiple related diseases. Currently, text mining techniques have lower accuracy and more coverage gaps than manual curation, but the potential for higher productivity. Recently, there has been some progress in developing text mining applications to tackle this enormous problem. However, there is still a need to assess these text mining applications and integrate the findings with data. Towards these ends, we manually created a rare disease variant data set which can be used to test and refine a text mining algorithm. We developed a manual curation workflow which incorporated searching for genetic variants on PubMed, listing symptoms and phenotypes associated with each variant, converting cDNA or protein to standardized genetic notation, obtaining annotations from public data sources, and presenting the details in an online interface for future release. Our project aims include summation of pathogenic variant frequency in populations to estimate birth prevalence for each rare disease. We created test data sets of known accuracy, coverage, and genotype phenotype associations that can be used to validate text mining approaches. Enhanced text mining will significantly decrease the amount of time necessary to gather data to molecularly characterize a rare disease and render it possible to mine rare disease phenotype genotype associations for the more than 7,000 rare disease genes in a timely fashion.
The Frederick National Laboratory for Cancer Research hosts a Data Coordinating Center and Toolset (the Metadata Designer plus Validator) hereafter referred to as the DCC that embodies a scalable, next-generation biological and cancer research data repository that is flexible, intuitive, and adaptive. The DCC provides integrated management of datasets across all deposited projects making its data more accessible and easily reusable by the cancer research community. The DCC stores and manages access to data, enabling researchers or data depositors to grant controlled access only to specific collaborators while maintaining a user-specified embargo on deposited datasets. The DCC enables a data access and sharing capability aimed to facilitate the development of new biological insights. The DCC implements the ISA (Investigation-Study-Assay) paradigm. The ISA framework provides a rich description of experimental metadata that is agnostic and irrespective of sample characteristic, technology or measurement type. It provides clear and simple sample-to-data relationships that enables resulting data and discoveries to be reproducible and reusable. These data are in the standard Investigation-Study-Assay tab-delimited format (ISA-TAB) format, which describes a scientific investigation, its study or studies, and each study's assay(s). The DCC portal is a public repository of experiment-related information describing cancer and biomedical research investigations. The portal can be used to browse, search, and access data from uploaded datasets. Through its stand-alone Toolset (Metadata Designer plus Validator), the DCC offers data depositors and researchers a simple interface to create and validate ISA-Tab compatible metadata associated with data generated in their research. Deriving new insights from aggregates of datasets and the need for reproducible research has never been more apparent. However, a fundamental requirement for these is a comprehensive metadata annotation and documentation of the research processes - an art that is elusive to researchers. Beyond the resultant publication is the data and metadata. These two basic ingredients are building blocks for a successful reproducible research project. In addition, appropriately curated data and metadata are invaluable resources for meta-analyses of seemingly disparate research studies. Further emphasizing the importance of these basic elements, are the principles of Findability, Accessibility, Interoperability, and Reusability (FAIR) - to guide data producers and publishers around obstacles and help maximize added-value gained from published studies. The DCC provides a much-needed resource to the research community for simplifying the sharing of data sets that meet these principles. The DCC portal is located at https://cssi-dcc.nci.nih.gov/cssiportal/ Citation Format: Paul Aiyetan, Paul Donovan, David Mott, Matthew Starr, Rajani Kuchipudi, Mahesh Yelisetti, Debra Hope, Corinne Zeitler, Uma Mudunuri, Andrew Quong. Enabling data access, sharing, collaborative and reproducible research: The Frederick National Laboratory for Cancer Research (FNLCR) data coordinating center [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 2489.
Journal Article A FAIR Principle Data Model for Focused Ion Beam Scanning Electron Microscopy (FIB-SEM) and the Frederick National Laboratory Data Coordinating Center Get access Paul Aiyetan, Paul Aiyetan Frederick National Laboratory for Cancer Research, Frederick, MD, USAGeorge Mason University School of Systems Biology, Manassas, VA, USA Corresponding author: paul.aiyetan@nih.gov Search for other works by this author on: Oxford Academic Google Scholar Kedar Narayan, Kedar Narayan Frederick National Laboratory for Cancer Research, Frederick, MD, USACenter for Molecular Microscopy, Center for Cancer Research, National Cancer Institute, National Institutes of Health, Frederick, MD, USA Search for other works by this author on: Oxford Academic Google Scholar David Mott, David Mott Frederick National Laboratory for Cancer Research, Frederick, MD, USA Search for other works by this author on: Oxford Academic Google Scholar Rajani Kuchipudi, Rajani Kuchipudi Frederick National Laboratory for Cancer Research, Frederick, MD, USA Search for other works by this author on: Oxford Academic Google Scholar Corinne Zeitler, Corinne Zeitler Frederick National Laboratory for Cancer Research, Frederick, MD, USA Search for other works by this author on: Oxford Academic Google Scholar Debra Hope, Debra Hope Frederick National Laboratory for Cancer Research, Frederick, MD, USA Search for other works by this author on: Oxford Academic Google Scholar Uma Mudunuri, Uma Mudunuri Frederick National Laboratory for Cancer Research, Frederick, MD, USA Search for other works by this author on: Oxford Academic Google Scholar Andrew Quong Andrew Quong Frederick National Laboratory for Cancer Research, Frederick, MD, USA Search for other works by this author on: Oxford Academic Google Scholar Microscopy and Microanalysis, Volume 25, Issue S2, 1 August 2019, Pages 1368–1369, https://doi.org/10.1017/S1431927619007578 Published: 01 August 2019
BACKGROUND:Network medicine aims to map molecular perturbations of any given diseases onto complex networks with functional interdependencies that underlie a pathological phenotype. Furthermore, investigating the time dimension of disease progression from a network perspective is key to gaining key insights to the disease process and to identify diagnostic or therapeutic targets. Existing platforms are ineffective to modularize the large complex systems into subgroups and consolidate heterogeneous data to web-based interactive animation.RESULTS:We have developed PanoromiX platform, a data-agnostic dynamic interactive visualization web application, enables the visualization of outputs from genome based molecular assays onto modular and interactive networks that are correlated with any pathophenotypic data (MRI, Xray, behavioral, etc.) over a time course all in one pane. As a result, PanoromiX reveals the complex organizing principles that orchestrate a disease-pathology from a gene regulatory network (nodes, edges, hubs, etc.) perspective instead of snap shots of assays. Without extensive programming experience, users can design, share, and interpret their dynamic networks through the PanoromiX platform with rich built-in functionalities.CONCLUSIONS:This emergent tool of network medicine is the first to visualize the interconnectedness of tailored genome assays to pathological networks and phenotypes for cells or organisms in a data-agnostic manner. As an advanced network medicine tool, PanoromiX allows monitoring of panel of biomarker perturbations over the progression of diseases, disease classification based on changing network modules that corresponds to specific patho-phenotype as opposed to clinical symptoms, systematic exploration of complex molecular interactions and distinct disease states via regulatory network changes, and the discovery of novel diagnostic and therapeutic targets.
Advanced high throughput technologies such as Next Generation Sequencing have created opportunities for scientists to explore the human microbiome more intensively and effectively. This enables us to understand the impact of imbalances in normal microbial flora through determining changes in the microbial composition and investigating how these correspond with changes occurring at the molecular level and the physiological level in the human body. Here, we developed the visualization explorer component for our in‐house automated microbiome analysis pipeline. The explorer takes enriched microbial composition per sample as input and provides a unique way to classify and visualize the microbiome data, organized under tabs. Under the relative abundance tab is visualization of the microbial composition across different samples with taxonomic annotation using different charts such as bar and radial charts. The PCA/Diversity index tab allows one to perform and visualize the principal component analysis and microbial diversity index comparison between treatment and control categories. Under the functional enrichment tab, the metabolite enrichment for the present microbial composition per sample is conducted and provides the predicted significant pathways between two categories in a network view. This visualization explorer was built using R open source language and Shiny server API. PICRUST performs the metabolite prediction and pathway enrichment. For the network visualization, explorer uses Graphviz open source software. The other features of this explorer include several visualization options such as heatmaps, histograms and line charts to compare differentially abundant bacterial species between disease and control samples. Support or Funding Information DHPRDTEGDF 374000GJ
How to visualize complex dynamic networks is a tremendous challenge in many scientific areas. For example, biologists want to show the changes of various molecules across time. The engineers would like to track the numerous measurements of industrial systems. Some desktop applications such as Cytoscape, Matlab, and Gephi, may be suitable for a small network to produce static images. However, the static images are almost impossible to read when the network complexity increases. To create a useful tool for visualizing complex dynamic networks, we developed Modular Explorer (MOE) to ascertain the three most desirable features as follows: 1. Scientifically informative: Has an interactive environment, modular layout, and animation. 2. Flexible: Has enough design options with minimal programming experience required. 3. Sharable: Safe sharing via email, compatible with most computers/ smart phones. To illustrate the efficacy of MOE, we used a prion protein (Prp) replication and accumulation network. The authors integrated the global gene expression in the brains of eight distinct mouse strain-prion strain combinations throughout the progression of the prion disease, and identified differentially expressed genes (DEGs) and pathways for different stages of the disease. A putative network was constructed based on the pooled DEGs primarily found in BL6 mice infected with infected prions across a time course of 6, 10, 14, 18, 20, and 22 weeks after inoculation. The DEGs shared by various prion-mouse combinations were carefully distinguished. The original network was developed using Cytoscape and took tremendous time and effort to organize the layout and carefully modify the node size and font styles one by one. Using MOE, it is rather straightforward to construct the full interactive network in a few seconds by uploading nodes and links information, and yet have the dynamic slideshow and many other functions, such as search capability, gene description, and focused interactions. In this paper, we introduce an easy-to-use web application MOE with rich functions for dynamic network visualization. It was specifically designed to foster scientifically informative, flexible, and sharable visualization that is desired by many scientists. With this tool, one can save a lot of manual labor and have informative, interactive networks that can be displayed and shared freely. We are not aware of any similar products in existence. Support or Funding Information DISCLAIMERS: Research was conducted in compliance with the Animal Welfare Act and all other Federal Requirements. The views expressed are those of the authors and do not constitute endorsement by the U.S. Army. A six-time-point dynamic network representing prion replication and accumulation process reproduced from Fig 4 in Hwang et al. 2009. In all, 16 pathways represented in modules are involved in the network; the node color denotes the relative fold change of the gene expression (red-upregulation, yellow-no change, green-downregulation). There are six prion-mice combinations. A six-time-point dynamic network representing prion replication and accumulation process reproduced from Fig 4 in Hwang et al. 2009. In all, 16 pathways represented in modules are involved in the network; the node color denotes the relative fold change of the gene expression (red-upregulation, yellow-no change, green-downregulation). There are six prion-mice combinations.
DNA damage in somatic cells originates from both environmental and endogenous sources, giving rise to mutations through multiple mechanisms. When these mutations affect the function of critical genes, cancer may ensue. Although identifying genomic subsets of mutated genes may inform therapeutic options, a systematic survey of tumor mutational spectra is required to improve our understanding of the underlying mechanisms of mutagenesis involved in cancer etiology. Recent studies have presented genome-wide sets of somatic mutations as a 96-element vector, a procedure that only captures the immediate neighbors of the mutated nucleotide. Herein, we present a 32 × 12 mutation matrix that captures the nucleotide pattern two nucleotides upstream and downstream of the mutation. A somatic autosomal mutation matrix (SAMM) was constructed from tumor-specific mutations derived from each of 909 individual cancer genomes harboring a total of 10,681,843 single-base substitutions. In addition, mechanistic template mutation matrices (MTMMs) representing oxidative DNA damage, ultraviolet-induced DNA damage, 5m CpG deamination, and APOBEC-mediated cytosine mutation, are presented. MTMMs were mapped to the individual tumor SAMMs to determine the maximum contribution of each mutational mechanism to the overall mutation pattern. A Manhattan distance across all SAMM elements between any two tumor genomes was used to determine their relative distance. Employing this metric, 89.5 % of all tumor genomes were found to have a nearest neighbor from the same tissue of origin. When a distance-dependent 6-nearest neighbor classifier was used, 86.9 % of all SAMMs were assigned to the correct tissue of origin. Thus, although tumors from different tissues may have similar mutation patterns, their SAMMs often display signatures that are characteristic of specific tissues.
UNLABELLED:As sequencing becomes cheaper and more widely available, there is a greater need to quickly and effectively analyze large-scale genomic data. While the functionality of AVIA v1.0, whose implementation was based on ANNOVAR, was comparable with other annotation web servers, AVIA v2.0 represents an enhanced web-based server that extends genomic annotations to cell-specific transcripts and protein-level functional annotations. With AVIA's improved interface, users can better visualize their data, perform comprehensive searches and categorize both coding and non-coding variants.AVAILABILITY AND IMPLEMENTATION:AVIA is freely available through the web at http://avia.abcc.ncifcrf.gov.CONTACT:Hue.Vuong@fnlcr.nih.govSUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
As the discipline of biomedical science continues to apply new technologies capable of producing unprecedented volumes of noisy and complex biological data, it has become evident that available methods for deriving meaningful information from such data are simply not keeping pace. In order to achieve useful results, researchers require methods that consolidate, store and query combinations of structured and unstructured data sets efficiently and effectively. As we move towards personalized medicine, the need to combine unstructured data, such as medical literature, with large amounts of highly structured and high-throughput data such as human variation or expression data from very large cohorts, is especially urgent. For our study, we investigated a likely biomedical query using the Hadoop framework. We ran queries using native MapReduce tools we developed as well as other open source and proprietary tools. Our results suggest that the available technologies within the Big Data domain can reduce the time and effort needed to utilize and apply distributed queries over large datasets in practical clinical applications in the life sciences domain. The methodologies and technologies discussed in this paper set the stage for a more detailed evaluation that investigates how various data structures and data models are best mapped to the proper computational framework.
SysBioCube is an integrated data warehouse and analysis platform for experimental data relating to diseases of military relevance developed for the US Army Medical Research and Materiel Command Systems Biology Enterprise (SBE). It brings together, under a single database environment, pathophysio-, psychological, molecular and biochemical data from mouse models of post-traumatic stress disorder and (pre-) clinical data from human PTSD patients.. SysBioCube will organize, centralize and normalize this data and provide an access portal for subsequent analysis to the SBE. It provides new or expanded browsing, querying and visualization to provide better understanding of the systems biology of PTSD, all brought about through the integrated environment. We employ Oracle database technology to store the data using an integrated hierarchical database schema design. The web interface provides researchers with systematic information and option to interrogate the profiles of pan-omics component across different data types, experimental designs and other covariates.
Single base substitutions constitute the most frequent type of human gene mutation and are a leading cause of cancer and inherited disease. These alterations occur non-randomly in DNA, being strongly influenced by the local nucleotide sequence context. However, the molecular mechanisms underlying such sequence context-dependent mutagenesis are not fully understood. Using bioinformatics, computational and molecular modeling analyses, we have determined the frequencies of mutation at G • C bp in the context of all 64 5'-NGNN-3' motifs that contain the mutation at the second position. Twenty-four datasets were employed, comprising >530,000 somatic single base substitutions from 21 cancer genomes, >77,000 germline single-base substitutions causing or associated with human inherited disease and 16.7 million benign germline single-nucleotide variants. In several cancer types, the number of mutated motifs correlated both with the free energies of base stacking and the energies required for abstracting an electron from the target guanines (ionization potentials). Similar correlations were also evident for the pathological missense and nonsense germline mutations, but only when the target guanines were located on the non-transcribed DNA strand. Likewise, pathogenic splicing mutations predominantly affected positions in which a purine was located on the non-transcribed DNA strand. Novel candidate driver mutations and tissue-specific mutational patterns were also identified in the cancer datasets. We conclude that electron transfer reactions within the DNA molecule contribute to sequence context-dependent mutagenesis, involving both somatic driver and passenger mutations in cancer, as well as germline alterations causing or associated with inherited disease.
The non-B DB, available at http://nonb.abcc.ncifcrf.gov, catalogs predicted non-B DNA-forming sequence motifs, including Z-DNA, G-quadruplex, A-phased repeats, inverted repeats, mirror repeats, direct repeats and their corresponding subsets: cruciforms, triplexes and slipped structures, in several genomes. Version 2.0 of the database revises and re-implements the motif discovery algorithms to better align with accepted definitions and thresholds for motifs, expands the non-B DNA-forming motifs coverage by including short tandem repeats and adds key visualization tools to compare motif locations relative to other genomic annotations. Non-B DB v2.0 extends the ability for comparative genomics by including re-annotation of the five organisms reported in non-B DB v1.0, human, chimpanzee, dog, macaque and mouse, and adds seven additional organisms: orangutan, rat, cow, pig, horse, platypus and Arabidopsis thaliana. Additionally, the non-B DB v2.0 provides an overall improved graphical user interface and faster query performance.
The accumulation of mutations is a contributing factor in the initiation of premalignant mammary lesions and their progression to malignancy and metastasis. We have used a mouse model in which the carcinogen is the mouse mammary tumor virus (MMTV) which induces clonal premalignant mammary lesions and malignant mammary tumors by insertional mutagenesis. Identification of the genes and signaling pathways affected in MMTV-induced mouse mammary lesions provides a rationale for determining whether genetic alteration of the human orthologues of these genes/pathways may contribute to human breast carcinogenesis. A high-throughput platform for inverse PCR to identify MMTV-host junction fragments and their nucleotide sequences in a large panel of MMTV-induced lesions was developed. Validation of the genes affected by MMTV-insertion was carried out by microarray analysis. Common integration site (CIS) means that the gene was altered by an MMTV proviral insertion in at least two independent lesions arising in different hosts. Three of the new genes identified as CIS for MMTV were assayed for their capability to confer on HC11 mouse mammary epithelial cells the ability for invasion, anchorage independent growth and tumor development in nude mice. Analysis of MMTV induced mammary premalignant hyperplastic outgrowth (HOG) lines and mammary tumors led to the identification of CIS restricted to 35 loci. Within these loci members of the Wnt, Fgf and Rspo gene families plus two linked genes (Npm3 and Ddn) were frequently activated in tumors induced by MMTV. A second group of 15 CIS occur at a low frequency (2-5 observations) in mammary HOGs or tumors. In this latter group the expression of either Phf19 or Sdc2 was shown to increase HC11 cells invasion capability. Foxl1 expression conferred on HC11 cells the capability for anchorage-independent colony formation in soft agar and tumor development in nude mice. The published transcriptome and nucleotide sequence analysis of gene expression in primary human breast tumors was interrogated. Twenty of the human orthologues of MMTV CIS associated genes are deregulated and/or mutated in human breast tumors.
Although the capability of DNA to form a variety of non-canonical (non-B) structures has long been recognized, the overall significance of these alternate conformations in biology has only recently become accepted en masse. In order to provide access to genome-wide locations of these classes of predicted structures, we have developed non-B DB, a database integrating annotations and analysis of non-B DNA-forming sequence motifs. The database provides the most complete list of alternative DNA structure predictions available, including Z-DNA motifs, quadruplex-forming motifs, inverted repeats, mirror repeats and direct repeats and their associated subsets of cruciforms, triplex and slipped structures, respectively. The database also contains motifs predicted to form static DNA bends, short tandem repeats and homo(purine•pyrimidine) tracts that have been associated with disease. The database has been built using the latest releases of the human, chimp, dog, macaque and mouse genomes, so that the results can be compared directly with other data sources. In order to make the data interpretable in a genomic context, features such as genes, single-nucleotide polymorphisms and repetitive elements (SINE, LINE, etc.) have also been incorporated. The database is accessed through query pages that produce results with links to the UCSC browser and a GBrowse-based genomic viewer. It is freely accessible at http://nonb.abcc.ncifcrf.gov.