Sharing data across research groups is an essential driver of biomedical research. In particular, biomedical databases with interactive query-answering systems allow users to retrieve information from the database using restricted types of queries. For example, medical data repositories allow researchers designing clinical studies to query how many patients in the database satisfy certain criteria, a workflow known as cohort discovery. In addition, genomic “beacon” services allow users to query whether or not a given genetic variant is observed in the database, a workflow we refer to as variant lookup. While these systems aim to facilitate the sharing of aggregate biomedical insights without divulging sensitive individual-level data, they can still leak private information about the individuals through the query answers. To address these privacy concerns, existing studies have proposed to perturb query results with a small amount of noise in order to reduce sensitivity to underlying individuals [1, 2]. However, these existing efforts either lack rigorous guarantees of privacy or introduce an excessive amount of noise into the system, limiting their effectiveness in practice.
Protein-DNA, -RNA and -peptide interactions drive nearly all cellular processes. Due to their high importance, high-throughput technologies using sequence libraries that cover all k-mers (i.e. words of length k) have been developed to measure them in a universal and unbiased manner [1]. These techniques all face a similar challenge: the space on the experimental device is limited, restricting the total sequence space that can be probed in a single experiment. While de Bruijn sequences cover all k-mers in the most compact manner, they remain |Σ|k characters long (where Σ is the alphabet, e.g. {A,C,G,T}). Here, we introduce a novel idea and algorithm for sequence design to cover all possible k-mers with a significantly smaller experimental sequence library by using joker characters, which represent all characters in the alphabet. Experimentally, such joker characters can be easily incorporated during oligonucleotide or peptide synthesis by using degenerate mixtures of nucleotides or amino acids, at no extra cost. However, joker characters introduce degeneracy which could potentially lower the statistical robustness of the measurements (as a measurement of a single oligonucleotide is now assigned to multiple sequences instead of just one). To address this challenge, we limit the use of joker characters to either one or two joker characters per k-mer, enabling the coverage of (k+2)-mers at the same cost and space of k-mers — a savings of a factor of |Σ|2 in sequence length (16 and 400 for DNA and amino acid alphabets, respectively). We validate that the library remains capable of de novo identification of high-affinity k-mers by testing it on known DNA-protein binding data for hundreds of proteins. The implementation of our algorithm is freely available at jokercake.csail.mit.edu.
BACKGROUND:RNAs within extracellular vesicles (EVs) have potential as diagnostic biomarkers for patients with cancer and are identified in a variety of biofluids. Glioblastomas (GBMs) release EVs containing RNA into cerebrospinal fluid (CSF). Here we describe a multi-institutional study of RNA extracted from CSF-derived EVs of GBM patients to detect the presence of tumor-associated amplifications and mutations in epidermal growth factor receptor (EGFR).METHODS:CSF and matching tumor tissue were obtained from patients undergoing resection of GBMs. We determined wild-type (wt)EGFR DNA copy number amplification, as well as wtEGFR and EGFR variant (v)III RNA expression in tumor samples. We also characterized wtEGFR and EGFRvIII RNA expression in CSF-derived EVs.RESULTS:EGFRvIII-positive tumors had significantly greater wtEGFR DNA amplification (P = 0.02) and RNA expression (P = 0.03), and EGFRvIII-positive CSF-derived EVs had significantly more wtEGFR RNA expression (P = 0.004). EGFRvIII was detected in CSF-derived EVs for 14 of the 23 EGFRvIII tissue-positive GBM patients. Conversely, only one of the 48 EGFRvIII tissue-negative patients had the EGFRvIII mutation detected in their CSF-derived EVs. These results yield a sensitivity of 61% and a specificity of 98% for the utility of CSF-derived EVs to detect an EGFRvIII-positive GBM.CONCLUSION:Our results demonstrate CSF-derived EVs contain RNA signatures reflective of the underlying molecular genetic status of GBMs in terms of wtEGFR expression and EGFRvIII status. The high specificity of the CSF-derived EV diagnostic test gives us an accurate determination of positive EGFRvIII tumor status and is essentially a less invasive "liquid biopsy" that might direct mutation-specific therapies for GBMs.
Sequence libraries that cover all k-mers enable universal, unbiased measurements of binding to both oligonucleotides and peptides. While the number of k-mers grows exponentially in k, space on all experimental platforms is limited. Here, we shrink k-mer library sizes by using joker characters, which represent all characters in the alphabet simultaneously. We present the JokerCAKE (joker covering all k-mers) algorithm for generating a short sequence such that each k-mer appears at least p times with at most one joker character per k-mer. By running our algorithm on a range of parameters and alphabets, we show that JokerCAKE produces near-optimal sequences. Moreover, through comparison with data from hundreds of DNA-protein binding experiments and with new experimental results for both standard and JokerCAKE libraries, we establish that accurate binding scores can be inferred for high-affinity k-mers using JokerCAKE libraries. JokerCAKE libraries allow researchers to search a significantly larger sequence space using the same number of experimental measurements and at the same cost.
Sequence libraries that cover all k‐mers enable universal and unbiased measurements of binding to both oligonucleotides and peptides. For these measurements, it is desirable to search the largest possible library of k‐mers, as this increases the information content of any returned motif and the ability to predict function in vivo. However, while the number of k‐mers required grows exponentially in k, space on all experimental platforms is limited. The optimal solution makes use of a de Bruijn sequence, which covers all k‐mers in the most compact manner but still requires the explicit occurrence of all k‐mers within the sequence. Here, we introduce a novel advance to shrink k‐mer library sizes further by using joker characters, which represent all characters in the alphabet simultaneously. In this work, we consider the practical problem of generating a minimum‐length sequence that covers each k‐mer at least p times with the caveat that at most one or two joker characters occur in each k‐mer, thereby limiting the potential degeneracy that joker characters could introduce. We present the first algorithmic solution to this problem, JokerCAKE (Joker Covering All K‐mErs). By running our algorithm on a range of parameters and alphabets, we show that JokerCAKE produces sequences that are very close to the theoretical lower bound. Moreover, through comparison with data from hundreds of DNA‐protein binding experiments and direct comparison between experimental results obtained for both standard de Bruijn k‐mer sequences and joker de Bruijn k‐mer sequences, we establish that accurate binding scores can be inferred for high‐affinity k‐mers using JokerCAKE libraries
Analysis of extracellular vesicles (EVs) derived from plasma or cerebrospinal fluid (CSF) has emerged as a promising biomarker platform for therapeutic monitoring in glioblastoma patients. However, the contents of the various subpopulations of EVs in these clinical specimens remain poorly defined. Here we characterize the relative abundance of miRNA species in EVs derived from the serum and cerebrospinal fluid of glioblastoma patients. EVs were isolated from glioblastoma cell lines as well as the plasma and CSF of glioblastoma patients. The microvesicle subpopulation was isolated by pelleting at 10,000× g for 30 min after cellular debris was cleared by a 2000× g (20 min) spin. The exosome subpopulation was isolated by pelleting the microvesicle supernatant at 120,000× g (120 min). qRT-PCR was performed to examine the distribution of miR-21, miR-103, miR-24, and miR-125. Global miRNA profiling was performed in select glioblastoma CSF samples. In plasma and cell line derived EVs, the relative abundance of miRNAs in exosome and microvesicles were highly variable. In some specimens, the majority of the miRNA species were found in exosomes while in other, they were found in microvesicles. In contrast, CSF exosomes were enriched for miRNAs relative to CSF microvesicles. In CSF, there is an average of one molecule of miRNA per 150–25,000 EVs. Most EVs derived from clinical biofluids are devoid of miRNA content. The relative distribution of miRNA species in plasma exosomes or microvesicles is unpredictable. In contrast, CSF exosomes are the major EV compartment that harbor miRNAs.
Glioblastoma cells secrete extra-cellular vesicles (EVs) containing microRNAs (miRNAs). Analysis of these EV miRNAs in the bio-fluids of afflicted patients represents a potential platform for biomarker development. However, the analytic algorithm for quantitative assessment of EV miRNA remains under-developed. Here, we demonstrate that the reference transcripts commonly used for quantitative PCR (including GAPDH, 18S rRNA, and hsa-miR-103) were unreliable for assessing EV miRNA. In this context, we quantitated EV miRNA in absolute terms and normalized this value to the input EV number. Using this method, we examined the abundance of miR-21, a highly over-expressed miRNA in glioblastomas, in EVs. In a panel of glioblastoma cell lines, the cellular levels of miR-21 correlated with EV miR-21 levels (p<0.05), suggesting that glioblastoma cells actively secrete EVs containing miR-21. Consistent with this hypothesis, the CSF EV miR-21 levels of glioblastoma patients (n=13) were, on average, ten-fold higher than levels in EVs isolated from the CSF of non-oncologic patients (n=13, p<0.001). Notably, none of the glioblastoma CSF harbored EV miR-21 level below 0.25 copies per EV in this cohort. Using this cut-off value, we were able to prospectively distinguish CSF derived from glioblastoma and non-oncologic patients in an independent cohort of twenty-nine patients (Sensitivity=87%; Specificity=93%; AUC=0.91, p<0.01). Our results suggest that CSF EV miRNA analysis of miR-21 may serve as a platform for glioblastoma biomarker development.
The discovery that tumor-derived proteins and nucleic acids can be detected in nano-sized vesicles in the plasma and cerebrospinal fluid of patients afflicted with brain tumors has expanded opportunities for biomarker and therapeutic discovery. Through delivery of their contents to surrounding cells, exosomes, microvesicles, and other nano-sized extracellular vesicles secreted by tumors modulate their environment to promote tumor growth and survival. In this review, we discuss the biological processes mediated by these extracellular vesicles and their applications in terms of brain tumor diagnosis, monitoring, and therapy. We review the normal physiology of these extracellular vesicles, their pertinence to tumor biology, and directions for research in this field.
Recent studies suggest both normal and cancerous cells secrete vesicles into the extracellular space. These extracellular vesicles (EVs) contain materials that mirror the genetic and proteomic content of the secreting cell. The identification of cancer-specific material in EVs isolated from the biofluids (e.g., serum, cerebrospinal fluid, urine) of cancer patients suggests EVs as an attractive platform for biomarker development. It is important to recognize that the EVs derived from clinical samples are likely highly heterogeneous in make-up and arose from diverse sets of biologic processes. This article aims to review the biologic processes that give rise to various types of EVs, including exosomes, microvesicles, retrovirus like particles, and apoptotic bodies. Clinical pertinence of these EVs to neuro-oncology will also be discussed.
2052 Background: Leptomeningeal metastasis (LM) from solid tumors is typically a late manifestation of disease with a median survival of weeks to a few months. Treatment is palliative, with no widely accepted standard of care. Options include intrathecal (IT) or systemic chemotherapy, radiation therapy or ventriculoperitoneal shunting. Randomized trials comparing single agent IT methotrexate to liposomal cytarabine have shown similar efficacy and tolerability. There is limited data, however, on the use of combination IT chemotherapy in solid tumor LM. Methods: We conducted a retrospective cohort study of 19 subjects treated for LM from solid tumors at a single institution. In addition to therapies directed at active solid tumor sites, each subject received IT liposomal cytarabine plus IT methotrexate injections every two weeks. Survival data and treatment-related toxicities were determined by systematic chart review. Results: LM was diagnosed by CSF cytology in 12/19 (63%), while the remainder were diagnosed by clinical and MRI findings. The most common cancer types were breast 7(37%), glioblastoma 6(32%) and lung 3(16%). The majority 18(95%) had active systemic or parenchymal brain disease at the time of treatment, requiring systemic chemotherapy 18(95%) or radiation therapy 13(68%). The median number of IT treatments was 4(range 1-9). Treatment was interrupted due to toxicity in 3(17%), while 7(37%) experienced ≥ CTCAE grade III toxicities, most commonly meningitis 3(16%). Treatment was stopped in 7/19(37%) following complete cytologic response 6/11(55%) or radiographic clearance 1/7(14%). The median overall survival was 96 days(n=6; range 29-158), median time to neurologic progression was 46 days (n=9; range 6-101) and the most common cause of death was progression of systemic disease 4(67%). Conclusions: Combination IT chemotherapy was reasonably well-tolerated, even in a population also receiving chemotherapy for progressive systemic disease. IT-related adverse events occurred at rates similar to previously reported single agent trials. Prospective evaluation is necessary to determine whether there is a survival benefit compared to single agent IT chemotherapy.
The modest benefits of chemotherapy for Glioblastoma (GBM) patients’ makes it urgent to find potent drugs or combinations that can inhibitor tumor cell growth and proliferation. We have devised drug screening based on the NIH clinical collection compounds. This collection included 446 compounds that all have history of use in human clinical trials. The positive compounds can be rapidly entered into clinical trials for GBM. After found 22 drugs/compounds that can inhibit U87 cell proliferation and we tested 15 available positive drugs with A172, LN443, U118 cell line. After confirming in established serum-grown cell lines, we tested 8 FDA approved drugs in GBM primary cells with stem cell like features isolated from fresh brain tumor tissue samples and cultured in stem cell media. The various drug responses amongst the different patients’ primary cells suggested that we may be able to personalize drugs to optimize outcomes. In addition, we found that certain combinations of FDA approved drug had significantly enhanced activity against GBM cells in vitro when compared to the same agents used alone. Our work reveals that screening drugs and combinations for GBM patient may have clinical implications. Citation Format: {Authors}. {Abstract title} [abstract]. In: Proceedings of the 103rd Annual Meeting of the American Association for Cancer Research; 2012 Mar 31-Apr 4; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2012;72(8 Suppl):Abstract nr 3680. doi:1538-7445.AM2012-3680
With advances in genomic profiling and sequencing technology, we are beginning to understand the landscape of the genetic events that accumulated during the neoplastic process. The insights gleamed from these genomic profiling studies with regards to glioblastoma etiology has been particularly satisfying because it cemented the clinical pertinence of major concepts in cancer biology-concepts developed over the past three decades. This article will review how the glioblastoma genomic data set serves as an illustrative platform for the concepts put forward by Hanahan and Weinberg on the cancer phenotype. The picture emerging suggests that most glioblastomas evolve along a multitude of pathways rather than a single defined pathway. In this context, the article will further provide a discussion of the subtypes of glioblastoma as they relate to key principles of developmental neurobiology.
Cassava (Manihot esculenta Crantz) is a staple food for over 600 million people in the tropics and subtropics and is increasingly used as an industrial crop for starch production. Cassava has a high growth rate under optimal conditions but also performs well in drought-prone areas and on marginal soils. To increase the tools for understanding and manipulating drought tolerance in cassava, we generated expressed sequence tags (ESTs) from normalized cDNA libraries prepared from dehydration-stressed and control well-watered tissues. Analysis of a total of 18,166 ESTs resulted in the identification of 8,577 unique gene clusters (5,383 singletons and 3,194 clusters). Functional categories could be assigned to 63% of the unigenes, while another ∼11% were homologous to hypothetical genes with unclear functions. The remaining ∼26% were not significantly homologous to sequences in public databases suggesting that some may be novel and putatively specific to cassava. The dehydration-stressed library uncovered numerous ESTs with recognized roles in drought-responses, including those that encode late-embryogenesis-abundant proteins thought to confer osmoprotective functions during water stress, transcription factors, heat-shock proteins as well as proteins involved in signal transduction and oxidative stress. The unigene clusters were screened for short tandem repeats for further development as microsatellite markers. A total of 592 clusters contained 646 repeats, representing 3.3% of the ESTs queried. The ESTs presented here are the first dehydration stress transcriptome of cassava and can be utilized for the development of microarrays and gene-derived molecular markers to further dissect the molecular basis of drought tolerance in cassava.