Collusion between students in online exams is a major problem that undermines the integrity of the exam results. Although there exist methods that use exam data to identify pairs of students who have likely copied each other's answers, these methods are restricted to specific formats of multiple-choice exams. Here we present a statistical algorithm, Q-SID, that efficiently detects groups of students who likely have colluded, i.e., collusion groups, with error quantification. Q-SID uses graded numeric question scores only, so it works for many formats of multiple-choice and non-multiple-choice exams. Q-SID reports two false-positive rates (FPRs) for each collusion group: (1) empirical FPR, whose null data are from 36 strictly proctored exam datasets independent of the user-input exam data and (2) synthetic FPR, whose null data are simulated from a copula-based probabilistic model, which is first fitted to the user-input exam data and then modified to have no collusion. On 34 unproctored exam datasets, including two benchmark datasets with true positives and negatives verified by textural analysis, we demonstrate that Q-SID is a collusion detection algorithm with powerful and robust performance across exam formats, numbers of questions and students, and exam complexity.
Background General translational cis-elements are present in the mRNAs of all genes and affect the recruitment, assembly, and progress of preinitiation complexes and the ribosome under many physiological states. These elements include mRNA folding, upstream open reading frames, specific nucleotides flanking the initiating AUG codon, protein coding sequence length, and codon usage. The quantitative contributions of these sequence features and how and why they coordinate to control translation rates are not well understood. Results Here, we show that these sequence features specify 42-81% of the variance in translation rates in Saccharomyces cerevisiae, Schizosaccharomyces pombe, Arabidopsis thaliana, Mus musculus, and Homo sapiens. We establish that control by RNA secondary structure is chiefly mediated by highly folded 25-60 nucleotide segments within mRNA 5 ' regions, that changes in tri-nucleotide frequencies between highly and poorly translated 5 ' regions are correlated between all species, and that control by distinct biochemical processes is extensively correlated as is regulation by a single process acting in different parts of the same mRNA. Conclusions Our work shows that general features control a much larger fraction of the variance in translation rates than previously realized. We provide a more detailed and accurate understanding of the aspects of RNA structure that directs translation in diverse eukaryotes. In addition, we note that the strongly correlated regulation between and within cis-control features will cause more even densities of translational complexes along each mRNA and therefore more efficient use of the translation machinery by the cell.
BIC-selected feature sets for “RNAfold” and “5′motifs” features. (XLSX 56 kb)
Identifying functional enhancers elements in metazoan systems is a major challenge. For example, large-scale validation of enhancers predicted by ENCODE reveal false positive rates of at least 70%. Here we use the pregrastrula patterning network of Drosophila melanogaster to demonstrate that loss in accuracy in held out data results from heterogeneity of functional signatures in enhancer elements. We show that two classes of enhancer are active during early Drosophila embryogenesis and that by focusing on a single, relatively homogeneous class of elements, over 98% prediction accuracy can be achieved in a balanced, completely held-out test set. The class of well predicted elements is composed predominantly of enhancers driving multi-stage, segmentation patterns, which we designate segmentation driving enhancers (SDE). Prediction is driven by the DNA occupancy of early developmental transcription factors, with almost no additional power derived from histone modifications. We further show that improved accuracy is not a property of a particular prediction method: after conditioning on the SDE set, naïve Bayes and logistic regression perform as well as more sophisticated tools. Applying this method to a genome-wide scan, we predict 1,640 SDEs that cover 1.6% of the genome, 916 of which are novel. An analysis of 32 novel SDEs using wholemount embryonic imaging of stably integrated reporter constructs chosen throughout our prediction rank-list showed >90% drove expression patterns. We achieved 86.7% precision on a genome-wide scan, with an estimated recall of at least 98%, indicating high accuracy and completeness in annotating this class of functional elements. Significance Statement We demonstrate a high accuracy method for predicting enhancers genome wide with > 85% precision as validated by transgenic reporter assays in Drosophila embryos. This is the first time such accuracy has been achieved in a metazoan system, allowing us to predict with high-confidence 1640 enhancers, 916 of which are novel. The predicted enhancers are demarcated by heterogeneous collections of epigenetic marks; many strong enhancers are free from classical indicators of activity, including H3K27ac, but are bound by key transcription factors. H3K27ac, often used as a one-dimensional predictor of enhancer activity, is an uninformative parameter in our data.
Translation rate per mRNA molecule correlates positively with mRNA abundance. As a result, protein levels do not scale linearly with mRNA levels, but instead scale with the abundance of mRNA raised to the power of an “amplification exponent”. Here we show that to quantitate translational control, the translation rate must be decomposed into two components. One, TRmD, depends on the mRNA level and defines the amplification exponent. The other, TRmIND, is independent of mRNA amount and impacts the correlation coefficient between protein and mRNA levels. We show that in S. cerevisiae TRmD represents ∼20% of the variance in translation and directs an amplification exponent of 1.20 with a 95% confidence interval [1.14, 1.26]. TRmIND constitutes the remaining ∼80% of the variance in translation and explains ∼5% of the variance in protein expression. We also find that TRmD and TRmIND are preferentially determined by different mRNA sequence features: TRmIND by the length of the open reading frame and TRmD both by a ∼60 nucleotide element that spans the initiating AUG and by codon and amino acid frequency. Our work provides more appropriate estimates of translational control and implies that TRmIND is under different evolutionary selective pressures than TRmD.
Numerous affinity purification-mass spectrometry (AP-MS) and yeast two-hybrid screens have each defined thousands of pairwise protein-protein interactions (PPIs), most of which are between functionally unrelated proteins. The accuracy of these networks, however, is under debate. Here, we present an AP-MS survey of the bacterium Desulfovibrio vulgaris together with a critical reanalysis of nine published bacterial yeast two-hybrid and AP-MS screens. We have identified 459 high confidence PPIs from D. vulgaris and 391 from Escherichia coli. Compared with the nine published interactomes, our two networks are smaller, are much less highly connected, and have significantly lower false discovery rates. In addition, our interactomes are much more enriched in protein pairs that are encoded in the same operon, have similar functions, and are reproducibly detected in other physical interaction assays than the pairs reported in prior studies. Our work establishes more stringent benchmarks for the properties of protein interactomes and suggests that bona fide PPIs much more frequently involve protein partners that are annotated with similar functions or that can be validated in independent assays than earlier studies suggested.
Identifying protein-protein interactions (PPIs) at an acceptable false discovery rate (FDR) is challenging. Previously we identified several hundred PPIs from affinity purification - mass spectrometry (AP-MS) data for the bacteria Escherichia coli and Desulfovibrio vulgaris. These two interactomes have lower FDRs than any of the nine interactomes proposed previously for bacteria and are more enriched in PPIs validated by other data than the nine earlier interactomes. To more thoroughly determine the accuracy of ours or other interactomes and to discover further PPIs de novo, here we present a quantitative tagless method that employs iTRAQ MS to measure the copurification of endogenous proteins through orthogonal chromatography steps. 5273 fractions from a four-step fractionation of a D. vulgaris protein extract were assayed, resulting in the detection of 1242 proteins. Protein partners from our D. vulgaris and E. coli AP-MS interactomes copurify as frequently as pairs belonging to three benchmark data sets of well-characterized PPIs. In contrast, the protein pairs from the nine other bacterial interactomes copurify two- to 20-fold less often. We also identify 200 high confidence D. vulgaris PPIs based on tagless copurification and colocalization in the genome. These PPIs are as strongly validated by other data as our AP-MS interactomes and overlap with our AP-MS interactome for D.vulgaris within 3% of expectation, once FDRs and false negative rates are taken into account. Finally, we reanalyzed data from two quantitative tagless screens of human cell extracts. We estimate that the novel PPIs reported in these studies have an FDR of at least 85% and find that less than 7% of the novel PPIs identified in each screen overlap. Our results establish that a quantitative tagless method can be used to validate and identify PPIs, but that such data must be analyzed carefully to minimize the FDR.
Transcription, not translation, chiefly determines protein abundance in mammals [Also see Research Article by Jovanovic et al. ]
In mathematical modeling of genetic regulation networks, the classical assumption is that the direction of regulation by a regulator (activation or repression) on a specific target gene is fixed. However experimental studies have suggested that the direction of gene regulation might be concentration-dependent. Earlier modeling work assumed only specific transcription factors behave in this concentration-dependent manner. Nevertheless, based on our in vivo DNA occupancy measurements of transcription factors and target gene mRNA expressions, we propose that concentration-dependent activation or repression might be a more general phenomenon. More specifically, we assume a regulator is activating at low concentrations and repressing at high concentrations, and we use the regulation of even-skipped stripe 2 (eve2) during Stage 5 of Drosophila embryogenesis as a case study. Our general approach is to describe the dynamic evolution of target mRNA concentration using a set of parametric Ordinary Differential Equations, which incorporate target gene mRNA transcription and degradation, and transcription factor expression and in vivo DNA occupancies. We show our model is better at predicting regulator importance and the effect of transcription factor mutations than two commonly used monotone models, where each regulator either activates or represses the target gene, but not both. In addition, our model generates spatial-temporal maps of factor activity, highlighting the times and spatial locations at which different transcription factors might regulate target gene expression levels. Finally, we use our model to design mutation experiments that help us to verify our assumption of concentration-dependent regulation by determining whether transcription factors that are currently thought to either only activate or repress eve2 in fact act in a concentration-dependent manner.
Large scale surveys in mammalian tissue culture cells suggest that the protein expressed at the median abundance is present at 8,000–16,000 molecules per cell and that differences in mRNA expression between genes explain only 10–40% of the differences in protein levels. We find, however, that these surveys have significantly underestimated protein abundances and the relative importance of transcription. Using individual measurements for 61 housekeeping proteins to rescale whole proteome data from Schwanhausser et al. (2011), we find that the median protein detected is expressed at 170,000 molecules per cell and that our corrected protein abundance estimates show a higher correlation with mRNA abundances than do the uncorrected protein data. In addition, we estimated the impact of further errors in mRNA and protein abundances using direct experimental measurements of these errors. The resulting analysis suggests that mRNA levels explain at least 56% of the differences in protein abundance for the 4,212 genes detected by Schwanhausser et al. (2011), though because one major source of error could not be estimated the true percent contribution should be higher. We also employed a second, independent strategy to determine the contribution of mRNA levels to protein expression. We show that the variance in translation rates directly measured by ribosome profiling is only 12% of that inferred by Schwanhausser et al. (2011), and that the measured and inferred translation rates correlate poorly (R2 = 0.13). Based on this, our second strategy suggests that mRNA levels explain ∼81% of the variance in protein levels. We also determined the percent contributions of transcription, RNA degradation, translation and protein degradation to the variance in protein abundances using both of our strategies. While the magnitudes of the two estimates vary, they both suggest that transcription plays a more important role than the earlier studies implied and translation a much smaller role. Finally, the above estimates only apply to those genes whose mRNA and protein expression was detected. Based on a detailed analysis by Hebenstreit et al. (2012), we estimate that approximately 40% of genes in a given cell within a population express no mRNA. Since there can be no translation in the absence of mRNA, we argue that differences in translation rates can play no role in determining the expression levels for the ∼40% of genes that are non-expressed.
We present an algorithm for the per-voxel semantic segmentation of a three-dimensional volume. At the core of our algorithm is a novel "pyramid context" feature, a descriptive representation designed such that exact per-voxel linear classification can be made extremely efficient. This feature not only allows for efficient semantic segmentation but enables other aspects of our algorithm, such as novel learned features and a stacked architecture that can reason about self-consistency. We demonstrate our technique on 3D fluorescence microscopy data of Drosophila embryos for which we are able to produce extremely accurate semantic segmentations in a matter of minutes, and for which other algorithms fail due to the size and high-dimensionality of the data, or due to the difficulty of the task.
Animals comprise dynamic three-dimensional arrays of cells that express gene products in intricate spatial and temporal patterns that determine cellular differentiation and morphogenesis. A rigorous understanding of these developmental processes requires automated methods that quantitatively record and analyze complex morphologies and their associated patterns of gene expression at cellular resolution. Here we summarize light microscopy-based approaches to establish permanent, quantitative datasets-atlases-that record this information. We focus on experiments that capture data for whole embryos or large areas of tissue in three dimensions, often at multiple time points. We compare and contrast the advantages and limitations of different methods and highlight some of the discoveries made. We emphasize the need for interdisciplinary collaborations and integrated experimental pipelines that link sample preparation, image acquisition, image analysis, database design, visualization, and quantitative analysis. (C) 2013 Wiley Periodicals, Inc.
SysPatho Workshop St. Petersburg, Tsarskoe Selo, Russia September 11-14, 2012 St. Petersburg, 2012 INTERNATIONAL PROGRAM COMMITTEE Prof. Dr. Roland Eils Heidelberg University, Germany Prof. Dr. Sergei Inge-Vechtomov St. Petersburg State University, Russia Prof. Dr. Nikolay Kolchanov Institute of Cytology and Genetics, Russian Academy of Sciences, Russia Prof. Dr. Lars Kaderali Dresden University of Technology, Germany Prof. Dr. Maria Samsonova St. Petersburg State Polytechnical University, Russia Prof. Dr. Alexander Samsonov The Ioffe Phycical Technical Institute, Russian Academy of Sciences, Russia Prof. Dr. Inna Lavrik Otto-von-Guericke-University, Magdeburg, Germany SPONSOR ORGANIZATIONS HOST SUPPORTERS SysPatho Project – New Algorithms for Host Pathogen Systems Biology EU Seventh Framework Program St. Petersburg State Polytechnical University Ioffe Physical-Technical Institute of the Russian Academy of Sciences
A Systematic Evolution of Ligands by EXponential enrichment (SELEX) experiment begins in round one with a random pool of oligonucleotides in equilibrium solution with a target. Over a few rounds, oligonucleotides having a high affinity for the target are selected. Data from a high throughput SELEX experiment consists of lists of thousands of oligonucleotides sampled after each round. Thus far, SELEX experiments have been very good at suggesting the highest affinity oligonucleotide, but modeling lower affinity recognition site variants has been difficult. Furthermore, an alignment step has always been used prior to analyzing SELEX data.We present a novel model, based on a biochemical parametrization of SELEX, which allows us to use data from all rounds to estimate the affinities of the oligonucleotides. Most notably, our model also aligns the oligonucleotides. We use our model to analyze a SELEX experiment containing double stranded DNA oligonucleotides and the transcription factor Bicoid as the target. Our SELEX model outperformed other published methods for predicting putative binding sites for Bicoid as indicated by the results of an in-vivo ChIP-chip experiment.
Three-dimensional gene expression PointCloud data generated by the Berkeley Drosophila TranscriptionNetwork Project (BDTNP) provides quantitative information about the spatial and temporal expression of genes in early Drosophila embryos at cellular resolution. The BDTNP team visualizes and analyzes Point- Cloud data using the software application PointCloudXplore (PCX). To maximize the impact of novel, complex data sets, such as PointClouds, the data needs to be accessible to biologists and comprehensible to developers of analysis functions.We address this challenge by linking PCX and Matlab® via a dedicated interface, thereby providing biologists seamless access to advanced data analysis functions and giving bioinformatics researchers the opportunity to integrate their analysis directly into the visualization application. To demonstrate the usefulness of this approach, we computationally model parts of the expression pattern of the gene even skipped using a genetic algorithm implemented in Matlab and integrated into PCX via our Matlab interface.
Immunoprecipitation of cross-linked chromatin in combination with microarrays (ChIP-chip) or ultra high-throughput sequencing (ChIP-seq) is widely used to map genome-wide in vivo transcription factor binding. Both methods employ initial steps of in vivo cross-linking, chromatin isolation, DNA fragmentation, and immunoprecipitation. For ChIP-chip, the immunoprecipitated DNA samples are then amplified, labeled, and hybridized to DNA microarrays. For ChIP-seq, the immunoprecipitated DNA is prepared for a sequencing library, and then the library DNA fragments are sequenced using ultra high-throughput sequencing platform. The protocols described here have been developed for ChIP-chip and ChIP-seq analysis of sequence-specific transcription factor binding in Drosophila embryos. A series of controls establish that these protocols have high sensitivity and reproducibility and provide a quantitative measure of relative transcription factor occupancy. The quantitative nature of the assay is important because regulatory transcription factors bind to highly overlapping sets of thousands of genomic regions and the unique regulatory specificity of each factor is determined by relative moderate differences in occupancy between factors at commonly bound regions.
In animals, each sequence-specific transcription factor typically binds to thousands of genomic regions in vivo. Our previous studies of 20 transcription factors show that most genomic regions bound at high levels in Drosophila blastoderm embryos are known or probable functional targets, but genomic regions occupied only at low levels have characteristics suggesting that most are not involved in the cis -regulation of transcription. Here we use transgenic reporter gene assays to directly test the transcriptional activity of 104 genomic regions bound at different levels by the 20 transcription factors. Fifteen genomic regions were selected based solely on the DNA occupancy level of the transcription factor Kruppel. Five of the six most highly bound regions drive blastoderm patterns of reporter transcription. In contrast, only one of the nine lowly bound regions drives transcription at this stage and four of them are not detectably active at any stage of embryogenesis. A larger set of 89 genomic regions chosen using criteria designed to identify functional cis -regulatory regions supports the same trend: genomic regions occupied at high levels by transcription factors in vivo drive patterned gene expression, whereas those occupied only at lower levels mostly do not. These results support studies that indicate that the high cellular concentrations of sequence-specific transcription factors drive extensive, low-occupancy, nonfunctional interactions within the accessible portions of the genome.
Because of the increasing diversity of data sets and measurement techniques in biology, a growing spectrum of modeling methods is being developed. It is generally recognized that it is critical to pick the appropriate method to exploit the amount and type of biological data available for a given system. Here, we describe a method for use in situations where temporal data from a network is collected over multiple time points, and in which little prior information is available about the interactions, mathematical structure, and statistical distribution of the network. Our method results in models that we term Nonparametric exterior derivative estimation Ordinary Differential Equation (NODE) model's. We illustrate the method's utility using spatiotemporal gene expression data from Drosophila melanogaster embryos. We demonstrate that the NODE model's use of the temporal characteristics of the network leads to quantifiable improvements in its predictive ability over nontemporal models that only rely on the spatial characteristics of the data. The NODE model provides exploratory visualizations of network behavior and structure, which can identify features that suggest additional experiments. A new extension is also presented that uses the NODE model to generate a comb diagram, a figure that presents a list of possible network structures ranked by plausibility. By being able to quantify a continuum of interaction likelihoods, this helps to direct future experiments.
John-Marc Chandonia合作论文数ASTRAL6
Maxim Shatsky合作论文数Beckman Coulter Diagnostics4