RNA sequencing has the potential to reveal many modalities of transcriptional regulation, such as various splicing phenotypes, but studies on gene regulation are often limited to gene expression due to the complexity of extracting and analyzing multiple RNA phenotypes. Here, we present Pantry, a framework to efficiently generate diverse RNA phenotypes from RNA sequencing data and perform downstream integrative analyses with genetic data. Pantry generates phenotypes from six modalities of transcriptional regulation (gene expression, isoform ratios, splice junction usage, alternative TSS/polyA usage, and RNA stability) and integrates them with genetic data via QTL mapping, TWAS, and colocalization testing. We apply Pantry to Geuvadis and GTEx data, finding that 4768 of the genes with no identified eQTL in Geuvadis have QTL in at least one other transcriptional modality, resulting in a 66% increase in genes over eQTL mapping. We further found that the QTL exhibit modality-specific functional properties that are further reinforced by joint analysis of different RNA modalities. We also show that generalizing TWAS to multiple RNA modalities approximately doubles the discovery of unique gene-trait associations, and enhances identification of regulatory mechanisms underlying GWAS signal in 42% of previously associated gene-trait pairs. Here, the authors present the Pantry framework, which extracts features from RNA sequencing data and performs multimodal genetic analyses. This type of analysis can increase gene-trait associations identified compared to using only expression levels.
Transcriptome data is commonly used to understand genome function via quantitative trait loci (QTL) mapping and to identify the molecular mechanisms driving genome wide association study (GWAS) signals through colocalization analysis and transcriptome-wide association studies (TWAS). While RNA sequencing (RNA-seq) has the potential to reveal many modalities of transcriptional regulation, such as various splicing phenotypes, such studies are often limited to gene expression due to the complexity of extracting and analyzing multiple RNA phenotypes. Here, we present Pantry (Pan-transcriptomic phenotyping), a framework to efficiently generate diverse RNA phenotypes from RNA-seq data and perform downstream integrative analyses with genetic data. Pantry currently generates phenotypes from six modalities of transcriptional regulation (gene expression, isoform ratios, splice junction usage, alternative TSS/polyA usage, and RNA stability) and integrates them with genetic data via QTL mapping, TWAS, and colocalization testing. We applied Pantry to Geuvadis and GTEx data, and found that 4,768 of the genes with no identified expression QTL in Geuvadis had QTLs in at least one other transcriptional modality, resulting in a 66% increase in genes over expression QTL mapping. We further found that QTLs exhibit modality-specific functional properties that are further reinforced by joint analysis of different RNA modalities. We also show that generalizing TWAS to multiple RNA modalities (xTWAS) approximately doubles the discovery of unique gene-trait associations, and enhances identification of regulatory mechanisms underlying GWAS signal in 42% of previously associated gene-trait pairs. We provide the Pantry code, RNA phenotypes from all Geuvadis and GTEx samples, and xQTL and xTWAS results on the web.
Expression Quantitative Trait Loci (eQTLs) are critical to understanding the mechanisms underlying disease-associated genomic loci. Nearly all protein-coding genes in the human genome have been associated with one or more eQTLs. Here we introduce a multi-variant generalization of allelic Fold Change (aFC), aFC-n, to enable quantification of the cis-regulatory effects in multi-eQTL genes under the assumption that all eQTLs are known and conditionally independent. Applying aFC-n to 458,465 eQTLs in the Genotype-Tissue Expression (GTEx) project data, we demonstrate significant improvements in accuracy over the original model in estimating the eQTL effect sizes and in predicting genetically regulated gene expression over the current tools. We characterize some of the empirical properties of the eQTL data and use this framework to assess the current state of eQTL data in terms of characterizing cis-regulatory landscape in individual genomes. Notably, we show that 77.4% of the genes with an allelic imbalance in a sample show 0.5 log2 fold or more of residual imbalance after accounting for the eQTL data underlining the remaining gap in characterizing regulatory landscape in individual genomes. We further contrast this gap across tissue types, and ancestry backgrounds to identify its correlates and guide future studies. Genetic variants that influence gene expression play a major role in human phenotypic variability and disease susceptibility. Here, the authors introduce a computational method to estimate the regulatory effect size in genes with multiple conditionally independent regulatory variants.
Regulation of transcript structure generates transcript diversity and plays an important role in human disease(1-7). The advent oflong-read sequencing technologies offers the opportunity to study the role of genetic variation in transcript structure(8-)(16). In this Article, we present a large human long-read RNA-seq dataset using the Oxford Nanopore Technologies platform from 88 samples from Genotype-Tissue Expression (GTEx) tissues and cell lines, complementing the GTEx resource. We identified just over 70,000 novel transcripts for annotated genes, and validated the protein expression of 10% of novel transcripts. We developed a new computational package, LORALS, to analyse the genetic effects of rare and common variants on the transcriptome by allele-specific analysis of long reads. We characterized allele-specific expression and transcript structure events, providing new insights into the specific transcript alterations caused by common and rare genetic variants and highlighting the resolution gained from long-read data. We were able to perturb the transcript structure upon knockdown of PTBP1, an RNA binding protein that mediates splicing, thereby finding genetic regulatory effects that are modified by the cellular environment. Finally, we used this dataset to enhance variant interpretation and study rare variants leading to aberrant splicing patterns.
Heterogeneous Stock (HS) rats are a genetically diverse outbred rat population that is widely used for studying genetics of behavioral and physiological traits. Mapping Quantitative Trait Loci (QTL) associated with transcriptional changes would help to identify mechanisms underlying these traits. We generated genotype and transcriptome data for five brain regions from 88 HS rats. We identified 21 392 cis-QTLs associated with expression and splicing changes across all five brain regions and validated their effects using allele specific expression data. We identified 80 cases where eQTLs were colocalized with genome-wide association study (GWAS) results from nine physiological traits. Comparing our dataset to human data from the Genotype-Tissue Expression (GTEx) project, we found that the HS rat data yields twice as many significant eQTLs as a similarly sized human dataset. We also identified a modest but highly significant correlation between genetic regulatory variation among orthologous genes. Surprisingly, we found less genetic variation in gene regulation in HS rats relative to humans, though we still found eQTLs for the orthologs of many human genes for which eQTLs had not been found. These data are available from the RatGTEx data portal (RatGTEx.org) and will enable new discoveries of the genetic influences of complex traits.
Human immunodeficiency virus (HIV) infection is associated with an increased risk of non-Hodgkin lymphoma (NHL). Even in the era of suppressive antiretroviral treatment, HIV-infected individuals remain at higher risk of developing NHL compared to the general population. To identify potential genetic risk loci, we performed case-control genome-wide association studies and a meta-analysis across three cohorts of HIV+ patients of European ancestry, including a total of 278 cases and 1924 matched controls. We observed a significant association with NHL susceptibility in the C-X-C motif chemokine ligand 12 (CXCL12) region on chromosome 10. A fine mapping analysis identified rs7919208 as the most likely causal variant (P = 4.77e-11), with the G>A polymorphism creating a new transcription factor binding site for BATF and JUND. These results suggest a modulatory role of CXCL12 regulation in the increased susceptibility to NHL observed in the HIV-infected population.
Despite rapid progress in characterizing the role of host genetics in SARS-Cov-2 infection, there is limited understanding of genes and pathways that contribute to COVID-19. Here, we integrated a genome-wide association study of COVID-19 hospitalization (7,885 cases and 961,804 controls from COVID-19 Host Genetics Initiative) with mRNA expression, splicing, and protein levels (n=18,502). We identified 27 genes related to inflammation and coagulation pathways whose genetically predicted expression was associated with COVID-19 hospitalization. We functionally characterized the 27 genes using phenome- and laboratory-wide association scans in Vanderbilt Biobank (BioVU; n=85,460) and identified coagulation-related clinical symptoms, immunologic, and blood-cell-related biomarkers. We replicated these findings across trans-ethnic studies and observed consistent effects in individuals of diverse ancestral backgrounds in BioVU, pan-UK Biobank, and Biobank Japan. Our study highlights putative causal genes impacting COVID-19 severity and symptomology through the host inflammatory response.
Fast and easy access to a wide range of documents in various languages, in conjunction with the wide availability of translation and editing tools, has led to the need to develop effective tools for detecting cross-lingual plagiarism. Given a suspicious document, cross-lingual plagiarism detection comprises two main subtasks: retrieving documents that are candidate sources for that document and analysing those candidates one by one to determine their similarity to the suspicious document. In this article, we examine the second subtask, also called the detailed analysis subtask, where the goal is to align plagiarised fragments from source and suspicious documents in different languages. Our proposed approach has two main steps: the first step tries to find candidate plagiarised fragments and focuses on high recall, followed by a more precise similarity analysis based on dynamic text alignment that will filter the results by finding alignments between the identified fragments. With these two steps, the proximity of the terms will be considered in different levels of granularity. In both steps, our approach uses a dictionary to obtain translations of individual terms instead of using a machine translation system to convert longer passages from one language to another. We used a weighting scheme to distinct multiple translations of the terms. Experimental results show that our method outperforms the methods used by the systems that achieved the best results in the PAN-2012 and PAN-2014 competitions.
Rapid growth of documents and the increased accessibility of electronic documents lead to the need to develop effective tools for detecting plagiarised texts. The task of plagiarism detection entails two main subtasks, suspicious candidate retrieval and pairwise document similarity analysis also called detailed analysis. In this paper we focus on the second subtask. We will report our monolingual plagiarism detection system which is used to process the Persian plagiarism corpus for the task of pairwise document similarity. To retrieve plagiarised passages this paper presents a pairwise plagiarism detection algorithm based on a vector space model considering the proximity of the terms. The proposed framework is applicable in any language and it could also adapted for cross language domain. We evaluate the performance in terms of precision, recall, granularity and Plagdet metrics.
The Web offers fast and easy access to a wide range of documents in various languages, and translation and editing tools provide the means to create derivative documents fairly easily. This leads to the need to develop effective tools for detecting cross-language plagiarism. Given a suspicious document, cross-language plagiarism detection comprises two main subtasks: retrieving documents that are candidate sources for that document and analyzing those candidates one by one to determine their similarity to the suspicious document. In this paper we focus on the second subtask and introduce a novel approach for assessing cross-language similarity between texts for detecting plagiarized cases. Our proposed approach has two main steps: a vector-based retrieval framework that focuses on high recall, followed by a more precise similarity analysis based on dynamic text alignment. Experiments show that our method outperforms the methods of the best results in PAN-2012 and PAN-2014 in terms of plagdet score. We also show that aligning n-gram units, instead of aligning complete sentences, improves the accuracy of detecting plagiarism.
The task of plagiarism detection entails two main steps, suspicious candidate retrieval and pairwise document similarity analysis also called detailed analysis. In this paper we focus on the second subtask. We will report our monolingual plagiarism detection system which is used to process the Persian plagiarism corpus for the task of pairwise document similarity. To retrieve plagiarised passages a plagiarism detection method based on vector space model, insensitive to context reordering, is presented. We evaluate the performance in terms of precision, recall, granularity and plagdet metrics.
The rapid growth of documents in different languages, the increased accessibility of electronic documents, and the availability of translation tools have caused cross-lingual plagiarism detection research area to receive increasing attention in recent years. The task of cross-language plagiarism detection entails two main steps: candidate retrieval and assessing pairwise document similarity. In this paper we examine candidate retrieval, where the goal is to find potential source documents of a suspicious text. Our proposed method for cross-language plagiarism detection is a keyword-focused approach. Since plagiarism usually happens in parts of the text, there is a requirement to segment the texts into fragments to detect local similarity. Therefore we propose a topic-based segmentation algorithm to convert the suspicious document to a set of related passages. After that, we use a proximity-based Model to retrieve documents with the best matching passages. Experiments show promising results for this important phase of cross-language plagiarism detection. (C) 2016 Elsevier Ltd. All rights reserved.
With advancements in industry and information technology, large volumes of electronic documents such as newspapers, emails, weblogs, and theses are produced daily. Producing electronic documents has considerable benefits such as easy organizing and data management. Therefore, existence of automatic systems such as spell and grammar-checker/correctors can help to improve their quality. In this article, the development of an automatic spelling, grammatical and real-word error checker for Persian (Farsi) language, named Vafa Spell-Checker, is explained. Different kinds of errors in a text can be categorized into spelling, grammatical, and real-word errors. Vafa Spell-Checker is a hybrid system in which both rule-based and statistical approaches are used to detect/correct whole types of errors. The detection and correction phases of spelling and real-word errors are fully statistical, while for the grammar-checker, a rule-based approach is proposed. Vafa Spell-Checker attempts to process these kinds of error types in an integrated system for Persian language. The results on the real-world collected test set indicate that continuing the work on grammar-checker requires statistical approaches. Evaluation results with respect to F0.5 measure for spell-checker, grammar-checker, and real-word error checker are about 0.908, 0.452, and 0.187, respectively. Moreover, several free-usable language resources for Persian that are generated during this project are demonstrated in this article. These resources could be used in the further research in Persian language.
Real-word errors or context sensitive spelling errors, are misspelled words that have been wrongly converted into another word of vocabulary. One way to detect and correct real-word errors is using Statistical Machine Translation (SMT), which translates a text containing some real-word errors into a correct text of the same language. In this paper, we improve the results of mentioned SMT system by employing some discourseaware features into a log-linear reranking method. Our experiments on a real-world test data in Persian show an improvement of about 9.5% and 8.5% in the recall of detection and correction respectively. Other experiments on standard English test sets also show considerable improvement of real-word checking results.
Producing electronic rather than paper documents has considerable benefits such as easier organizing and data management. Therefore, existence of automatic writing assistance tools such as spell and grammar checker/correctors can increase the quality of electronic texts by removing noise and correcting the erroneous sentences. Different kinds of errors in a text can be categorized into spelling, grammatical and real‐word errors. In this article, we present a language‐independent approach based on a statistical machine translation framework to develop a proofreading tool, which detects grammatical errors as well as context‐sensitive spelling mistakes (real‐word errors). A hybrid model for grammar checking is suggested by combining the mentioned approach with an existing rule‐based grammar checker. Experimental results on both English and Persian languages indicate that the proposed statistical method and the rule‐based grammar checker are complementary in detecting and correcting syntactic errors. The results of the hybrid grammar checker, applied to some English texts, show an improvement of about 24% with respect to the recall metric with almost similar value for precision. Experiments on real‐world data set show that state‐of‐the‐art results are achieved for grammar checking and context‐sensitive spell checking for Persian language. Copyright © 2012 John Wiley & Sons, Ltd.
Existence of automatic writing assistance tools such as spell and grammar checker/corrector can help in increasing electronic texts with higher quality by removing noises and cleaning the sentences. Different kinds of errors in a text can be categorized into spelling, grammatical and real-word errors. In this article, the concepts of an automatic grammar checker for Persian (Farsi) language, is explained. A statistical grammar checker based on phrasal statistical machine translation (SMT) framework is proposed and a hybrid model is suggested by merging it with an existing rule-based grammar checker. The results indicate that these two approaches are complimentary in detecting and correcting syntactic errors, although statistical approach is able to correct more probable errors. The state-of-the-art results on Persian grammar checking are achieved by using the hybrid model. The obtained recall is about 0.5 for correction and about 0.57 for detection with precision about 0.63.
With improvements in industry and information technology, large volumes of electronic texts such as newspapers, emails, weblogs, books and thesis are produced daily. Producing electrical documents has considerable benefits such as easy organizing and data management. Therefore, existence of automatic systems such as spell and grammar checker/corrector can help in reducing costs and increasing the electronic texts and it will improve the quality of electronic texts. You can input your text and the computer program will point out to you the spelling errors. It may also help with your grammar. Grammatical errors are described as wrong relation between words like subject-verb disagreement or wrong sequence of words like using plural noun where a single noun is needed. Grammar checking phase starts after spell checking is finished. This paper briefly describes the concepts and definition of grammar checkers in general followed by developing the first Persian (Farsi) grammar checker leading to an overview of the error types of Persian language. The proposed system detects and corrects about 20 frequent Persian grammar errors and tested on a sample dataset, retrieved about 70% and 83% accuracy respect to precision and recall metrics.