As the primary genetic determinant of immune recognition of self and non‐self, the hyperpolymorphic HLA genes play key roles in disease association and transplantation. The large, variably sized HLA class II genes have historically been less well characterized than the shorter HLA class I genes. Here, we have used Pacific Biosciences Single Molecule Real‐Time (SMRT®) DNA sequencing to perform four‐field resolution HLA typing of HLA‐DRB1/3/4/5, ‐DQA1, ‐DQB1, ‐DPA1 and ‐DPB1 from a panel of 181 B‐lymphoblastoid cell lines from the International HLA and Immunogenetics Workshops. By interrogating all exons, introns, and the untranslated regions of these important reference cells, we have improved their HLA typing resolution on the IPD‐IMGT/HLA database. We observed widespread non‐coding polymorphism, with over twice as many unique genomic sequences identified compared with coding sequences (CDS). We submitted 263 unique sequences to the IPD‐IMGT/HLA Database, often from multiple cell lines, including 114 confirmations of existing alleles, of which 30 were also extensions to full‐length genomic sequences where only CDS was available previously. A total of 149 novel alleles were identified, largely differing from their closest reference allele sequences by a single nucleotide polymorphism (SNP). However, some highly divergent alleles were deemed to be recombinants, only detectable by full‐length sequencing with long, phased reads. The fourth‐field variation we observed allowed fine mapping of linkage disequilibrium patterns and haplotypes to particular ancestries. This study has highlighted the under‐appreciated non‐coding diversity in HLA class II genes, with potential implications for population genetic and clinical studies.
We have developed a genotyping assay that produces fully phased, unambiguous HLA-E genotyping using Pacific Biosciences' single molecule real-time DNA sequencing. In total 212 cell lines were genotyped, including the panel of 107 established at the 10th International Histocompatibility Workshop. Our results matched the previously known HLA-E genotype in 94 (44.3%) cell lines, in all cases either improving or equalling previous genotyping resolution. Three (1.4%) cells had discrepant HLA-E genotyping data and 115 (54.2%) had no previous HLA-E data. The HLA-E genotypes for four (1.9%) cell lines resulted in a change of zygosity by identifying two distinct haplotypes. We discovered eight novel HLA-E alleles, extended the known reference sequence of seven and confirmed the existence of a further 10.
While the success of allogeneic stem cell transplantation depends on a high degree of HLA compatibility between donor and patient, finding a suitable donor remains challenging due to the hyperpolymorphic nature of HLA genes. We calculated high‐resolution allele, haplotype and phenotype frequencies for HLA‐A, ‐C, ‐B, ‐DRB1 and ‐DQB1 for 10 subpopulations of the Anthony Nolan (AN) register using an in‐house expectation‐maximisation (EM) algorithm run on mixed resolution HLA data, covering 676 155 individuals. Sample sizes range from 599 410 for British/Irish North West European (BINWE) individuals, the largest subpopulation in the United Kingdom to 1105 for the British Bangladeshi population. Calculation of genetic distance between the subpopulations based on haplotype frequencies shows three broad clusters, each following a major continental group: European, African and Asian. We further analysed the HLA haplotype and phenotype diversity of each subpopulation, and found that 35.52% of BINWE individuals ranging to 98.34% of Middle Eastern individuals on the register had a unique phenotype within their subpopulation. These analyses and the allele, haplotype and phenotype frequency data of the subpopulation on the AN register are a valuable resource in understanding the HLA diversity in the United Kingdom and can be used to improve the accuracy of match likelihoods and to inform future donor recruitment strategies.
Within an infected individual, influenza virus exists as a heterogeneous population of variants. When representing the viral population as a consensus sequence, information about minority variants is lost. However, using next generation sequencing (NGS), it is possible to identify nucleotide substitutions which segregate at low frequencies in the viral population, and can give insight into the within-host processes that drive the virus’s evolution, and is a step towards understanding the dynamics of the disease. During the course of an infection, mutations may occur, and at each segregating site, the frequency of the derived allele in the population will fluctuate. We develop a method which can use information about the relative frequencies of mutations in NGS data from a viral population sampled at multiple time points, to infer past population dynamics with a Bayesian skyline model. By using coalescent theory, we analytically derive the joint allele frequency spectrum for a population across multiple time points, and relate this to the coalescent intervals generated from the skyline model. We demonstrate the model on data taken from populations of equine influenza virus sampled during an infection, and show that it is possible to infer a posterior distribution of effective viral population size through time. We also show how the model can be used to infer the probability that a mutation occurred within-host, as opposed to being an ancestral mutation which occurred prior to infection.Author Summary When a host is infected by a virus, many particles of the infecting agent enter the body of the host. This viral population is composed of many closely related viruses that continue diversifying by mutating while reproducing in the host. New sequencing technologies allow the quantifying of the proportion of the different variants present in the host at a particular time. Unfortunately, the data resulting from such sequencing techniques are difficult to interpret as they consist of many unlinked copies of relatively small fragments of genetic code distributed along the genome of the virus.We designed a method combining models of virus genealogies and frequency of mutations appearing in the data to reconstruct the variation of the viral population inside the host. It also allows us to time the apparition of particular variants. This could be useful to detect if a particular mutation (e.g. providing drug resistance) has appeared in host or was circulating before. We applied our method to data of within-host evolution of equine influenza.
INTRODUCTION The introduction of a novel approach to HLA typing, incorporating the use of Single Molecule Real Time (SMRT®) Sequencing on the Pacific Biosciences RS II platform raises a number of bioinformatics challenges. This platform enables the sequencing of fully phased long reads that are capable of spanning the full length of HLA class I and the majority of class II genes to produce allele resolution typing. Throughput rate is increased further by barcoding samples, allowing a high multiplex set-up to reduce the financial and environmental cost. The high throughput produces very large data-sets, known typically as “big data”, which creates challenges for conventional bioinformatics analysis and interpretation.
Kaposi’s sarcoma (KS) is a mesenchymal tumor, caused by Human herpesvirus 8 (HHV8) with molecular and cytogenetic changes poorly understood. To gain further insight on the underlying molecular changes in KS, we performed microRNA (miRNA) microarray analysis of 17 Kaposi’s sarcoma specimens. Three normal skin specimens were used as controls. The most significant differentially expressed miRNA were confirmed by quantitative reverse transcriptase polymerase chain reaction (RT-PCR). We detected in KS versus normal skin 185 differentially expressed miRNAs, 76 were upregulated and 109 were downregulated. The most significantly downregulated miRNAs were miR-99a , miR-200 family, miR-199b-5p , miR-100 and miR-335 , whereas kshv-miR-K12-4-3p , kshv-miR-K12-1, kshv-miR-K12-2, kshv-miR-K12-4-5p and kshv-miR-K12-8 were significantly upregulated. High expression levels of kshv-miR-K12-1 ( p = 0.004) and kshv-miR-K12-4-3p ( p = 0.001) was confirmed by RT-PCR. The predicted target genes for differentially expressed miRNAs included genes which are involved in a variety of cellular processes such as angiogenesis (i.e. THBS1) and apoptosis (i.e. CASP3, MCL1), suggesting a role for these miRNAs in Kaposi’s sarcoma pathogenesis.
In many domains data items are represented by vectors of counts; count data arises, for example, in bioinformatics or analysis of text documents represented as word count vectors. However, often the amount of data available from an interesting data source is too small to model the data source well. When several data sets are available from related sources, exploiting their similarities by transfer learning can improve the resulting models compared to modeling sources independently. We introduce a Bayesian generative transfer learning model which represents similarity across document collections by sparse sharing of latent topics controlled by an Indian buffet process. Unlike a prominent previous model, hierarchical Dirichlet process (HDP) based multi-task learning, our model decouples topic sharing probability from topic strength, making sharing of low-strength topics easier. In experiments, our model outperforms the HDP approach both on synthetic data and in first of the two case studies on text collections, and achieves similar performance as the HDP approach in the second case study.
Background Xenografts have been shown to provide a suitable source of tumor tissue for molecular analysis in the absence of primary tumor material. We utilized ES xenograft series for integrated microarray analyses to identify novel biomarkers. Method Microarray technology (array comparative genomic hybridization (aCGH) and micro RNA arrays) was used to screen and identify copy number changes and differentially expressed miRNAs of 34 and 14 passages, respectively. Incubated cells used for xenografting (Passage 0) were considered to represent the primary tumor. Four important differentially expressed miRNAs (miR-31, miR-31*, miR-145, miR-106) were selected for further validation by real time polymerase chain reaction (RT-PCR). Integrated analysis of aCGH and miRNA data was performed on 14 xenograft passages by bioinformatic methods. Results The most frequent losses and gains of DNA copy number were detected at 9p21.3, 16q and at 8, 15, 17q21.32-qter, 1q21.1-qter, respectively. The presence of these alterations was consistent in all tumor passages. aCGH profiles of xenograft passages of each series resembled their corresponding primary tumors (passage 0). MiR-21, miR-31, miR-31*, miR-106b, miR-145, miR-150*, miR-371-5p, miR-557 and miR-598 showed recurrently altered expression. These miRNAS were predicted to regulate many ES-associated genes, such as genes of the IGF1 pathway, EWSR1, FLI1 and their fusion gene ( EWS-FLI1 ). Twenty differentially expressed miRNAs were pinpointed in regions carrying altered copy numbers. Conclusion In the present study, ES xenografts were successfully applied for integrated microarray analyses. Our findings showed expression changes of miRNAs that were predicted to regulate many ES associated genes, such as IGF1 pathway genes, FLI1, EWSR1 , and the EWS-FLI1 fusion genes.
Count data arises for example in bioinformatics or analy- sis of text documents represented as word count vectors. With several data sets available from related sources, exploiting their similarities by transfer learning can improve models compared to modeling sources in- dependently. We introduce a Bayesian generative transfer learning model which represents similarity across document collections by sparse sharing of latent topics controlled by an Indian Buffet Process. Unlike Hierarchi- cal Dirichlet Process based multi-task learning, our model decouples topic sharing probability from topic strength, making sharing of low-strength topics easier, and outperforms the HDP approach in experiments.
Given a learning task for a data set, learning it together with related tasks (data sets) can improve performance. Gaussian process models have been applied to such multi-task learning scenarios, based on joint priors for functions underlying the tasks. In previous Gaussian process approaches, all tasks have been assumed to be of equal importance, whereas in transfer learning the goal is asymmetric: to enhance performance on a target task given all other tasks. In both settings, transfer learning and joint modelling, negative transfer is a key problem: performance may actually decrease if the tasks are not related closely enough. In this paper, we propose aGaussian processmodel for the asymmetric setting, which learns to "explain away" non-related variation in the additional tasks, in order to focus on improving performance on the target task. In experiments, our model improves performance compared to single-task learning, symmetric multi-task learning using hierarchical Dirichlet processes, and transfer learning based on predictive structure learning.
Sauf mention contraire ci-dessus, le contenu de cette notice bibliographique peut être utilisé dans le cadre d’une licence CC BY 4.0 Inist-CNRS/Unless otherwise stated above, the content of this bibliographic record may be used under a CC BY 4.0 licence by Inist-CNRS/A menos que se haya señalado antes, el contenido de este registro bibliográfico puede ser utilizado al amparo de una licencia CC BY 4.0 Inist-CNRS
Helena Aidos Antti Ajanki Alessia Albanese Sheng Hua Bao Andras Benczur Domenico Beneventano Tim Brailsford Guillaume Cabanac Ke Ke Cai Bin Cao Mark Carman Paolo Casoto Yuming Chen Chien Chin Chen Yi Cheng Flavio Chierichetti Hanachi Chihab Helder Coelho Enrique Munos de Cote Faezeh Ensan Timur Fayruzov Alessio Ferone Piero Fraternali Shima Gerani Jean Marie Gilliot Roberto De Prisco Matteo Di Gioia Giorgio Maria Di Nunzio Anastasios Gounaris Francisco Grimaldo Allel Hadjali Rabab Hayek Nicolas Hernandez Derek Hao Hu Jeroen Janssen Melih Kandemir Arto Klami Monica Landoni Jens Lechtenbörger Danielle Lee Gayle Leen Cane Wing-ki Leung Peng Li Wen Li Bo Liu Tomek Loboda Eric Louie Hiep Luong Antonio Maratea José Martinez Motohiro Mase Li Meng Duoqian Miao Luis Moniz Maurizio Montagnuolo Luis Morgado Takeshi Morita Guillermo Morales Luna Masayuki Okabe Nicola Orio Fabrizio Orlandi Salvatore Orlando Mirko Orsini Denis Parra Marco Pellegrini Wei Peng Quang-Khai Pham Fabien Picarougne Antoine Pigeau John A. Piorkowski Olivier Pivert Ajith Kodakateri Pudhiyaveetil Daniel Rocacher Régis Saint-Paul Antonio Sala Giuseppe Salvi Yacine Sam Karen Sauvagnat Steven Schockaert Bo Shao Gianmaria Silvello Fabrizio Silvestri Janne Sinkkonen Laurianne Sitbon Gavin Smith Serena Sorrentino Mirco Speretta Sebastian Stein Yu-Wei Sung Gunnar Thies Paulo Trigo Luca Vassena Patricia Victor Maurizio Vincini Xin Wang Qiang Wang Dingding Wang Chi Wang Xian Wu Nanhong Ye Erliang Zeng Xiao Xun Zhang Judy Zhao Jie Zhou
This tutorial first covers a variety of aspects of text and document modeling in order to motivate the use of topic models. It will briefly consider aspects of named entity recognition, natural language parsing, information retrieval, and look at some standard algorithms for addressing these problems. This will not be an extensive coverage but rather a brief look at some of the issues, problems, methods and software. Second, the tutorial will present the recent area of topic models, for instance Latent Dirichlet Allocation, Non-negative Matrix Factorisation, and related methods. These will be reviewed and examples given using the author's software DCA. Some basic experimental methods and variations will also be covered. This will draw on material from the author's papers here as well as the broader literature. Short Biography: Dr. Wray Buntine relocated from Helsinki Finland to Canberra Australia in April 2007 to join NICTA. Recently of University of Helsinki and HIIT, and previously NASA Ames Research Center, UC Berkeley, and Google, he is known for his theoretical and applied work in machine learning, probabilistic methods, and discrete component methods for text analysis. He also has a development background, and has produced two GPL'd suites of machine learning software, IND for decision trees in the 90's and MPCA for component analysis mid 2000's. His PhD.\ in the area of decision trees in machine learning was awarded in 1992, and a Docentship at University of Helsinki in 2006. He is co-programme chair of ECMLPKDD 2009, is on the editorial board for the Data Mining and Knowledge Discovery, New Generation Computing, and Statistics and Computing and has served on a wide variety of conference committees.
C Fyfe合作论文数6
Jaakko Peltonen合作论文数Neural Networks Research Centre, Helsinki University of Technology, P.O. Box 5400, FI-02015 HUT, Finland3