This paper addresses the clustering and classification of active genes during the process of cell division. Cell division ensures the proliferation of cells, but it becomes increasingly abnormal in cancer cells. The genes studied here are described by their expression profiles (i.e. time series) during the cell division cycle. This work focuses on evaluating the efficiency of four major metrics for clustering and classifying genes expression profiles and is based on a random-periods model for the expression of cell-cycle genes. The model accounts for the observed attenuation in cycle amplitude or duration, variations in the initial amplitude, and drift in the expression profiles.
The biological problem of identifying the active genes during the cell division process is addressed. The cell division ensures the proliferation of cells, which is drastically aberrant in cancer cells. The studied genes are described by their expression profiles during the cell division cycle. Commonly, the identification process is a supervised approach based on an a priori set of reference genes, assumed as well-characterizing the cell cycle phases. Each studied gene is then classified by its peak similarity to one pre-specified reference gene. This classical approach suffers from two limitations. On the one hand, there is no consensus between biologists about the set of reference genes to consider for the identification process. On the other hand, the proximity measures used for genes expression profiles are unjustified and mainly based on the expression values regardless of the genes expression behavior. To identify genes expression profiles, a new adaptive clustering approach is proposed which consists of two main points. First, it allows in an unsupervised way the selection of a well-justified set of reference genes, to be compared with the pre-specified ones. Secondly, it enables the users to learn the appropriate proximity measure to use for genes expression data, a measure which will cover both proximity on values and on behavior. The adaptive clustering method is compared to a correlation-based approach through public and simulated genes expression data.
New technologies and equipment allow for mass treatment of samples and research teams share acquired data on an always larger scale. In this context scientists are facing a major data exploitation problem. More precisely, using these data sets through data mining tools or introducing them in a classical experimental approach require a preliminary understanding of the information space, in order to direct the process. But acquiring this grasp on the data is a complex activity, which is seldom supported by current software tools. The goal of this paper is to introduce a solution to this scientific data grasp problem. Illustrated in the Tissue MicroArrays application domain, the proposal is based on the synthesis notion, which is inspired by Information Retrieval paradigms. The envisioned synthesis model gives a central role to the study the researcher wants to conduct, through the task notion. It allows for the implementation of a task-oriented Information Retrieval prototype system. Cases studies and user studies were used to validate this prototype system. It opens interesting prospects for the extension of the model or extensions towards other application domains.
This paper addresses the clustering and classification of active genes during the process of cell division. Cell division ensures the proliferation of cells, but becomes drastically aberrant in cancer cells. The studied genes are described by their expression profiles (i.e. time series) during the cell division cycle. This work focuses on evaluating the efficiency of four major metrics for clustering and classifying gene expression profiles. The study is based on a random-periods model for the expression of cell-cycle genes. The model accounts for the observed attenuation in cycle amplitude or duration, variations in the initial amplitude, and drift in the expression profiles.
Ce travail s'inscrit dans le cadre de l'etude de la division cellulaire assurant la proliferation des cellules. Une meilleure comprehension de ce phenomene biologique necessite l'identification des genes caracterisant chaque phase du cycle cellulaire. Le procede d'identification est generalement base sur un ensemble de genes dits genes de reference, selectionnes experimentalement et consideres comme caracterisant les phases du cycle cellulaire. Les niveaux d'expression des genes etudies sont mesures durant le cycle de la division cellulaire et permettent de construire des profils d'expression. Chaque gene etudie est affecte a la phase du cycle cellulaire correspondant au groupe de genes de reference le plus similaire. Cette approche classique souffre de deux limites. D'une part, les mesures de proximite les plus couramment utilisees entre profils d'expression de genes sont basees sur les ecarts en valeurs sans tenir compte de la forme des profils. D'autre part, dans la litterature il n'y a pas consensus quant a l'ensemble des genes de reference a considerer. Dans cet article, notre but est de proposer une classification adaptative, basee sur un indice de dissimilarite incluant les proximites en valeurs et en forme des profils d'expression de genes, permettant d'identifier les phases d'expression des genes etudies, et de presenter un nouvel ensemble de genes de reference valide par une connaissance biologique.
DNA microarray technology allows to monitor simultaneously the expression levels of thousands of genes during important biological processes and across collections of related experiments. Clustering and classification techniques have proved to be helpful to understand gene function, gene regulation, and cellular processes. However the conventional proximity measures between genes expression data, used for clustering or classification purpose, do not fit gene expression specifications as they are based on the closeness of the expression magnitudes regardless of the overall gene expression profile (shape). We propose in this paper an adaptive dissimilarity index which would cover both values and behavior proximity. The effectiveness of the adaptive dissimilarity index is illustrated through a classification process for identification of genes cell cycle phases.
This paper focuses on the cell division cycle insuring the proliferation of cells and which is dras- tically aberrant in cancer cells. The aim of this biological problem is the identification of genes characterizing each cell cycle phase. The identification process is commonly based on a prior set of well-characterized cell cycle genes, called reference genes. The expression levels of the studied genes are measured during the cell division cycle. Each studied gene is assigned a cell cycle phase by it's peak similarity to the reference genes. This classical approach suffers of two limitations. On the one hand, the most widely used proximity measures between gene expression profiles are based on the closeness of the values regardless to the similarity with respect to (w.r.t.) the genes expression behavior. On the other hand, many different ill-founded sets of reference genes are pro- posed in the literature, and biologists are not agree about those of genes best characterizing the observed cell cycle phases. Our aim in this paper is twice. We propose a new dissimilarity index for gene expression profiles to include both proximity measures w.r.t. values and w.r.t. behavior. An adaptive unsupervised classification, based on the proposed dissimilarity index, is then performed to identify the cell cycle phases of the studied genes. Finally, we propose a new set of reference genes, well-assessed by a biological knowledge.
*869CA9C6EF69F6E9AFC6EF6&C+F FC6 9EDDF69F6 A 6 #(B+F 6E,8E8#C8C 69#BF"F.Une première approche pour caractériser ce processus est proposée, et les connaissances nécessaires sont introduites.Nous présentons ensuite l'architecture du moteur d'adaptation envisagé.Une application au domaine médical, plus précisément à la technologie des Tissue Microarrays, permet d'illustrer l'intérêt de cette démarche.Des éléments de validation sont fournis dans ce domaine à l'aide de l'outil TreeMaps de l'Université du Maryland.ABSTRACT.Beyond classical techniques to explore a document collection such as Information Retrieval and graphical visualisation we present in this paper an analytic view which combines the expression of a goal-centered query and the construction of a virtual multimedia document.Building this synthetic document is considered as a complex adaptation problem.We propose a first approach to characterise this process and introduce the required knowledge.An application to the medical field and in particular the Tissue MicroArrays technology allows us to illustrate the interest of this approach.A preliminary validation of the proposed concepts is performed by means of the TreeMaps tool from the University of Maryland.
The amount of documents available on the Internet makes information retrieval a difficult task. In parallel some efforts towards the Internet personalization allow web pages user-adaptivity. In order to fulfil those objectives current works rely on adaptive information systems and documents summarizing concepts. In this context most systems intend to represent a diversity through a synthesis or a narrative coherence through spatial organisation. We intend to extract a representative collective presenting a thematic coherence according several analytic points of view in order to generate a synthetic document which could be used for data mining. The considered system will be at first designed for the medical field and in particular in the Tissue Microarray technology field.
In oncology research, Tissue Microarray (TMA) technology allows for the mass treatment of hundreds of tissue samples and rapid visualisation of molecular targets. Since this technique is relatively new, there are very few dedicated information systems and little formalised knowledge about the technique is available. We therefore intend to set up an integrated system around TMA technology that is accessible from the Internet. In particular we intend to set up a multimedia document generation system to assist with TMA design.
Le domaine biomedical dispose de plus en plus de techniques generatrices d'enormes quantites de donnees. C'est le cas de la recherche en oncologie, ou la technique des Tissue MicroArrays (TMA) permet le traitement en masse d'un grand nombre d'echantillons de tissus. Du fait de la jeunesse de cette technique, il existe peu de systemes informatiques adaptes. Nous nous proposons donc de realiser une plate-forme integree, nommee TMA-Explorer, regroupant un ensemble d'outils autour de la technologie des TMA. Cette plate-forme vise a etre accessible par l'Internet pour les personnes utilisant la technique des TMA, afin de permettre une capitalisation de connaissances. Entre autres, nous nous proposons de mettre en place un systeme de generation d'un ensemble de documents multimedia, pour aider a la conception des TMA et accompagner la fouille de donnees.