Background Thanks to the wider spread of high-throughput experimental techniques, biologists are accumulating large amounts of datasets which often mix quantitative and qualitative variables and are not always complete, in particular when they regard phenotypic traits. In order to get a first insight into these datasets and reduce the data matrices size scientists often rely on multivariate analysis techniques. However such approaches are not always easily practicable in particular when faced with mixed datasets. Moreover displaying large numbers of individuals leads to cluttered visualisations which are difficult to interpret. Results We introduced a new methodology to overcome these limits. Its main feature is a new semantic distance tailored for both quantitative and qualitative variables which allows for a realistic representation of the relationships between individuals (phenotypic descriptions in our case). This semantic distance is based on ontologies which are engineered to represent real-life knowledge regarding the underlying variables. For easier handling by biologists, we incorporated its use into a complete tool, from raw data file to visualisation. Following the distance calculation, the next steps performed by the tool consist in (i) grouping similar individuals, (ii) representing each group by emblematic individuals we call archetypes and (iii) building sparse visualisations based on these archetypes. Our approach was implemented as a Python pipeline and applied to a rosebush dataset including passport and phenotypic data. Conclusions The introduction of our new semantic distance and of the archetype concept allowed us to build a comprehensive representation of an incomplete dataset characterised by a large proportion of qualitative data. The methodology described here could have wider use beyond information characterizing organisms or species and beyond plant science. Indeed we could apply the same approach to any mixed dataset.
Background Seedling growth is an early phase of plant development highly susceptible to environmental factors such as soil nitrogen (N) availability or presence of seed-borne pathogens. Whereas N plays a central role in plant-pathogen interactions, its role has never been studied during this early phase for the interaction between Arabidopsis thaliana and Alternaria brassicicola , a seed-transmitted necrotrophic fungus. The aim of the present work was to develop an in vitro monitoring system allowing to study the impact of the fungus on A. thaliana seedling growth, while modulating N nutrition. Results The developed system consists of square plates placed vertically and filled with nutrient agar medium allowing modulation of N conditions. Seeds are inoculated after sowing by depositing a droplet of conidial suspension. A specific semi-automated image analysis pipeline based on the Ilastik software was developed to quantify the impact of the fungus on seedling aerial development, calculating an index accounting for every aspect of fungal impact, namely seedling death, necrosis and developmental delay. The system also permits to monitor root elongation. The interest of the system was then confirmed by characterising how N media composition [0.1 and 5 mM of nitrate (NO 3 − ), 5 mM of ammonium (NH 4 + )] affects the impact of the fungus on three A. thaliana ecotypes. Seedling development was strongly and negatively affected by the fungus. However, seedlings grown with 5 mM NO 3 − were less susceptible than those grown with NH 4 + or 0.1 mM NO 3 − , which differed from what was observed with adult plants (rosette stage). Conclusions The developed monitoring system allows accurate determination of seedling growth characteristics (both on aerial and root parts) and symptoms. Altogether, this system could be used to study the impact of plant nutrition on susceptibility of various genotypes to fungi at the seedling stage.
Urbanisation is usually modelled to account for the trade-off between the rent from an agricultural and urban land-use in a location.In this article, we propose a model that includes a characterisation of the land in respect of not only its economic and physical aspects, but also using variables in relation to landscape perception.To that end, we develop an original two-stage approach consisting of estimating a probability of urbanisation and then taking the uncertainty of urbanisation into account using an internal meta-regression method.The landscape descriptors, constructed based on a textual analysis of the Landscape Atlases, are introduced in this second stage.The application of this method to the urban area of Angers shows the importance of these elements in analysing urbanisation.
HAL is a multi-disciplinary open access archive for the deposit and dissemination of scientific research documents, whether they are published or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et à la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. ELTerm: a terminology module for a plant data management system Lysiane Hauguel, Tanguy Lallemand, Rayan Eid, Fabrice Dupuis, Sylvain Gaillard, Florian Blessing, Sandra Pelletier, Julie Bourbeillon
Visualizing metadata change in networks and / or clusters BIDefI Different visualizations produced: • Two groups of alluvial diagrams representing experiment metadata and gene annotation changes. Representation using: ○ Nodes as rectangle. Their height represents the percentage of representation of this term in relation to all items mapped on this part of GO for each iteration. ○ Links between nodes. Their colors are based on statistical status. • One force directed graph deployed on user request, to display Gene Ontology graph with a given GO Slim term as root and all genes with this GO Slim reference mapped on their GO term. Number of genes mapped on each term is represented by node size. Results Visualizations • To represent high volume of sequential metadata in two dimensions few approaches are possible [6] including: ○ Networks diagram can deal with sequential data but not with an efficiency and readability (iterations using colors or shapes codes). ○ Sequence sunburst is more compact but also less visual. In fact there is no links displayed between nodes. ○ Alluvial diagrams are sometime used to visualize tree paths. Inspired by this approach, this type of visualization has been adopted to show changes of matrix at each iteration. • Final choices: ○ Alluvial diagrams to display changes in networks. ○ Force directed graph deployed to explore a given iteration • Alluvial diagrams built using D3.js [7] version 5 and in particular the sankey module version 0.12.1. Some functions were rewritten to fit with project requirements and in particular variable width of links. Methods • Client side application, need to precompute as possible. Need to minify all JSON datas to limit size of transferred files. • Using adapted pre-existent visualizations, this project allows to visualize at same place and with efficiency, results of an heterogeneous matrix bi-clustering using associated metadata from different reference ontologies. • Use of alluvial diagrams to represent changes in networks allowing evolution display of a matrix composed by genetic and experimental metadata in an interpretative way with statistical information. • This project allows to cross-analyze different types of information in the same place with a visual and efficient presentation. Conclusions [1] Ashburner M, Ball CA, Blake JA, et al. Gene ontology: tool for the unification of biology.
L'emergence de technologies d'imagerie (acquisition et analyse d'images) en biologie a permis un changement d'echelle dans le nombre d'individus pouvant etre observes conjointement et dans la frequence a laquelle les mesures peuvent etre realisees. Par exemple, dans les etudes de germination de semences, la ou un nombre reduit de lots, contenant un nombre reduit de graines, pouvaient etre evalues, a des intervalles de temps assez longs, ce sont des centaines de lots incluant des centaines de graines pour lesquels des descripteurs peuvent etre extraits a un rythme qui se rapproche de plus en plus du temps reel. Pour le biologiste, ce changement d'echelle pose un veritable challenge pour capturer et decrire finement les comportements des lots de semences. Il represente egalement une veritable opportunite non encore exploitee de prendre en compte l'heterogeneite des individus constituant les lots. Prealablement a l'analyse de l'effet d'un ou plusieurs facteurs (genotype, traitement, etc.) sur les differentes variables quantitatives acquises lors des experiences, les analyses de correlation et analyses multivariees telles que l'analyse en composantes principales (ACP) sont desormais employees en routine par les biologistes. Dans le contexte d'acquisitions a grande echelle structurees en lots, une pratique usuelle consiste a representer chaque lot par une valeur calculee unique pour chacune des variables mesurees (moyenne, mediane, etc.). Cette approche permet de simplifier l'analyse et ameliorer la lisibilite de la visualisation en reduisant le nombre de points representes. L'heterogeneite au sein d'un lot pourrait toutefois etre potentiellement caracteristique du comportement de certains systemes biologiques, telles que les semences, et etre liee a d'autres facteurs, identifies ou non lors de la realisation de l'analyse. Il est donc important de voir a quel point il pourrait etre pertinent de prendre en compte cette heterogeneite au niveau de l'ACP. Afin d'evaluer les approches alternatives de representation des lots, nous avons choisi de comparer les resultats obtenus avec 3 methodes de type ACP sur un jeu de donnees reelles acquises au cours du projet ANR REGULEG.
The demand by biologists to integrate heterogeneous and large datasets from omics and phenotyping activites is rapidly increasing. However, methods automating this approach are still at its infancy and to our knowledge, no operational and user-friendly software yet exists. Experiments are performed independently and resulting data are cross-analysed manually and a-posteriori by scientists. For instance, the biology teams from the IRHS (Institut de Recherche en Horticulture et Semences) in Angers have been accumulating datasets of different natures (transcriptomic, biochemistry, physical measures, sensory analysis, etc.) regarding perennial, annual and biannual plants. These datasets are described using reference ontolo-gies enriched with in-house knowledge and stored in a Laboratory Information Management System (LIMS) which is developed and distributed by the IRHS Bioinformatics team. The main objective of the DIVIS (Data Integration and VISualization) project is to develop a directly usable prototype of a new data analysis tool, by combining the most promising integration and visualisation approaches that are publicly available using the heterogeneous large scale datasets stored in our LIMS.
The management of biological resources collection is a major key for quality of research results. It is also mandatory to keep trace from data to studied organism through every experimental steps including sampling, culture conditions and description of the subject of the study. To achieve this goal in the context of plant science, there are some notable software such as Doriane Labkey and GreenGlobal. After evaluating these solutions with other laboratories we ended with the conclusion that, for different reasons, neither fulfill our needs especially when dealing with perennial plants such as fruit trees (apple or pear trees) or ornamental bush (rose). We then decided to provide our own solution based on the needs and feedbacks from the different teams of the IRHS in the fields of biological resources management, breeding, genetics, molecular biology, physiology and phenotyping. We introduce here two pieces of software:-ELVIS: an information system that takes care of data management and provides an extensible set of python libraries to interact with. These libraries are used here to provide a JSON-RPC API exposed as web services. This is the foundation of the LIMS of our laboratory.-PREMS: a dynamic web interface for biological resources management. This is one of the interfaces developed to interact with ELVIS. ELVIS The laboratory information system Main features:-3 levels of storage: Variety, Accession, Lot (as plant or seeds)-Notations on each level-Grouping of lot for actions (notation, transfer)-Multi-criteria search-Usergroup level data partitioning-Single value or batch importation for material introduction-Single value or batch notations with the ability of exporting field notation forms that can be uploaded back in the database Technology: Qooxdoo (Javascrip framwork) Bioinformatic team HTTP(S) Postgresql JSON-RPC Web API Pyhon libraries ELVIS Qooxdoo libraries PREMS PREMS The plant resource management system Multi-criteria Variety search gives a list of Varieties. A selection in that list displays a detailed view of the Lots and Accessions for that Variety. This view provides quick access to denomination and passport data. On Lots, the view gives access to phenotyping notations and location of the plants or the seeds bag. The user can also enter new notations. ELVIS and PREMS are made available to the community through the SourceSup forge under CeCILL license. ELVIS can provide access to the data by implementing compatible API to interact with other tools. For instance, we are deploying a PHIS installation with the integration of IRHS team in the PHENOM project and are working on ways to make the two applications interact. https://sourcesup.renater.fr/projects/elvis/ https://sourcesup.renater.fr/projects/prems/
Background: Image analysis is increasingly used in plant phenotyping. Among the various imaging techniques that can be used in plant phenotyping, chlorophyll fluorescence imaging allows imaging of the impact of biotic or abiotic stresses on leaves. Numerous chlorophyll fluorescence parameters may be measured or calculated, but only a few can produce a contrast in a given condition. Therefore, automated procedures that help screening chlorophyll fluorescence image datasets are needed, especially in the perspective of high-throughput plant phenotyping.Results: We developed an automatic procedure aiming at facilitating the identification of chlorophyll fluorescence parameters impacted on leaves by a stress. First, for each chlorophyll fluorescence parameter, the procedure provides an overview of the data by automatically creating contact sheets of images and/or histograms. Such contact sheets enable a fast comparison of the impact on leaves of various treatments, or of the contrast dynamics during the experiments. Second, based on the global intensity of each chlorophyll fluorescence parameter, the procedure automatically produces radial plots and box plots allowing the user to identify chlorophyll fluorescence parameters that discriminate between treatments. Moreover, basic statistical analysis is automatically generated. Third, for each chlorophyll fluorescence parameter the procedure automatically performs a clustering analysis based on the histograms. This analysis clusters images of plants according to their health status. We applied this procedure to monitor the impact of the inoculation of the root parasitic plant Phelipanche ramosa on Arabidopsis thaliana ecotypes Col-0 and Ler.Conclusions: Using this automatic procedure, we identified eight chlorophyll fluorescence parameters discriminating between the two ecotypes of A. thaliana, and five impacted by the infection of Arabidopsis thaliana by P. ramosa. More generally, this procedure may help to identify chlorophyll fluorescence parameters impacted by various types of stresses. We implemented this procedure at http://www.phenoplant.org freely accessible to users of the plant phenotyping community.
Un des problemes actuels en bioinformatique est de comprendre les mecanismes de regulation au sein d'une cellule. Notre travail concerne l'etude des reseaux de genes, avec la particularite d'y integrer les acteurs encore mal connus que sont les ARN anti-sens. Plusieurs etudes pointent l'importance de la regulation par les anti-sens dans les phenomenes de reponses aux stress. Pour etudier l'impact des anti-sens dans les reseaux de genes, nous etudions leur comportement dans deux contextes experimentaux d'un processus biologique de stress.
Le projet Paytal (Paysages et étalement urbain) est un travail pluridisciplinaire qui propose une méthodologie reproductible d’analyse des liens entre paysages et urbanisation. Ces liens sont ambivalents puisque en habitant des paysages qu’elles valorisent, les populations les modifient. Les politiques publiques visant à améliorer les paysages et le cadre de vie auront donc des effets pervers si elles ne contrôlent pas également le développement urbain. Pour mesurer et caractériser ces liens, nous développons une méthodologie en trois étapes, appliquée à l’aire urbaine d’Angers. D’abord, nous caractérisons finement l’urbanisation par télédétection. Ensuite, pour prendre en compte les aspects sensibles des paysages, nous générons une information spatialisée sur la perception des paysages à partir d’une ontologie. Celle-ci est basée sur l’information contenue dans les atlas des paysages tout en pouvant être généralisée à d’autres sources d’information. Enfin, nous expliquons l’urbanisation par un modèle de conversion des terres tenant compte des grands déterminants de l’urbanisation (proximité aux centres d’emploi, etc.), des caractéristiques physiques des paysages et de leurs dimensions sensibles. Nous montrons que ces dimensions jouent un rôle dans les phénomènes de périurbanisation.
Le projet PAYTAL (Paysages et Etalement Urbain) est un travail pluridisciplinaire visant a proposer une methodologie reproductible d’analyse des liens entre paysages et urbanisation. Ces liens sont importants du fait que les populations valorisent particulierement certains paysages et sont pretes a payer pour le droit d’habiter proches de ces paysages. En habitant ces paysages, elles les modifient. Les politiques publiques qui visent a ameliorer les paysages et le cadre de vie, auront donc des effets pervers si elles ne controlent pas egalement le developpement urbain. Pour mesurer et caracteriser les liens entre urbanisation et paysages, nous procedons en trois etapes, que nous appliquons a l’aire urbaine d’Angers. D’abord, nous caracterisons finement l’urbanisation par teledetection. Ensuite, nous reconnaissons pleinement les aspects sensibles des paysages. Nous en tenons compte en generant une information spatialisee sur la perception des paysages a partir d’une ontologie. Celle-ci est basee sur l’information contenue dans les Atlas des Paysages mais pourrait etre generalisee a d’autres sources d’information. Enfin, nous expliquons l’urbanisation par un modele de conversion des terres tenant compte des grands determinants de l’urbanisation (proximite aux centres d’emploi, etc.), des caracteristiques physiques des paysages et de leurs dimensions sensibles. Ce faisant, nous montrons que ces deux dimensions jouent un role dans les phenomenes de periurbanisation, au moins a Angers, et devraient donc etre prises en compte dans les politiques d’amenagement et de planification urbaine.
The quality of ornamental plants can be appraised with several types of criteria: tolerance to biotic and abiotic stresses, development potentialities and aesthetics. This last criterion, aesthetic quality, is specific to ornamental plants and objective measurements are required. Three methodologies for measuring aesthetic quality have been proposed. The first involves classical measurements of morphological features, such as flower number and diameter or leaf size. The second is based on sensory methods recently adapted to ornamental plants. The third, used by the International Union for the Protection of New Varieties of Plants (UPOV) for distinctness, uniformity and stability (DUS) tests, is based on morphological characteristics calibrated on specific reference varieties. The aim of this work was to compare these three methodologies for assessing some flowering and foliage characteristics of rosebushes. Six plants from 10 rose varieties identified by UPOV as reference varieties were cultivated for two years in a greenhouse and outdoors in Angers, France. They were measured and photographed weekly during flowering. Photographs of the plants in full bloom were submitted to a panel of judges for sensory assessment. The results of the three assessment methodologies were compared. Sensory and morphometric measurements were highly correlated and sensory measurements confirmed UPOV scales, whereas some morphometric measures diverged slightly from UPOV scales. We discuss the advantages, disadvantages and complementarity of these three methodologies.
Selon la Convention Europeenne du Paysage, le paysage est une partie de territoire telle que percue par les populations, dont le caractere resulte de l'action de facteurs naturels et/ou humains et de leurs interrelations . Nous nous interessons a la conception d'une ontologie axee perception par les populations en exploitant les aspects perceptifs disponibles dans les atlas de paysage. Pour palier les faibles frequences de ce vocabulaire, nous presentons une approche semi-automatique qui nous a permis de construire une premiere version de l'ontologie.
Un Atlas de paysages est un document destine a dresser l'etat des lieux des paysages et des dynamiques qui les transforment (Luginbuhl 1994). Il est realise a l'echelle du departement ou de la region, le plus souvent a l'initiative de la Direction Regionale de l'Environnement, et par une equipe generalement pluridisciplinaire. Afin que les Atlas de paysages constituent la formulation d'un etat de reference partage (Brunet-Vinck 2004), la Direction de l'Architecture et le l'Urbanisme a propose en 1994 un cadre methodologique explicitant la facon de structurer la maitrise d'ouvrage et donnant des etapes directrices claires pour analyser le paysage (Bourget 2011). Selon la methodologie proposee, la caracterisation des paysages ne doit pas se contenter d'une description de l'utilisation du sol, mais tenter de restituer les caracteres de l'ensemble de l'aspect du territoire et de ce que l'on en percoit , notamment par le biais d'une analyse de la dimension sensible du paysage. A ce titre l'atlas de Paysages est un objet hybride ou la subjectivite de l'auteur est clairement mise en avant (Davodeau 2009). Utiliser ce type d'objet pour la qualification des paysages pose donc question et suppose de s'appuyer sur une methodologie rigoureuse ou la part de la variabilite introduite par l'auteur puisse etre extraite. Les differents outils d'analyse textuelle (Tropes / ALCESTE / IRAMUTEQ) reposent sur une hypothese de stabilite interne du corpus (Rousseliere & Vezina 2009) derivant de celle de repetition (Reinert 2003). Ceci pose un probleme fondamental pour traiter de ce type de donnees (a l'instar de celles provenant d'entretiens semi-directifs, probleme bien souligne par Dalud-Vincent (2011)). Outre la methode traditionnelle en analyse de contenu de comparaison entre une classification experte et une classification automatique (Jenny 1997), nous avons mobilise deux methodes complementaires derivant des methodes d'evaluation econometrique des programmes (Imbens & Wooldridge 2009). Tout en renouant avec les objectifs initiaux de la statistique textuelle (voir Lebart & Salem 1994), notre proposition transposant cette litterature dans un cadre d'analyse de donnees textuelles est a notre connaissance originale. Une premiere dite d'experience naturelle s'appuie sur le fait que de nombreuses unites paysageres debordent les frontieres administratives departementales, alors que les atlas des paysages sont propres a un departement. En raison de cette exogeneite administrative, la difference de vocabulaire utilise peut etre alors raisonnablement attribuee a la personnalite des auteurs. Une deuxieme methode consiste a apparier sur des caracteristiques observables des territoires (type d'occupation du sol) les unites paysageres proches et de faire une typologie des vocabulaires utilises d'apres les caracteristiques des producteurs d'atlas des paysages. Pour ces deux methodes, differents indicateurs de subjectivite sont proposes.
This is a proposal developed within the community as an important first step in formalizing standards in reporting the production and properties of protein binding reagents, such as antibodies, developed and sold for the identification and detection of specific proteins present in biological samples. It defines a checklist of required information, intended for use by producers of affinity reagents, quality-control laboratories, users and databases. We envision that both commercial and freely available affinity reagents, as well as published studies using these reagents, could include a MIAPAR-compliant document describing the product's properties with every available binding partner.
Protein affinity reagents (PARs), most commonly antibodies, are essential reagents for protein characterization in basic research, biotechnology, and diagnostics as well as the fastest growing class of therapeutics. Large numbers of PARs are available commercially; however, their quality is often uncertain. In addition, currently available PARs cover only a fraction of the human proteome, and their cost is prohibitive for proteome scale applications. This situation has triggered several initiatives involving large scale generation and validation of antibodies, for example the Swedish Human Protein Atlas and the German Antibody Factory. Antibodies targeting specific subproteomes are being pursued by members of Human Proteome Organisation (plasma and liver proteome projects) and the United States National Cancer Institute (cancer-associated antigens). ProteomeBinders, a European consortium, aims to set up a resource of consistently quality-controlled protein-binding reagents for the whole human proteome. An ultimate PAR database resource would allow consumers to visit one on-line warehouse and find all available affinity reagents from different providers together with documentation that facilitates easy comparison of their cost and quality. However, in contrast to, for example, nucleotide databases among which data are synchronized between the major data providers, current PAR producers, quality control centers, and commercial companies all use incompatible formats, hindering data exchange. Here we propose Proteomics Standards Initiative (PSI)-PAR as a global community standard format for the representation and exchange of protein affinity reagent data. The PSI-PAR format is maintained by the Human Proteome Organisation PSI and was developed within the context of ProteomeBinders by building on a mature proteomics standard format, PSI-molecular interaction, which is a widely accepted and established community standard for molecular interaction data. Further information and documentation are available on the PSI-PAR web site.
New technologies and equipment allow for mass treatment of samples and research teams share acquired data on an always larger scale. In this context scientists are facing a major data exploitation problem. More precisely, using these data sets through data mining tools or introducing them in a classical experimental approach require a preliminary understanding of the information space, in order to direct the process. But acquiring this grasp on the data is a complex activity, which is seldom supported by current software tools. The goal of this paper is to introduce a solution to this scientific data grasp problem. Illustrated in the Tissue MicroArrays application domain, the proposal is based on the synthesis notion, which is inspired by Information Retrieval paradigms. The envisioned synthesis model gives a central role to the study the researcher wants to conduct, through the task notion. It allows for the implementation of a task-oriented Information Retrieval prototype system. Cases studies and user studies were used to validate this prototype system. It opens interesting prospects for the extension of the model or extensions towards other application domains.
David J. Sherman合作论文数Computer Science Department of the ENSEIRB,2