The MIRABEL project has developed a groundbreaking ICT system that fits the future liberalised energy sector and enables the integration of much higher rates of distributed and renewable energy sources (RES) into the electricity grid. This is done using a radical novel approach for active demand (and supply) side management in which electricity consumers and producers issue explicit so-called flex-offers indicating their available flexibilities in time and electricity amount. The MIRABEL system processes large amounts of flex-offers in order to balance electricity supply and demand in near real-time and thus supports the integration of non-schedulable renewable energy sources much better than earlier approaches.
Nowadays, Renewable Energy Sources (RES) are attracting more and more interest. Thus, many countries aim to increase the share of green energy and have to face with several challenges (e.g., balancing, storage, pricing). In this paper, we address the balancing challenge and present the MIRABEL project which aims to prototype an Energy Data Management System (EDMS) which takes benefit of flexibilities to efficiently balance energy demand and supply. The EDMS consists of millions of heterogeneous nodes that each incorporates advanced components (e.g., aggregation, forecasting, scheduling, negotiation). We describe each of these components and their interaction. Preliminary experimental results confirm the feasibility of our EDMS.
Background: Ontology term labels can be ambiguous and have multiple senses. While this is no problem for human annotators, it is a challenge to automated methods, which identify ontology terms in text. Classical approaches to word sense disambiguation use co-occurring words or terms. However, most treat ontologies as simple terminologies, without making use of the ontology structure or the semantic similarity between terms. Another useful source of information for disambiguation are metadata. Here, we systematically compare three approaches to word sense disambiguation, which use ontologies and metadata, respectively.Results: The 'Closest Sense' method assumes that the ontology defines multiple senses of the term. It computes the shortest path of co-occurring terms in the document to one of these senses. The 'Term Cooc' method defines a log-odds ratio for co-occurring terms including co-occurrences inferred from the ontology structure. The 'MetaData' approach trains a classifier on metadata. It does not require any ontology, but requires training data, which the other methods do not. To evaluate these approaches we defined a manually curated training corpus of 2600 documents for seven ambiguous terms from the Gene Ontology and MeSH. All approaches over all conditions achieve 80% success rate on average. The 'MetaData' approach performed best with 96%, when trained on high-quality data. Its performance deteriorates as quality of the training data decreases. The 'Term Cooc' approach performs better on Gene Ontology (92% success) than on MeSH (73% success) as MeSH is not a strict is-a/part-of, but rather a loose is-related-to hierarchy. The 'Closest Sense' approach achieves on average 80% success rate.Conclusion: Metadata is valuable for disambiguation, but requires high quality training data. Closest Sense requires no training, but a large, consistently modelled ontology, which are two opposing conditions. Term Cooc achieves greater 90% success given a consistently modelled ontology. Overall, the results show that well structured ontologies can play a very important role to improve disambiguation.Availability: The three benchmark datasets created for the purpose of disambiguation are available in Additional file 1.
EU Directive 86/609/EEC for the protection of laboratory animals obliges scientists to consider whether a planned animal experiment can be replaced, reduced or refined (3Rs principle). To meet this regulatory obligation, scientists must consult the relevant scientific literature prior to any experimental study using laboratory animals. More than 50 million potentially 3Rs relevant documents are spread over the World Wide Web, biomedical literature and patent databases. In April 2008, the beta version of Go3R ("www.Go3R.org":http://www.Go3R.org), the first knowledge-based semantic search engine for alternative methods to animal experiments, was released. Go3R is free of charge and enables scientists and regulatory authorities involved in the planning, authorisation and performance of animal experiments to determine the availability of alternative methods in a fast and comprehensive manner. The technical basis of this search engine is specific 3Rs expert knowledge captured within the Go3R Ontology containing 87,218 labels and synonyms. A total of 16,620 concepts were structured in 28 branches, where 1,227 concepts were newly defined to specifically describe directly 3Rs relevant knowledge. Additionally relevant headings from MeSH where referenced to reflect the topics associated with the definition of Animal Testing Alternatives. Therefore it is distinguished between thematic-defining and directly 3Rs relevant branches. In addition to the assignment of direct parent-child relationships, further relationship types were introduced to allow to model 3Rs relevant domain knowledge. Examples for such knowledge are e.g. (1) the characteristics of cell culture tests methods, which usually utilize “specific cell types” or “cell lines” and are associated with a specific “endpoint” and “endpoint detection method” or (2) named test methods like “PREDISAFE™”, which replaces an animal test namely the “eye irritation test” in rabbits and uses specific cells namely “SIRC Cells” or (3) the “Haemagglutinin-Neuraminidase Protein Assay”, which detects a protein of the “Newcastle disease virus”. Thereby, an article in which e.g. a specific 3Rs method is not explicitly mentioned could still be recognized as relevant for the specific topic searched for in an indirect manner, for example if it mentions specific cells, endpoints or endpoint detection methods, which are relevant for the respective application. The search engine Go3R with its novel ontology is already well recognized by the 3Rs community and will be further maintained and developed.
Consideration and incorporation of all available scientific information is an important part of the planning of any scientific project. As regards research with sentient animals, EU Directive 86/609/EEC for the protection of laboratory animals requires scientists to consider whether any planned animal experiment can be substituted by other scientifically satisfactory methods not entailing the use of animals or entailing less animals or less animal suffering, before performing the experiment. Thus, collection of relevant information is indispensable in order to meet this legal obligation. However, no standard procedures or services exist to provide convenient access to the information required to reliably determine whether it is possible to replace, reduce or refine a planned animal experiment in accordance with the 3Rs principle. The search engine Go3R, which is available free of charge under http://Go3R.org, runs up to become such a standard service. Go3R is the world-wide first search engine on alternative methods building on new semantic technologies that use an expert-knowledge based ontology to identify relevant documents. Due to Go3R's concept and design, the search engine can be used without lengthy instructions. It enables all those involved in the planning, authorisation and performance of animal experiments to determine the availability of non-animal methodologies in a fast, comprehensive and transparent manner. Thereby, Go3R strives to significantly contribute to the avoidance and replacement of animal experiments.
Searching relevant information on the web is a main occupation of researchers nowadays. Classical keyword-based search engines have limits. Inconsistent vocabulary used by authors is not handled. Relevant information spread over multiple documents can not be found. An overview over an entire document collection can not be given by the means of ranked lists. Question answering requiring semantic disambiguation of occurring terminology is not possible. Trends in the literature can not be followed if vocabulary is evolving over time. GoPubMed is a semantic search engine using the background knowledge of ontologies to index the biomedical literature. In this chapter we discuss how semantic search can contribute to overcome the limits of classical search paradigms.
With the ever increasing size of scientific literature, finding relevant documents and answering questions has become even more of a challenge. Recently, ontologies—hierarchical, controlled vocabularies—have been introduced to annotate genomic data. They can also improve the question and answering and the selection of relevant documents in the literature search. Search engines such as GoPubMed.org use ontological background knowledge to give an overview over large query results and to answer questions. We review the problems and solutions underlying these next-generation intelligent search engines and give examples of the power of this new search paradigm.
With the ever increasing size of scientific literature, finding relevant documents and answering questions has become even more of a challenge. Recently, ontologies — hierarchical, controlled vocabularies — have been introduced to annotate genomic data. They can also improve the question answering and the selection of relevant documents in the literature search. Search engines such as GoPubMed.org use ontological background knowledge to give an overview over large query results and to answer questions. Here we give an overview over GoPubMed. We show how it can answer questions using the GeneOntology and the Medical subject Headings as background knowledge. We also demonstrate that GoPubMed is general by applying it to the problem of associating genes, tissues, and developmental stages, as described in the Edinburgh Mouse Atlas. GoPubMed builds on background knowledge in the form of ontologies, which are given for the previous two applications. We describe a method to automatically generate the vocabulary for ontologies and compare our method to 3 other approaches in the context of a lipid metabolism ontology. The deliverable comprises three sections. GoPubMed is described in Section 1, MousePubMed in Section 2 and ontology generation in Section 3.
The biomedical literature can be seen as a large integrated, but unstructured data repository. Extracting facts from literature and making them accessible is approached from two directions: manual curation efforts develop ontologies and vocabularies to annotate gene products based on statements in papers. Text mining aims to automatically identify entities and their relationships in text using information retrieval and natural language processing techniques. Manual curation is highly accurate but time consuming, and does not scale with the ever increasing growth of literature. Text mining as a high-throughput computational technique scales well, but is error-prone due to the complexity of natural language. How can both be married to combine scalability and accuracy? Here, we review the state-of-the-art text mining approaches that are relevant to annotation and discuss available online services analysing biomedical literature by means of text mining techniques, which could also be utilised by annotation projects. We then examine how far text mining has already been utilised in existing annotation projects and conclude how these techniques could be tightly integrated into the manual annotation process through novel authoring systems to scale-up high-quality manual curation.
Background Currently there is a strong need for methods that help to obtain an accurate description of protein interfaces in order to be able to understand the principles that govern molecular recognition and protein function. Many of the recent efforts to computationally identify and characterize protein networks extract protein interaction information at atomic resolution from the PDB. However, they pay none or little attention to small protein ligands and solvent. They are key components and mediators of protein interactions and fundamental for a complete description of protein interfaces. Interactome profiling requires the development of computational tools to extract and analyze protein-protein, protein-ligand and detailed solvent interaction information from the PDB in an automatic and comparative fashion. Adding this information to the existing one on protein-protein interactions will allow us to better understand protein interaction networks and protein function. Description SCOWLP ( S tructural C haracterization O f W ater, L igands and P roteins) is a user-friendly and publicly accessible web-based relational database for detailed characterization and visualization of the PDB protein interfaces. The SCOWLP database includes proteins, peptidic-ligands and interface water molecules as descriptors of protein interfaces. It contains currently 74,907 protein interfaces and 2,093,976 residue-residue interactions formed by 60,664 structural units (protein domains and peptidic-ligands) and their interacting solvent. The SCOWLP web-server allows detailed structural analysis and comparisons of protein interfaces at atomic level by text query of PDB codes and/or by navigating a SCOP-based tree. It includes a visualization tool to interactively display the interfaces and label interacting residues and interface solvent by atomic physicochemical properties. SCOWLP is automatically updated with every SCOP release. Conclusion SCOWLP enriches substantially the description of protein interfaces by adding detailed interface information of peptidic-ligands and solvent to the existing protein-protein interaction databases. SCOWLP may be of interest to many structural bioinformaticians. It provides a platform for automatic global mapping of protein interfaces at atomic level, representing a useful tool for classification of protein interfaces, protein binding comparative studies, reconstruction of protein complexes and understanding protein networks. The web-server with the database and its additional summary tables used for our analysis are available at http://www.scowlp.org .
The adoption of agent technologies and multi-agent systems constitutes an emerging area in bioinformatics. In this article, we report on the activity of the Working Group on Agents in Bioinformatics (BIOAGENTS) founded during the first AgentLink III Technical Forum meeting on the 2nd of July, 2004, in Rome. The meeting provided an opportunity for seeding collaborations between the agent and bioinformatics communities to develop a different (agent-based) approach of computational frameworks both for data analysis and management in bioinformatics and for systems modelling and simulation in computational and systems biology. The collaborations gave rise to applications and integrated tools that we summarize and discuss in context of the state of the art in this area. We investigate on future challenges and argue that the field should still be explored from many perspectives ranging from bio-conceptual languages for agent-based simulation, to the definition of bio-ontology-based declarative languages to be used by information agents, and to the adoption of agents for computational grids.
The life sciences are a promising application area for semantic web technologies as there are large online structured and unstructured data repositories and ontologies, which structure this knowledge. We briefly give an overview over biomedical ontologies and show how they can help to locate, retrieve, and integrate biomedical data. Annotating literature with ontology terms is an important problem to support such ontology-based searches. We review the steps involved in this text mining task and introduce the ontology-based search engine GoPubMed. As the underlying data sources evolve, so do the ontologies. We give a brief overview over different approaches supporting the semi-automatic evolution of ontologies.
The biomedical literature grows at a tremendous rate and PubMed comprises already over 15 000 000 abstracts. Finding relevant literature is an important and difficult problem. We introduce GoPubMed, a web server which allows users to explore PubMed search results with the Gene Ontology (GO), a hierarchically structured vocabulary for molecular biology. GoPubMed provides the following benefits: first, it gives an overview of the literature abstracts by categorizing abstracts according to the GO and thus allowing users to quickly navigate through the abstracts by category. Second, it automatically shows general ontology terms related to the original query, which often do not even appear directly in the abstract. Third, it enables users to verify its classification because GO terms are highlighted in the abstracts and as each term is labelled with an accuracy percentage. Fourth, exploring PubMed abstracts with GoPubMed is useful as it shows definitions of GO terms without the need for further look up. GoPubMed is online at www.gopubmed.org. Querying is currently limited to 100 papers per query.
Ontologies, which are structured, hierarchical vocabularies, are widely used in molecular biology to annotate sequence and structure data. One such ontology, the GeneOntology, contains some 18000 terms on biological processes, molecular function, and cellular components. GeneOntology is available as flat file, in web formats such as XML and RDF, and as database. Using these formats we compare three different reasoners to query the GeneOntology. Prolog is the classical logic programming approach to reason over the ontology, Prova is a rule-based Java scripting language, and Xcerpt a query language for XML and RDF. We conclude by discussing the strengths and weaknesses of the three approaches.
This deliverable specifies use cases based on bioinformatics research carried out by members ofA2. The use cases involve the use of rules to reason over ontologies and pathways (Dresden,Edinburgh, ...
Bioinformatics is an important application area for semantic web technologies as much of the data is online and accessible in XML format, as some sites already support web services, and as ontologies are widely used to annotate data. In this deliverable, we give a survey over 18 of the most important bioinformatics resources and discuss their availability and accessibility, which are two of the main criteria for these resources to act as bases for later demonstrators.
He Tan合作论文数Dept. of Computer and Information Science
Linköpings universitet
3
Bogdan Filipič合作论文数Department of Intelligent Systems
Jozef Stefan Institute1