With the advent of the COVID-19 pandemic, there was the hope that data science approaches could help discover means for understanding, mitigating, and treating the disease. This manifested itself in the creation of the COVID-19 Open Research Dataset (CORD-19) which aggregated COVID-19- related scientific literature for use by the data mining community. As a group of interdisciplinary informatics researchers at NIST, we embarked on an effort to use our experience and previously developed systems to explore whether we could enhance the CORD-19 data set and facilitate its use. This effort produced a prototype scientific informatics system that extended data curation, data repository, resource registry, term extraction and indexing systems and resulted in a repackaging of CORD-19 as a Python data package. This paper documents our efforts, provides lessons learned, and proposes a general architecture for these types of systems.
This short paper describes a web resource-the NIST CORD-19 Web Resource-for community explorations of the COVID-19 Open Research Dataset (CORD-19). The tools for exploration in the web resource make use of the NIST-developed Root- and Rule-based method, which exploits underlying linguistic structures to create terms that represent phrases in a corpus. The method allows for auto-suggesting-related terms to discover terms to refine the search of a COVID-19 heterogenous document base. The method also produces taxonomic structures in the target domain as well as providing semantic information about the relationships between terms. This term structure can serve as a basis for creating topic modeling and trend analysis tools. In this paper, we describe use of a novel search engine to demonstrate some of the capabilities above.
Root- and rule-based terms are structured representations of natural language phrases that can be automatically generated using a combination of statistical and symbolic methods. These terms are able to represent and normalize syntactic information about natural language phrases, making them richer than basic n-grams while greatly reducing the vocabulary size. In this paper, we discuss the use of root- and rule-based terms for information retrieval. We represent documents and queries as collections of root- and rule-based terms and show that this improves conventional information retrieval methods such as Latent Semantic Indexing and Latent Direchlet Allocation. Root- and rule-based terms improve on state of the art evaluation scores for the TREC 2016 clinical decision support track.
This short paper describes a web resource—the NIST CORD-19 Web Resource—for community explorations of the COVID-19 Open Research Dataset (CORD-19). The tools for exploration in the web resource make use of the NIST-developed Root- and Rule-based method, which exploits underlying linguistic structures to create terms that represent phrases in a corpus. The method allows for auto-suggesting-related terms to discover terms to refine the search of a COVID-19 heterogenous document base. The method also produces taxonomic structures in the target domain as well as providing semantic information about the relationships between terms. This term structure can serve as a basis for creating topic modeling and trend analysis tools. In this paper, we describe use of a novel search engine to demonstrate some of the capabilities above.
Literature search is usually done by a common vocabulary provided by authors, or at times by an ontology to generate information across categories. Web of Science also employs citation indexing to enhance search capabilities such as to search across disciplines. We have developed a new novel Natural Language Procedure that has a potential to automate the identification and indexing of vocabulary for use in a disciplinary community. In this method, vocabulary is created de novo from the documents. The length of the term created by this method could be single terms (nouns) representing a generic concept such as “Material” that would be related to large number of documents or terms that are longer such as “Dielectric material or “high-voltage dielectric material”, conjugated compound terms with qualifiers and qualified (root) terms, that are found in a smaller set of documents. Our method uses some of the novel concepts for creating compound words of desired semantics found in Indo-European languages, in particular Sanskrit. The method is called Root and Rule (R&R, DOI: 10.1007/s11837-015-1487-4). The created terms are specific to the usage in a particular domain and hence providing a means to let the user choose the term her is searching for without needing to develop search strategy as in general purpose search and indexing models . We have been testing our method in many disciplines including Material science, Biology, Cyber security, Engineering and Bibliographic data. One of our test case is the articles held by IUCr where we also use the IUCr common vocabulary as an add on to the R&R terms. Primary focus of my presentation is to get a community feedback on the vocabulary we created to search IUCr documents.
Motivated by the need for flexible, intuitive, reusable, and normalized terminology for guiding search and building ontologies, we present a general approach for generating sets of such terminologies from natural language documents. The terms that this approach generates are root- and rule-based terms, generated by a series of rules designed to be flexible, to evolve, and, perhaps most important, to protect against ambiguity and standardize semantically similar but syntactically distinct phrases to a normal form. This approach combines several linguistic and computational methods that can be automated with the help of training sets to quickly and consistently extract normalized terms. We discuss how this can be extended as natural language technologies improve and how the strategy applies to common use-cases such as search, document entry and archiving, and identifying, tracking, and predicting scientific and technological trends.
The finite element method (FEM) has very broad applications in a lot of research areas, and isogeometric analysis (IGA) is a new advancement based on FEM to integrate design with analysis. This chapter reviews the basic algorithm of finite element analysis (FEA) and its new developments, including IGA, extended FEM and immersed FEM. As a popular and powerful numerical method to solve partial differential equations over complex domains, FEM has been developed rapidly and used in many research areas including computational medicine, biology and engineering. The FEM is a general technique to solve boundary value problems with uniformly and non-uniformly spaced grids or meshes. In the implementation, the element stiffness matrix and element load vector are computed element by element, and then assembled together into the global stiffness matrix and global load vector. FEA has been applied …
Crystallographic studies of ligands bound to biological macromolecules (proteins and nucleic acids) represent an important source of information concerning drug-target interactions, providing atomic level insights into the physical chemistry of complex formation between macromolecules and ligands. Of the more than 115,000 entries extant in the Protein Data Bank (PDB) archive, similar to 75% include at least one non-polymeric ligand. Ligand geometrical and stereochemical quality, the suitability of ligand models for in silico drug discovery and design, and the goodness-of-fit of ligand models to electron-density maps vary widely across the archive. We describe the proceedings and conclusions from the first Worldwide PDB/Cambridge Crystallographic Data Center/Drug Design Data Resource (wwPDB/CCDC/D3R) Ligand Validation Workshop held at the Research Collaboratory for Structural Bioinformatics at Rutgers University on July 30-31, 2015. Experts in protein crystallography from academe and industry came together with non-profit and for-profit software providers for crystallography and with experts in computational chemistry and data archiving to discuss and make recommendations on best practices, as framed by a series of questions central to structural studies of macromolecule-ligand complexes. What data concerning bound ligands should be archived in the PDB? How should the ligands be best represented? How should structural models of macromolecule-ligand complexes be validated? What supplementary information should accompany publications of structural studies of biological macromolecules? Consensus recommendations on best practices developed in response to each of these questions are provided, together with some details regarding implementation. Important issues addressed but not resolved at the workshop are also enumerated.
Intuitive, flexible, and evolving terminology plays a significant role in capitalizing on recommended knowledge representation models for materials engineering applications. In this article, we present a proposed rules-based approach with initial examples from a growing corpus of materials terms in the National Institute of Standards and Technology (NIST) Materials Data Repository. Our method aims to establish a common, consistent, and evolving set of rules for creating or extending terminology as needed to describe materials data. The rules are intended to be simple and generalizable for users to understand and extend as well as for groups to apply to their own repositories. The rules generate terms that facilitate machine processing and decision making.
Background: There are significant challenges associated with the building of ontologies for cell biology experiments including the large numbers of terms and their synonyms. These challenges make it difficult to simultaneously query data from multiple experiments or ontologies. If vocabulary terms were consistently used and reused across and within ontologies, queries would be possible through shared terms. One approach to achieving this is to strictly control the terms used in ontologies in the form of a pre-defined schema, but this approach limits the individual researcher's ability to create new terms when needed to describe new experiments.Results: Here, we propose the use of a limited number of highly reusable common root terms, and rules for an experimentalist to locally expand terms by adding more specific terms under more general root terms to form specific new vocabulary hierarchies that can be used to build ontologies. We illustrate the application of the method to build vocabularies and a prototype database for cell images that uses a visual data-tree of terms to facilitate sophisticated queries based on a experimental parameters. We demonstrate how the terminology might be extended by adding new vocabulary terms into the hierarchy of terms in an evolving process. In this approach, image data and metadata are handled separately, so we also describe a robust file-naming scheme to unambiguously identify image and other files associated with each metadata value. The prototype database http://sbd.nist.gov/ consists of more than 2000 images of cells and benchmark materials, and 163 metadata terms that describe experimental details, including many details about cell culture and handling. Image files of interest can be retrieved, and their data can be compared, by choosing one or more relevant metadata values as search terms. Metadata values for any dataset can be compared with corresponding values of another dataset through logical operations.Conclusions: Organizing metadata for cell imaging experiments under a framework of rules that include highly reused root terms will facilitate the addition of new terms into a vocabulary hierarchy and encourage the reuse of terms. These vocabulary hierarchies can be converted into XML schema or RDF graphs for displaying and querying, but this is not necessary for using it to annotate cell images. Vocabulary data trees from multiple experiments or laboratories can be aligned at the root terms to facilitate query development. This approach of developing vocabularies is compatible with the major advances in database technology and could be used for building the Semantic Web.
Efficient and user friendly ontologies are crucial for the effective use of chemical structural data on compounds. This paper describes an automated technique to create a structural ontology for compounds like ligands, co-factors and inhibitors of protein and DNA molecules using a technique developed from Perl scripts, which use a relational database for input and output, called Chem-BLAST (Chemical Block Layered Alignment of Substructure Technique). This technique recursively identifies substructures using rules that operate on the atomic connectivity of compounds. Substructures obtained from the compounds are compared to generate a data model expressed as triples. A chemical ontology of the substructures is made up of numerous interconnected ‘hubs-and-spokes’ is generated in the form of a data tree. This data-tree is used in a Web interface to allow users to zoom into compounds of interest by stepping through the hubs from the top to the bottom of the data-tree. The technique has been applied for (a) 2-D and 3-D structural data for AIDS1; (b) ~60,000 structures from the PDB 2,3. Recently, this technique has been applied to approximately 3,000,000 compounds from PubChem4,5,6. Plausible ways to use this data model for the Semantic Web are also discussed.
The AIDS HIV structural databases (HIVSDB, http://bioinfo.nist.gov/SemanticWeb_pr2d/chemblast.do ), the Protein Data Bank (http://www.rcsb.org/pdb/home/home.do) and the PubChem ( http://pubchem.ncbi.nlm.nih.gov/ ) distribute one of the largest comprehensive collections of structural data on inhibitors, drug leads and clinical drugs for many diseases including AIDS. These databases contain info on several thousand biologically active compounds from many classes (HIV PR, RT, CCR5, Integrase) of FDA approved drugs for AIDS. Efficient and yet user friendly rule-based data management systems that support state-of-the-art annotation, visualization and query capabilities are crucial for the effective use of data for fragment based structural pharmacology and rational drug design. Semantic Web is the vision of the World Wide Web Consortium for enabling seamless integration of electronic data for data mining and knowledge generation across the Web. Robust and functionally relevant ontology plays a critical role in developing the data elements for a Semantic Web. Presentation will illustrate how Rule-based Semantic Web concepts are used for novel annotation, data integration, storage, and query to manage and display structural (fragments, 2-D images and text-based) biological, and pre-clinical data (http://xpdb.nist.gov/chemblast/pdb.html ). The technique is called (Chem-BLAST – Chemical Block Layered Alignment of Substructure Technique) and it allows rapid comparison and exchange of compound information using automated rulebased methods to develop structural ontology for use in drug discovery process. The methods and the results that will be presented are probably the first of its kind in biological world that uses entirely rule-based event processing methods to build, integrate and exchange information. The ontology developed by the methods is amenable to be presented either as an RDF or XML or to be used in a relational database. We call it a structural facebook as it is driven by ontology of structures of ‘shared features’ and presents them using visual images for rapid comparison.
The technique of co-crystallisation as a means to produce a solid dosage form of an active pharmaceutical ingredient (API) has been gaining popularity within the pharmaceutical industry.Co-crystals have the potential for greater flexibility and diversity in the created forms than salts or hydrates due to their compatibility with non-ionisable APIs and the large range of pharmaceutically acceptable coformers that are potentially available.The effective use of co-crystallisation for this purpose is clearly affected by the capacity with which the solid forms produced can be controlled and predicted.The structural knowledge stored within the Cambridge Structural Database (CSD, [1]) is highly relevant to all aspects of crystal engineering and the understanding of intermolecular interactions at a fundamental level.The study of appropriate structural data is of prime importance when investigating these areas.Recent software advances [2] have broadened the opportunities for using the CSD for design and control of crystal forms.This study focuses on the analysis of a large family of pharmaceutical co-crystals containing a consistent API component [3] to provide insight into intermolecular interactions in general as well as co-crystal design in particular.
Per-deuteration of proteins is becoming more common place and the assumption is in general that deuteration does not affect protein structure.It should be noted that functional changes upon deuteration are known, e.g.D2O is toxic to living systems, reaction kinetics change [1], proteins stiffen in D2O [2] and ferroelectrics alter their properties [3].There are two exceptions known to us for protein structures.Firstly in Kuhn et al [4] they observe a difference in position of a critical H versus D atom for subtilisin versus trypsin respectively.Secondly in haloalkane dehydrogenase Liu et al 2007 [5] observe a rotation of an Asp towards a His for the deuterated enzyme.In the absence of a statistical database of protein neutron AND ultra-high resolution X-ray crystal structures we have instead examined the effect of deuteration on structure by data-mining of the Cambridge Structural Database [6] for deuterated and hydrogenated pairs of small molecule structures which have been analysed by neutron or X-ray crystallography.There are, mainly, examples of isomorphous crystal pairs but also some non-isomorphous pairs.Differences between these structures for both types have been calculated and their statistical significance assessed.There are precisely enough measured structural differences but we find that they are in each case small enough that they do not upset the general assumption that deuteration does not affect protein structure.1.
Biology has been an experimental science until the recent prominence of Bioinformatics and Computational Biology. With the discovery of the DNA sequence the protein structure determination is now an emerging challenge. The protein structure is closely coupled to the function. Today, with the given increase in available computing power and no-cost storage, the ability to do computational experiments is emerging as a core competence necessary for rapid discovery in the future. The ability to include various complex physics such as electrostatics and hydrophobic interactions in realistic simulations has increased. The discovery of structure of proteins is the next frontier for a number of convergent areas in science. Hence the combination of geometry and physics becomes very critical to do realistic computational experiments.
There is considerable interest in RDF (Resource Description Framework) as a data representation standard for the growing information technology needs of drug discovery. Though several efforts towards this goal have been reported, most of the reported efforts have focused on text-based data. Structural data of chemicals are a key component of drug discovery and molecular images may offer certain advantages over text-based representations for them. Here we discuss the steps that we used to develop and search chemical Resource Description Framework (RDF) using text and image for structures of relevant to Acquired Immune Deficiency Syndrome (AIDS). These steps are (a) acquisition of the data on drugs, (b) definition of the framework to establish RDF on drugs using commonly asked questions during a drug discovery effort, (c) annotation of the structural data on drugs into RDF using the framework established in step (b), (d) validation of the annotation methods using Semantic Web concepts and tools, (e) design and development of public Web to distribute data to the public, (f) generation and distribution of data using OWL (Web Ontology Language). This paper describes this effort, discusses our observations and announces the availability of the OWL model at the W3C Web site (http://esw.w3.org/topic/HCLS/ChemicalTaxonomiesUseCase). The style of this paper is chosen so as to cover a broad audience including structural biologists, medicinal chemists, and information technologists and at times may appear say to the obvious for certain experts. A full discussion of our method and its comparison to other published methods is beyond the scope of this publication.
One of the difficulty in applying the concepts of semantic Web technology to a drug discovery is the creation of a software generated URI suitable both for ontology independent and ontology dependent properties of drugs. We propose the use of two types of URI; structure invariants (URI) to identify the ontology independent features and structure semi-invariants (OURI) to identify ontology dependent features. URI are defined by the most fundamental properties (chemical connectivity, bond type, etc.). OURI are defined from context dependent (semantic) properties (binding mode, chemical core).
Ira Monarch合作论文数5