Many would argue that the currency of research is citations; however, researchers and funding organizations alike are lacking tools with which they can explore how this currency translates to funding opportunities. Motivated by this need, in this paper we address one of the fundamental problems facing the development of such a tool, namely the problem of automatically extracting funding information from scientific articles. For this purpose, we experiment with a two-stage framework which ingests text, filters paragraphs which contain funding information, and then combines sequential learning methods to detect named entities in a novel ensemble approach. We present a comparative analysis of each independent component of this pipeline, named FundingFinder, the results of which indicate that the said pipeline can extract the funding organizations and the associated grants, from scientific articles, accurately and efficiently.
In commercial research and development projects, public disclosure of new chemical compounds often takes place in patents. Only a small proportion of these compounds are published in journals, usually a few years after the patent. Patent authorities make available the patents but do not provide systematic continuous chemical annotations. Content databases such as Elsevier's Reaxys provide such services mostly based on manual excerptions, which are time-consuming and costly. Automatic text-mining approaches help overcome some of the limitations of the manual process. Different text-mining approaches exist to extract chemical entities from patents. The majority of them have been developed using sub-sections of patent documents and focus on mentions of compounds. Less attention has been given to relevancy of a compound in a patent. Relevancy of a compound to a patent is based on the patent's context. A relevant compound plays a major role within a patent. Identification of relevant compounds reduces the size of the extracted data and improves the usefulness of patent resources (e.g. supports identifying the main compounds). Annotators of databases like Reaxys only annotate relevant compounds. In this study, we design an automated system that extracts chemical entities from patents and classifies their relevance. The gold-standard set contained 18 789 chemical entity annotations. Of these, 10% were relevant compounds, 88% were irrelevant and 2% were equivocal. Our compound recognition system was based on proprietary tools. The performance (F-score) of the system on compound recognition was 84% on the development set and 86% on the test set. The relevancy classification system had an F-score of 86% on the development set and 82% on the test set. Our system can extract chemical compounds from patents and classify their relevance with high performance. This enables the extension of the Reaxys database by means of automation.
Chemical patents are an important resource for chemical information. However, few chemical Named Entity Recognition (NER) systems have been evaluated on patent documents, due in part to their structural and linguistic complexity. In this paper, we explore the NER performance of a BiLSTM-CRF model utilising pre-trained word embeddings, character-level word representations and contextualized ELMo word representations for chemical patents. We compare word embeddings pre-trained on biomedical and chemical patent corpora. The effect of tokenizers optimized for the chemical domain on NER performance in chemical patents is also explored. The results on two patent corpora show that contextualized word representations generated from ELMo substantially improve chemical NER performance w.r.t. the current state-of-the-art. We also show that domain-specific resources such as word embeddings trained on chemical patents and chemical-specific tokenizers have a positive impact on NER performance.
This work is an exploratory study of how we could progress a step towards an AI assisted peer- review system. The proposed approach is an ambitious attempt to automate the Desk-Rejection phenomenon prevalent in academic peer review. In this investigation we first attempt to decipher the possible reasons of rejection of a scientific manuscript from the editors desk. To seek a solution to those causes, we combine a flair of information extraction techniques, clustering, citation analysis to finally formulate a supervised solution to the identified problems. The projected approach integrates two important aspects of rejection: i) a paper being rejected because of out of scope and ii) a paper rejected due to poor quality. We extract several features to quantify the quality of a paper and the degree of in-scope exploring keyword search, citation analysis, reputations of authors and affiliations, similarity with respect to accepted papers. The features are then fed to standard machine learning based classifiers to develop an automated system. On a decent set of test data our generic approach yields promising results across 3 different journals. The study inherently exhibits the possibility of a redefined interest of the research community on the study of rejected papers and inculcates a drive towards an automated peer review system.
We present a first version of a system for selecting chemical publications for inclusion in a chemistry information database. This database, Reaxys ( https://www.elsevier.com/solutions/reaxys ), is a portal for the retrieval of structured chemistry information from published journals and patents. There are three challenges in this task: (i) Training and input data are highly imbalanced; (ii) High recall ( ≥95% ) is desired; and (iii) Data offered for selection is numerically massive but at the same time, incomplete. Our system successfully handles the imbalance with the undersampling technique and achieves relatively high recall using chemical named entities as features. Experiments on a real-world data set consisting of 15,822 documents show that the features of chemical named entities boost recall by 8% over the usual n-gram features being widely used in general document classification applications. For fostering research on this challenging topic, a part of the data set compiled in this paper can be requested.
In this paper we present a solution for tagging funding bodies and grants in scientific articles using a combination of trained sequential learning models, namely conditional random fields (CRF), hidden markov models (HMM) and maximum entropy models (MaxEnt), on a benchmark set created in-house. We apply the trained models to address the BioASQ challenge 5c, which is a newly introduced task that aims to solve the problem of funding information extraction from scientific articles. Results in the dry-run data set of BioASQ task 5c show that the suggested approach can achieve a micro-recall of more than 85% in tagging both funding bodies and grants.
SKOS and OWL are quite different but complimentary languages. SKOS is targeted at "cognitive" or "navigational" representations, that is, thesauri, controlled vocabularies, and the like. OWL is targeted at logical representations of conceptual knowledge. To a first approximation, SKOS vocabularies try to capture useful relations between concepts, whereas OWL ontologies aim to capture true relations between concepts. Now, of course, the true is sometimes useful and the useful often true, thus SKOS and OWL overlap to some degree. However, there are applications where we need to know true relations (e.g., generating multiple choice questions). Furthermore, SKOS relations are not precisely specified (by design). For example, many different ways of being useful can be covered by the same SKOS relation, but only one way of being useful is actually applicable to some application.In this paper, we present a case study of modifying a large, existing SKOS vocabulary partially into OWL. This lifting is motivated by an application (generating multiple choice questions) that requires more precision in the representation than SKOS alone supports.
Modelling a science domain for the purposes of thematically categorizing the research work and enabling better browsing and search can be a daunting task, especially if a specialized taxonomy or ontology does not exist for this domain. Elsevier, the largest academic publisher, faces this challenge often, for the needs of supporting the journals submission system, but also for supplying ScienceDirect and Scopus, two flagship platforms of the company, with sufficient metadata, such as conceptual labels that characterize the research works, which can improve the user experience in browsing and searching the literature. In this paper we describe an Elsevier in-use case study of learning appropriate domain labels from a collection of 6, 357 full text articles in the neurology domain, exploring different document representations and clustering mechanisms. Besides the baseline approaches for document representation (e.g., bag-of-words) and their variations (e.g., n-grams), we employ a novel in-house methodology which produces conceptual fingerprints of the research articles, starting from a general domain taxonomy, such as the Medical Subject Headings (MeSH). A thorough empirical evaluation is presented, using a variety of clustering mechanisms and several validity indices to evaluate the resulting clusters. Our results summarize the best practices in modelling this specific domain and we report on the advantages and disadvantages of using the different clustering mechanisms and document representations that were examined, with the aim to learn appropriate conceptual labels for this domain.
Online communities, or groups, have largely been defined based on links, page rank, and eigenvalues. In this paper we explore identifying abstract groups, groups where member's interests and online footprints are similar but they are not necessarily connected to one another explicitly. We use a combination of structural information and content information from posts and their comments to build a footprint for groups. We find that these variables do a good job at identifying groups, placing members within a group, and help determine the appropriate granularity for group boundaries.
When an internet user clicks on a result in a search engine, a request is submitted to the destination web server that includes a referrer field containing the search terms given by the user. Using this information, website owners can analyze the search terms leading to their websites to better understand their visitors needs. This work explores some of the features that can be used for classification-based analysis of such referring search terms. We present initial results for the example task of classifying HTTP requests countries of origin. A system that can accurately predict the country of origin from query text may be a valuable complement to IP lookup methods which are susceptible to the obfuscation of dereferrers or proxies. We suggest that the addition of semantic features improves classifier performance in this example application. We begin by looking at related work and presenting our approach. After describing initial experiments and results, we discuss paths forward for this work.
Cognitive information processing at higher conceptual levels requires a computational approach to knowledge representation and analysis. Semantic network analysis bridges the gap between probabilistic pattern recognition techniques and symbolic representations by replacing cumbersome and computationally complex forms of logic-based semantic inference common in symbolic approaches with mathematical metrics on graph representations of labelled, directed semantic networked data. These metrics in turn support assessment of evidentiary support for the presence of patterns of interest in which entities play specified roles in complex event scenarios. The resulting system allows patterns to be specified at higher levels of conceptual abstraction while also remaining robust to conflicting and incomplete information.
We have developed methods to identify online communities, or groups, using a combination of structural information variables and content information variables from weblog posts and their comments to build a characteristic footprint for groups. We have worked with both explicitly connected groups and 'abstract' groups, in which the connection between individuals is in interest (as determined by content based features) and behavior (metadata based features) as opposed to explicit links. We find that these variables do a good job at identifying groups, placing members within a group, and helping determine the appropriate granularity for group boundaries. The group footprint can then be used to identify differences between the online groups. In the work described here we are interested in determining how an individual's online behavior is influenced by their membership in more than one group. For example, individuals belong to a certain culture; they may belong as well to a demographic group, and other 'chosen' groups such as churches or clubs. There is a plethora of evidence surrounding the culturally sensitive adoption, use, and behavior on the Internet. In this work we begin to investigate how culturally defined internet behaviors may influence behaviors of subgroups. We do this through amore » series of experiments in which we analyze the interaction between culturally defined behaviors and the behaviors of the subgroups. Our goal is to (a) identify if our features can capture cultural distinctions in internet use, and (b) determine what kinds of interaction there are between levels and types of groups.« less