Introduction: Patients with heart failure with reduced ejection fraction (HFrEF) are phenotypically heterogeneous, and functional status and prognosis associated with complex phenotypes are poorly ...
Background: Patients with heart failure (HF) are heterogeneous with multiple complex phenotypes across the ejection fraction (EF) spectrum. Phenotype-specific response to various treatments have not been well-described. Hypothesis: Clinical response to specific interventions will vary according to HF phenotype. Methods: Using latent class analysis, six cluster-based HF phenotypes across the EF spectrum were previously identified using patient data from 2130 patients enrolled in HF-ACTION (LVEF ≤ 35%) and 1767 patients enrolled from the Americas in TOPCAT (LVEF ≥ 45%) based on age, sex, race, CAD, BMI, hyperlipidemia, hypertension, diabetes mellitus, atrial fibrillation, COPD, anemia, and renal function, but not LVEF. Response to aerobic exercise training vs. usual care (HF-ACTION) and spironolactone vs. placebo (TOPCAT) were quantified by phenotype. The primary outcome was a composite of cardiovascular mortality (CVM) or HF hospitalization (HFH). Secondary outcomes included CVM, HFH, and all-cause mortality (ACM). Change in peak VO2 at 3 and 12 months were also analyzed in HF-ACTION. Results: Of the established phenotypes, the phenotype composed of elderly non-ischemic patients as well as the non-white/non-ischemic/hypertensive phenotype experienced improvement in combined CVM and HFH, ACM and exercise capacity (28% vs 38%, HR: 0.66 [0.46-0.94], 8% vs 16% HR: 0.49 [0.27-0.91], change in VO2: 1.06±3.01 vs 0.04±3.14, p<0.05). Elderly patients with non-ischemic HFrEF enrolled in HF-ACTION randomized to therapeutic exercise program demonstrated significantly improved exercise capacity compared to usual care (change in VO2: 1.45±2.82 vs -0.09±2.49 and 1.25±3.18 vs 0.66±3.64 respectively, p<0.05). Elderly non-ischemic patients treated with spironolactone in TOPCAT had a lower risk of the primary outcome CVM and HFH (20% vs 27%, HR: 0.67 [0.48-0.95]), driven mostly by reduced CVM (9% vs 17%, HR: 0.52 [0.32-0.84]). Conclusions: Response to varied treatments such as exercise training and spironolactone varies among complex HF phenotypes in both HFpEF and HFrEF. Additional investigation which further characterizes phenotype-specific treatments may help select specific interventions most likely to benefit specific phenotypes.
This paper presents the idea of applying an open source, web-based platform – Gonito.net – for hosting challenges for researchers in the field of natural language processing. Researchers are encouraged to compete in well-defined tasks by developing tools and running them on provided test data. The researcher who submits the best results becomes the winner of the challenge. Apart from the competition, Gonito.net also enables collaboration among researchers by means of source code sharing mechanisms. Gonito.net itself is fully open source, i.e. its source is available for download and compilation, as well as a running instance of the system, available at gonito.net. The key design feature of Gonito.net is using Git for managing solutions of problems submitted by competitors. This allows for research transparency and reproducibility.
There is currently a crisis in science related to highly publicized failures to reproduce large numbers of published studies. The current work proposes, by way of case studies, a methodology for moving the study of reproducibility in computational work to a full stage beyond that of earlier work. Specifically, it presents a case study in attempting to reproduce the reports of two R libraries for doing text mining of the PubMed/MEDLINE repository of scientific publications. The main findings are that a rational paradigm for reproduction of natural language processing papers can be established; the advertised functionality was difficult, but not impossible, to reproduce; and reproducibility studies can produce additional insights into the functioning of the published system. Additionally, the work on reproducibility lead to the production of novel user-centered documentation that has been accessed 260 times since its publication-an average of once a day per library.
This case study examines strategies used to leverage the library’s existing journal licenses to obtain a large collection of full-text journal articles in extensible markup language (XML) format; the right to text mine the collection; and the right to use the collection and the data mined from it for grant-funded research to develop biomedical natural language processing (BNLP) tools. Researchers attempted to obtain content directly from PubMed Central (PMC). This attempt failed due to limits on use of content in PMC. Next researchers and their library liaison attempted to obtain content from contacts in the technical divisions of the publishing industry. This resulted in an incomplete research data set. Then researchers, the library liaison, and the acquisitions librarian collaborated with the sales and technical staff of a major science, technology, engineering, and medical (STEM) publisher to successfully create a method for obtaining XML content as an extension of the library’s typical acquisition process for electronic resources. Our experience led us to realize that text mining rights of full-text articles in XML format should routinely be included in the negotiation of the library’s licenses.
BACKGROUND:Ontological concepts are useful for many different biomedical tasks. Concepts are difficult to recognize in text due to a disconnect between what is captured in an ontology and how the concepts are expressed in text. There are many recognizers for specific ontologies, but a general approach for concept recognition is an open problem.RESULTS:Three dictionary-based systems (MetaMap, NCBO Annotator, and ConceptMapper) are evaluated on eight biomedical ontologies in the Colorado Richly Annotated Full-Text (CRAFT) Corpus. Over 1,000 parameter combinations are examined, and best-performing parameters for each system-ontology pair are presented.CONCLUSIONS:Baselines for concept recognition by three systems on eight biomedical ontologies are established (F-measures range from 0.14-0.83). Out of the three systems we tested, ConceptMapper is generally the best-performing system; it produces the highest F-measure of seven out of eight ontologies. Default parameters are not ideal for most systems on most ontologies; by changing parameters F-measure can be increased by up to 0.4. Not only are best performing parameters presented, but suggestions for choosing the best parameters based on ontology characteristics are presented.
Each year in Canada and the United States, thousands of talented teenaged hockey players are faced with a life changing decision. They must make a choice between playing in the Canadian Hockey League (CHL) or going to college and playing NCAA DI (National College Athletic Association, Division I) hockey. This is such a big decision because each of these choices can help lead a player to a professional career, but in very different ways. The CHL is structured more like the NHL in the number of games they play and the day to day schedule of practice and games. College hockey plays fewer games and focuses more on the development of individuals both on the ice and in school. That being said, the purpose of this research was to determine which path is more effective at preparing young hockey players for a professional hockey career in the future. In order to answer this research question, I looked into the past rosters of college and CHL teams. I took a random sample of players from each path and looked into how far those players got in their hockey careers. This helped to show just how well the path they took prepared them for the future. Document Type Undergraduate Project Professor's Name Katharine Burakowski Subject Categories Sports Management This undergraduate project is available at Fisher Digital Publications: https://fisherpub.sjfc.edu/sport_undergrad/73 1 RUNNING HEAD: CHL OR NCAA HOCKEY Major Junior or NCAA Hockey: Which Path Should Be Chosen? Christopher Roeder St. John Fisher College 2 RUNNING HEAD: CHL OR NCAA HOCKEY Abstract Each year in Canada and the United States, thousands of talented teenaged hockey players are faced with a life changing decision. They must make a choice between playing in the Canadian Hockey League (CHL) or going to college and playing NCAA DI (National College Athletic Association, Division I) hockey. This is such a big decision because each of these choices can help lead a player to a professional career, but in very different ways. The CHL is structured more like the NHL in the number of games they play and the day to day schedule of practice and games. College hockey plays fewer games and focuses more on the development of individuals both on the ice and in school. That being said, the purpose of this research was to determine which path is more effective at preparing young hockey players for a professional hockey career in the future. In order to answer this research question, I looked into the past rosters of college and CHL teams. I took a random sample of players from each path and looked into how far those players got in their hockey careers. This helped to show just how well the path they took prepared them for the future. Introduction Sports can have a big role in the lives of people in many nations across the world. Different areas of the world have different sports that are most important to them. In the United States, there are four major sports. They are baseball, football, basketball, and hockey. In all but one of these sports, there is a basic path one must take to get to the professional levels. First, the athlete will play in high school, and then go to college, then get drafted and play in the “minor leagues”, and then finally get to the pros. In hockey, there is another path that is available to the professional hockey hopeful. It involves 3 RUNNING HEAD: CHL OR NCAA HOCKEY going to a major junior hockey league instead of college. Both of these paths have the ability to help take a young hockey player to the National Hockey League (NHL). And both paths will help prepare a young man for the rest of his life. But one is probably more effective at doing this than the other. Playing major junior hockey has the benefit of at least 68 games a season (depending on playoffs) and a similar every day schedule to the NHL (Kennedy, 2011). Going to college will give an individual the chance to get a degree, allows the very unique experience of being in college, and can let the athlete spend more time at the gym while waiting for the next game (Chong, 2011). The objective of this research was to determine if one path is better than the other. The criteria used to determine this was which path produces more players that make it to professional hockey after playing in that league. Determining this helped to answer the research question of: Which path (Major Junior Hockey or NCAA Division I Hockey) is more effective at preparing a young hockey player for a professional hockey career in the future? Personal development was also looked at in the literature review. However, for this research, personal development was not looked at when answering the research question. This issue of which path is better is important for several reasons. First, it is very important to the player who is trying to make a decision between the two. Hundreds of high school aged hockey players have to make this decision every year. And there are many people who support each path whole-heartedly. The 60 CHL teams are allowed to have 24 players on their roster. This means there are 1440 kids playing in the CHL. In 2011 there were a total of 1,568 Division I hockey players playing for the 59 schools (Podnieks, 2011). This means there are at the very least, there were 3,008 hockey players 4 RUNNING HEAD: CHL OR NCAA HOCKEY who have had to make this decision in the last four years (however this number is certainly higher because players frequently leave college or the CHL before their eligibility is up in order to follow other possible opportunities. Another reason this number is actually higher is that not everyone who goes to college is on the roster for games or could even be cut from the team). Each young man could very well be getting advice several people who support different paths. This can make it extremely difficult for someone to decide between the two. If research were done to find out which path was statistically better for the individual, it would make the decision a whole lot easier. Another reason this is important is that it could potentially help NHL scouts and coaches. It could help them with choosing players to draft or sign. This would be most helpful for picking boarder line players. Depending on which path seems to be better, NCAA or major junior hockey, a scout might pick a lower end player who took one path rather than the other. Literature Review There are many different opinions on the question of which path is better. Many people form the Untied States believe that going the college hockey route is better (Dilks, 2013). More people in Canada think that major junior hockey is the better route (Custance, 2011). Either way, there are hundreds of people giving their opinion on the question, but not much actual research into it. This is evident by the number of blogs and articles there are on the subject. A simple search on Google about the CHL and NCAA hockey brings up many thousands of results. And by looking through many of these results, it would appear that Canadian based articles favor major junior hockey while the opposite is true for American based articles (Bourne, 2012). Either way, this is what is 5 RUNNING HEAD: CHL OR NCAA HOCKEY known for sure. There are 59 Division I hockey teams and 60 major junior league teams (Chong, 2011). This is convenient because it makes the two paths that much more comparable. Major junior hockey is made up of three different leagues. They are the Western Hockey League (WHL), Ontario Hockey League (OHL), and the Quebec Major Junior Hockey League (QMJHL) (Fitzpatrick, 2012). These three leagues are brought together by an umbrella organization called the Canadian Hockey League (CHL). The CHL has a regular season schedule consisting of 68 games. The winner of each league goes on to play for the Memorial Cup, which is basically the major junior hockey championship. Players in this league are aged from 16 to 20 (21 year olds are allowed to play in special cases) and the vast majority are from Canada and the United States (Fitzpatrick, 2012). The 59 Division I hockey programs are split into five different conferences. These NCAA hockey teams are allowed to play a maximum of 34 regular season games (Fitzpatrick, 2012). The reason a player must choose between college and major junior is that the NCAA considers all players who play major junior hockey ineligible to play hockey in the NCAA (King, 2012). The reason behind this is not because these players are “paid”. Players get billeted with local families and receive a small weekly allowance from the team (Fitzpatrick, 2012). The main reason the NCAA considers CHL teams to be professional is that there are a few players in the league who are under NHL contracts. Sometimes, a NHL team will sign a player, and then end up sending him down to play in the CHL because they feel he is not yet ready to play in the NHL or American Hockey League (AHL) (Peters, 2012). The NCAA has a rule that basically says that any player who is given more than is necessary to play, is considered a professional. And any 6 RUNNING HEAD: CHL OR NCAA HOCKEY league that has professional players (those who are being paid by a professional team) competing in it is considered a professional league (King, 2012). Just because players can’t play college hockey after having played in the CHL, doesn’t mean some players don’t do just the opposite. The CHL welcomes players from the NCAA and in fact actively try to recruit them to switch paths. The CHL really has an advantage over the NCAA in recruiting players for several reasons. First of all, there are basically no rules when it comes to CHL recruiting. CHL teams can start recruiting players at any age, while college teams have to wait until the individual is of a certain year in school (Peters, 2012). Division I college hockey coaches are not allowed to initiate contact with prospective student athletes until June 15 of their sophomore year (10th grade) in high school (Peters, 2012). The
BACKGROUND:We introduce the linguistic annotation of a corpus of 97 full-text biomedical publications, known as the Colorado Richly Annotated Full Text (CRAFT) corpus. We further assess the performance of existing tools for performing sentence splitting, tokenization, syntactic parsing, and named entity recognition on this corpus.RESULTS:Many biomedical natural language processing systems demonstrated large differences between their previously published results and their performance on the CRAFT corpus when tested with the publicly available models or rule sets. Trainable systems differed widely with respect to their ability to build high-performing models based on this data.CONCLUSIONS:The finding that some systems were able to train high-performing models based on this corpus is additional evidence, beyond high inter-annotator agreement, that the quality of the CRAFT corpus is high. The overall poor performance of various systems indicates that considerable work needs to be done to enable natural language processing systems to work well when the input is full-text journal articles. The CRAFT corpus provides a valuable resource to the biomedical natural language processing community for evaluation and training of new models for biomedical full text publications.
We approached the problems of event detection, argument identification, and negation and speculation detection in the BioNLP'09 information extraction challenge through concept recognition and analysis. Our methodology involved using the OpenDMAP semantic parser with manually written rules. The original OpenDMAP system was updated for this challenge with a broad ontology defined for the events of interest, new linguistic patterns for those events, and specialized coordination handling. We achieved state-of-the-art precision for two of the three tasks, scoring the highest of 24 teams at precision of 71.81 on Task 1 and the highest of 6 teams at precision of 70.97 on Task 2. We provide a detailed analysis of the training data and show that a number of trigger words were ambiguous as to event type, even when their arguments are constrained by semantic class. The data is also shown to have a number of missing annotations. Analysis of a sampling of the comparatively small number of false positives returned by our system shows that major causes of this type of error were failing to recognize second themes in two-theme events, failing to recognize events when they were the arguments to other events, failure to recognize nontheme arguments, and sentence segmentation errors. We show that specifically handling coordination had a small but important impact on the overall performance of the system. The OpenDMAP system and the rule set are available at http://bionlp.sourceforge.net.
BACKGROUND:Bio-molecular event extraction from literature is recognized as an important task of bio text mining and, as such, many relevant systems have been developed and made available during the last decade. While such systems provide useful services individually, there is a need for a meta-service to enable comparison and ensemble of such services, offering optimal solutions for various purposes.RESULTS:We have integrated nine event extraction systems in the U-Compare framework, making them intercompatible and interoperable with other U-Compare components. The U-Compare event meta-service provides various meta-level features for comparison and ensemble of multiple event extraction systems. Experimental results show that the performance improvements achieved by the ensemble are significant.CONCLUSIONS:While individual event extraction systems themselves provide useful features for bio text mining, the U-Compare meta-service is expected to improve the accessibility to the individual systems, and to enable meta-level uses over multiple event extraction systems such as comparison and ensemble.
Summary: The UIMA framework and Web Services are emerging as useful tools for integrating biomedical text mining tools. This application note describes our work, which makes the NCBO Annotator available to UIMA workflows as a UIMA component. Availability: This wrapper is freely available on the web at http://bionlp-uima.sourceforge.net/ as part of the center’s UIMA tools distribution. It has been implemented in Java for support on Mac OS X, Linux and MS Windows Contact: chris.roeder@ucdenver.edu Integration and ease of installation are increasingly important concerns as the field of biomedical text mining tools grows in size and complexity. Issues include complex installation and integrating with other tools. Many tools are deployed as Web Services to avoid installation altogether. The NCBO’s Annotator (Jonquet) is one such tool. It integrates many ontologies into an annotation service available on the web. Incremental users don’t need to install it for themselves, just access the web service. UIMA (Ferruci) is an integration framework that makes combining disparate tools much easier. It provides a common user interface, common data representation and tool integration. The Center for Computational Pharmacology at the University of Colorado/SOM has adapted the NCBO Annotator to UIMA, making it available to UIMA projects. The NCBO Annotator “automatically processes a piece of raw text to annotate (or tag) it with relevant ontology concepts and return the annotations. “ (Jonquet) It makes use of much more than a single database or ontology and involves significant effort to integrate and maintain the data. Installing software and data locally would be cost prohibitive compared to remotely accessing an established instance. The annotator utilizes over 100 ontologies. They can be thought of as enriched term lists that *To whom correspondence should be addressed. include relationships and synonyms. One of the ontologies available is the Gene Ontology1 and can be used to find references to cell components, biological processes, and molecular function. The annotator finds terms in submitted text that is related to the concepts in the ontologies and returns annotations describing them. For example, if “mitochondrion” appeared in the submitted text, the start and end character indexes of the word, the GO Onotology id, “39917”, and an id for the concept in GO, “GO:0005739” would be returned. Such matches, direct matches, are found using the MGREP (Xuan) tool. The annotator makes use of the hierarchical nature of the ontologies as well as the UMLS2 Thesaurus to provide more functionality. It can climb the ontology’s hierarchy and report on more general concepts that relate to a particular word. “Intracellular membrane-enclosed organelle”, “intracellular organelle”, “organelle”, and “cell component” are the higher members of the hierarchy starting with “mitochondrion”. Such matches are called “is-a” matches. The UMLS also allows the annotator to navigate between ontologies and produce a broader range of results, including those from different forms of the word such as plurals (“mitochondria” would match “mitochondrion” for example) producing “mapping” matches. This functionality is available over the web to users as a web service. A web service is similar to a website, but written for the use of computer programs. In this case, it is accessible to both humans and computers through an available web page3 1 http://www.geneontology.org/ . For NLP projects that would make use of the functionality the annotator provides, the web service spares software developers the effort of procurement, installation, and maintenance of the code, data 2 http://umlsinfo.nlm.nih.gov/ 3 http://rest.bioontology.org/test_oba.html
We introduce a system developed for the BioCreativeII.5 community evaluation of information extraction of proteins and protein interactions. The paper focuses primarily on the gene normalization task of recognizing protein mentions in text and mapping them to the appropriate database identifiers based on contextual clues. We outline a "fuzzy" dictionary lookup approach to protein mention detection that matches regularized text to similarly regularized dictionary entries. We describe several different strategies for gene normalization that focus on species or organism mentions in the text, both globally throughout the document and locally in the immediate vicinity of a protein mention, and present the results of experimentation with a series of system variations that explore the effectiveness of the various normalization strategies, as well as the role of external knowledge sources. While our system was neither the best nor the worst performing system in the evaluation, the gene normalization strategies show promise and the system affords the opportunity to explore some of the variables affecting performance on the BCII.5 tasks.
BACKGROUND:An increase in work on the full text of journal articles and the growth of PubMedCentral have the opportunity to create a major paradigm shift in how biomedical text mining is done. However, until now there has been no comprehensive characterization of how the bodies of full text journal articles differ from the abstracts that until now have been the subject of most biomedical text mining research.RESULTS:We examined the structural and linguistic aspects of abstracts and bodies of full text articles, the performance of text mining tools on both, and the distribution of a variety of semantic classes of named entities between them. We found marked structural differences, with longer sentences in the article bodies and much heavier use of parenthesized material in the bodies than in the abstracts. We found content differences with respect to linguistic features. Three out of four of the linguistic features that we examined were statistically significantly differently distributed between the two genres. We also found content differences with respect to the distribution of semantic features. There were significantly different densities per thousand words for three out of four semantic classes, and clear differences in the extent to which they appeared in the two genres. With respect to the performance of text mining tools, we found that a mutation finder performed equally well in both genres, but that a wide variety of gene mention systems performed much worse on article bodies than they did on abstracts. POS tagging was also more accurate in abstracts than in article bodies.CONCLUSIONS:Aspects of structure and content differ markedly between article abstracts and article bodies. A number of these differences may pose problems as the text mining field moves more into the area of processing full-text articles. However, these differences also present a number of opportunities for the extraction of data types, particularly that found in parenthesized text, that is present in article bodies but not in article abstracts.
Summary: The Unstructured Information Management Architecture (UIMA) framework and web services are emerging as useful tools for integrating biomedical text mining tools. This note describes our work, which wraps the National Center for Biomedical Ontology (NCBO) Annotator—an ontology-based annotation service—to make it available as a component in UIMA workflows. Availability: This wrapper is freely available on the web at http://bionlp-uima.sourceforge.net/ as part of the UIMA tools distribution from the Center for Computational Pharmacology (CCP) at the University of Colorado School of Medicine. It has been implemented in Java for support on Mac OS X, Linux and MS Windows. Contact: chris.roeder@ucdenver.edu
This paper describes an effort to build a corpus of full-text journal articles in which every co-referring noun phrase is annotated. The identity and appositive relations were marked up. Several annotation schemas were evaluated and are described here; the OntoNotes guidelines were selected. Biomedical journal articles required a number of adaptations to the OntoNotes guidelines—mainly doing away with the notion of generics, which also had implications for the handling of nominal modifiers. Domain experts and linguists were evaluated with respect to their ability to function as annotators, and both were found to be effective. Progress is reported with about one third of the project done; inter-annotator agreement at this stage is 0.684 by the MUC metric.
Systems that locate mentions of concepts from ontologies in free text are known as ontology concept recognition systems. This paper describes an approach to the evaluation of the workings of ontology concept recognition systems through use of a structured test suite and presents a publicly available test suite for this purpose. It is built using the principles of descriptive linguistic field work and of software testing. More broadly, we also seek to investigate what general principles might inform the construction of such test suites. The test suite was found to be effective in identifying performance errors in an ontology concept recognition system. The system could not recognize 2.1% of all canonical forms and no non-canonical forms at all. Regarding the question of general principles of test suite construction, we compared this test suite to a named entity recognition test suite constructor. We found that they had twenty features in total and that seven were shared between the two models, suggesting that there is a core of feature types that may be applicable to test suite construction for any similar type of application.