MOTIVATION:In the age of big data, the amount of scientific information available online dwarfs the ability of current tools to support researchers in locating and securing access to the necessary materials. Well-structured open data and the smart systems that make the appropriate use of it are invaluable and can help health researchers and professionals to find the appropriate information by, e.g., configuring the monitoring of information or refining a specific query on a disease.METHODS:We present an automated text classifier approach based on the MEDLINE/MeSH thesaurus, trained on the manual annotation of more than 26 million expert-annotated scientific abstracts. The classifier was developed tailor-fit to the public health and health research domain experts, in the light of their specific challenges and needs. We have applied the proposed methodology on three specific health domains: the Coronavirus, Mental Health and Diabetes, considering the pertinence of the first, and the known relations with the other two health topics.RESULTS:A classifier is trained on the MEDLINE dataset that can automatically annotate text, such as scientific articles, news articles or medical reports with relevant concepts from the MeSH thesaurus.CONCLUSIONS:The proposed text classifier shows promising results in the evaluation of health-related news. The application of the developed classifier enables the exploration of news and extraction of health-related insights, based on the MeSH thesaurus, through a similar workflow as in the usage of PubMed, with which most health researchers are familiar.
With the rapid spread of the COVID-19 pandemic, the novel Meaningful Integration of Data Analytics and Services (MIDAS) platform quickly demonstrates its value, relevance and transferability to this new global crisis. The MIDAS platform enables the connection of a large number of isolated heterogeneous data sources, and combines rich datasets including open and social data, ingesting and preparing these for the application of analytics, monitoring and research tools. These platforms will assist public health author ities in: (i) better understanding the disease and its impact; (ii) monitoring the different aspects of the evolution of the pandemic across a diverse range of groups; (iii) contributing to improved resilience against the impacts of this global crisis; and (iv) enhancing preparedness for future public health emergencies. The model of governance and ethical review, incorporated and defined within MIDAS, also addresses the complex privacy and ethical issues that the developing pandemic has highlighted, allowing oversight and scrutiny of more and richer data sources by users of the system.
Amazing things have been achieved in a wide range of application domains by exploiting a multitude of small connected devices, defined as the Internet of Things. Managing of these devices and their resources is a task for the underlying Fog technology that enables building of smart and efficient applications. Currently, the Fog is not implemented to the extent that we can submit application requirements to a Fog provider, select returned resources and deploy an application on them. A widely adopted workaround is to deploy Cloud applications that exploit the functionality of IoT and Fog devices. Although Clouds provide virtually unlimited computation power, they could present a bottleneck and unnecessary communication overhead when a huge number of devices needs to be controlled, read or written to. Therefore, it is reasonable to formulate use cases that will exploit the Edge and Fog functionality and define a set of basic requirements for Fog providers.
There is an ever increasing number of data sources that potentially could be used to gain new insights into areas such as disease prevention, policy formulation/ evaluation and personalised medicine, but these are not optimised for use within an analytics type user interface. The MIDAS project was funded under a call for ‘Big Data supporting Public Health policies’ to develop a big data platform that facilitates the utilisation of healthcare data beyond existing isolated systems, making that data amenable to enrichment with open and social data [1]. This aligns closely with a number of themes in Knowledge Discovery in Databases (KDD) in that the platform enables the integration of heterogeneous data sources, providing privacy-preserving analytics, forecasting tools and visualisation modules to deliver actionable information. Policy makers as a result will have the capability to perform data-driven evaluations of the efficiency and effectiveness of proposed policies in terms of expenditure, delivery, wellbeing, and health and socio-economic inequalities, thus improving current policy formulation, delivery risk stratification and evaluation. This H2020 project has a total of 15 partners from 5 EU countries as well as Arizona State University (ASU). The partners are Universities, SMEs and health departments in governmental institutions.
We present 3XL News, a multi-lingual news aggregation application for iPad that provides real-time, comprehensive, global and multilingual news coverage. Using methods, developed within the XLike project, for semantic data extraction from news articles and linking of news stories we are able to construct a concise, yet in-depth view of current news stories and their semantic relation. This enables users real-time monitoring of current global events and analysis of diverse reporting in different languages and navigation across related news stories.
For most events of at least moderate significance, there are likely tens, often hundreds or thousands of online articles reporting on it, each from a slightly different perspective. If we want to understand an event in depth, from multiple perspectives, we need to aggregate multiple sources and understand the relations between them. However, current news aggregators do not offer this kind of functionality. As a step toward a solution, we propose DiversiNews, a real‐time news aggregation and exploration platfom whose main feature is a novel set of controls that allow users to contrast reports of a selected event based on topical emphases, sentiment differences, and/or publisher geolocation. News events are presented in the form of a ranked list of articles pertaining to the event and an automatically generated summary. Both the ranking and the summary are interactive and respond in real time to user's change of controls. We validated the concept and the user interface through user tests with positive results.
This paper presents a system that uses semantic data to improve cross-lingual linking of news article clusters. Two approaches are compared. The first based on two different Canonical Correlation Analysis (CCA) feature vector definitions: MAX-CCA and SUM-CCA, whereas the second one has been developed using a better-performed CCA approach in combination with Entity vectors. The aim of the comparison was to determine whether taking into account the semantic aspect of news increases performance and improves linking. Evaluations of the aforementioned techniques on a news corpus, both against Google News and manual, revealed good performance of our system. The overall gain in precision and recall when using entity vectors was significant.
In May 2011, an outbreak of enterohemorrhagic Escherichia coli (EHEC) occurred in northern Germany. The Shiga toxin-producing strain O104:H4 infected several thousand people, frequently leading to haemolytic uremic syndrome (HUS) and gastroenteritis (GI). First reports about the outbreak appeared in the German media on Saturday 21st of May 2011; the media attention rose to high levels in the following two weeks, with up to 2000 articles categorized per day by the automatic threat detection system MedISys (Medical Information System). In this article, we illustrate how MedISys detected the sudden increase in reporting on E. coli on 21st of May and how automatic analysis of the reporting provided epidemic intelligence information to follow the event. Categorization, filtering and clustering allowed identifying different aspects within the unfolding news event, analyzing general media and official sites in parallel.