Knowledge Base Population (KBP) aims to populate structured databases with facts extracted from text, encompassing tasks such as named entity recognition, coreference resolution, relation extraction, and entity linking. Traditional pipeline-based approaches sequentially chain modular components, leading to error propagation and unidirectional information flow. Additionally, black-box components often lack transparency and interpretability. In this paper, we propose a probabilistic pipeline framework for joint inference in end-to-end KBP. Our approach enables globally consistent decision-making by integrating local component feedback and external background knowledge. A key advantage is its ability to seamlessly incorporate knowledge about pipeline components, ontology constraints, linguistic patterns, and corpus characteristics. We evaluate our framework on two core KBP tasks: exhaustive relation extraction and entity linking.
Requirements Engineering (RE) is a critical yet time-consuming phase in software development, often hindered by inconsistencies and inaccuracies. We propose a novel framework that combines Foundation Models (FMs) and Multi-Agent Systems (MAS) to enhance RE efficiency, accuracy, and quality. Our framework aims to automate routine tasks and provide intelligent assistance throughout all RE phases. In a preliminary Proof-of-Concept (PoC) experiment, we explored the capabilities of Large Language Models (LLMs) in RE, evaluating their performance on various interlinked tasks across multiple phases. Our results showed that LLMs can achieve human-like accuracy in assessing requirements deliverables, highlighting the importance of suitable LLMs for each phase. We discuss the implications of our findings and outline a research agenda to address the limitations of FM-based MAS in RE, focusing on key objectives such as data availability, model calibration, and human-AI collaboration. Our goal is to create a more efficient, collaborative, and reliable RE process.
In recent years, data science and machine learning (ML) has become common across sectors and industries. Project methodologies are aimed at supporting projects and try catching up with ML trends and paradigm shifts. However, they are hardly successful, since still 80% of data science projects never reach deployment. The latest paradigm shift in the area of ML - the trend of generative AI and foundation models - changes the nature of data science projects and is not yet addressed by existing project methodologies. In this work, we present novel requirements that arise from real-world projects incorporating foundation models based on 29 case studies from the NLU domain. Furthermore, we assess existing data science methodologies and identify their shortcomings. Finally, we provide guidance on adapting projects to address the new challenges in the development and operation of foundation model based solutions.
Background: The healthcare sector is currently undergoing a significant transformation, driven by an increased utilization of data. In this evolving landscape, surveys are of pivotal importance to the comprehension of patient needs and preferences. Moreover, the digital affinity of patients and physicians within the healthcare system is reforming the manner in which healthcare services are accessed and delivered. The utilization and donation of data are influencing the future of medical research and treatment, while artificial intelligence (AI) is empowering patients and physicians with knowledge and improving healthcare delivery. Methods: In order to evaluate the opinions of patients and physicians regarding the management of personal health data and the functionality of upcoming data management devices in the context of healthcare digitization, we conducted an exploratory study and designed a survey. The survey focused on a number of key areas, including demographics, experience with digitization, data handling, the identification of needs for upcoming digitization, and AI in healthcare. Results: A total of 40 patients and 15 physicians participated in the survey. The results indicate that data security, timesaving/administrative support, and digital communication are aspects that patients associate with patient-friendly digitization. Based on the responses provided by physicians, it might be concluded that future digital platforms should prioritize usability, time efficacy, data security, and interoperability. Conclusions: In terms of expectations for future digital platforms, there is a notable overlap between the needs expressed by patients and those identified by physicians, particularly in relation to usability, time management, data security, and digital communication. This suggests that the requirements of different stakeholders can be combined in a future system, although individual issues may still require attention.
Achieving a good outcome for a person with Psoriatic Arthritis (PsA) is made difficult by late diagnosis, heterogenous clinical disease expression and in many cases, failure to adequately suppress inflammatory disease features. Single-centre studies have certainly contributed to our understanding of disease pathogenesis, but to adequately address the major areas of unmet need, multi-partner, collaborative research programmes are now required. HIPPOCRATES is a 5-year, Innovative Medicines Initiative (IMI) programme which includes 17 European academic centres experienced in PsA research, 5 pharmaceutical industry partners, 3 small-/medium-sized industry partners and 2 patient-representative organizations. In this review, the ambitious programme of work to be undertaken by HIPPOCRATES is outlined and common approaches and challenges are identified. It is expected that, when completed, the results will ultimately allow for changes in the approaches to diagnosing, managing and treating PsA allowing for better short-term and long-term outcomes.
The definitive diagnosis and early treatment of many immune-mediated inflammatory diseases (IMIDs) is hindered by variable and overlapping clinical manifestations. Psoriatic arthritis (PsA), which develops in ~30% of people with psoriasis, is a key example. This mixed-pattern IMID is apparent in entheseal and synovial musculoskeletal structures, but a definitive diagnosis often can only be made by clinical experts or when an extensive progressive disease state is apparent. As with other IMIDs, the detection of multimodal molecular biomarkers offers some hope for the early diagnosis of PsA and the initiation of effective management and treatment strategies. However, specific biomarkers are not yet available for PsA. The assessment of new markers by genomic and epigenomic profiling, or the analysis of blood and synovial fluid/tissue samples using proteomics, metabolomics and lipidomics, provides hope that complex molecular biomarker profiles could be developed to diagnose PsA. Importantly, the integration of these markers with high-throughput histology, imaging and standardized clinical assessment data provides an important opportunity to develop molecular profiles that could improve the diagnosis of PsA, predict its occurrence in cohorts of individuals with psoriasis, differentiate PsA from other IMIDs, and improve therapeutic responses. In this review, we consider the technologies that are currently deployed in the EU IMI2 project HIPPOCRATES to define biomarker profiles specific for PsA and discuss the advantages of combining multi-omics data to improve the outcome of PsA patients.
The COVID-19 pandemic and the high numbers of infected individuals pose major challenges for public health departments. To overcome these challenges, the health department in Cologne has developed a software called DiKoMa. This software offers the possibility to track contact and index persons, but also provides a digital symptom diary. In this work, the question of whether these can also be used for diagnostic purposes will be investigated. Machine learning makes it possible to identify infections based on early symptom profiles and to distinguish between the predominant dominant variants. Focusing on the occurrence of the symptoms in the first week, a decision tree is trained for the differentiation between contact and index persons and the prevailing dominant variants (Wildtype, Alpha, Delta, and Omicron). The model is evaluated, using sex- and age-stratified cross-validation and validated by symptom profiles of the first 6 days. The variants achieve an AUC-ROC from 0.89 for Omicron and 0.6 for Alpha. No significant differences are observed for the results of the validation set (Alpha 0.63 and Omicron 0.87). The evaluation of symptom combinations using artificial intelligence can determine the individual risk for the presence of a COVID-19 infection, allows assignment to virus variants, and can contribute to the management of epidemics and pandemics on a national and international level. It can help to reduce the number of specific tests in times of low labor capacity and could help to early identify new virus variants.
Zusammenfassung Hintergrund und Ziele Schon in der frühen Phase der global sehr verschieden verlaufenden COVID-19-Pandemie zeigten sich Hinweise auf den Einfluss sozioökonomischer Faktoren auf die Ausbreitungsdynamik der Erkrankung, die vor allem ab der zweiten Phase (September 2020) Menschen mit geringerem sozioökonomischen Status stärker betraf. Solche Effekte können sich auch innerhalb einer Großstadt zeigen. Die vorliegende Studie visualisiert und untersucht die zeitlich-räumliche Verbreitung aller in Köln gemeldeten COVID-19-Fälle (Februar 2020–Oktober 2021) auf Stadtteilebene und deren mögliche Assoziation mit sozioökonomischen Faktoren. Methoden Pseudonymisierte Daten aller in Köln gemeldeten COVID-19-Fälle wurden geocodiert, deren Verteilung altersstandardisiert auf Stadtteilebene über 4 Zeiträume kartiert und mit der Verteilung von sozialen Faktoren verglichen. Der mögliche Einfluss der ausgewählten Faktoren wird zudem in einer Regressionsanalyse in einem Modell mit Fallzuwachsraten betrachtet. Ergebnisse Das kleinräumige lokale Infektionsgeschehen ändert sich im Pandemieverlauf. Stadtteile mit schwächeren sozioökonomischen Indizes weisen über einen großen Teil des pandemischen Verlaufs höhere Inzidenzzahlen auf, wobei eine positive Korrelation zwischen den Armutsrisikofaktoren und der altersstandardisierten Inzidenz besteht. Die Stärke dieser Korrelation ändert sich im zeitlichen Verlauf. Schlussfolgerung Die zeitnahe Beobachtung und Analyse der lokalen Ausbreitungsdynamik lassen auch auf der Ebene einer Großstadt die positive Korrelation von nachteiligen sozioökonomischen Faktoren auf die Inzidenzrate von COVID-19 erkennen und können dazu beitragen, lokale Eindämmungsmaßnahmen zielgerecht zu steuern.
Methods from explainable machine learning are increasingly applied. However, evaluation of these methods is often anecdotal and not systematic. Prior work has identified properties of explanation quality and we argue that evaluation should be based on them. In this position paper, we provide an evaluation process that follows the idea of property testing. The process acknowledges the central role of the human, yet argues for a quantitative approach for the evaluation. We find that properties can be divided into two groups, one to ensure trustworthiness, the other to assess comprehensibility. Options for quantitative property tests are discussed. Future research should focus on the standardization of testing procedures.
Statistical models are inherently uncertain. Quantifying or at least upper-bounding their uncertainties is vital for safety-critical systems such as autonomous vehicles. While standard neural networks do not report this information, several approaches exist to integrate uncertainty estimates into them. Assessing the quality of these uncertainty estimates is not straightforward, as no direct ground truth labels are available. Instead, implicit statistical assessments are required. For regression, we propose to evaluate uncertainty realism -- a strict quality criterion -- with a Mahalanobis distance-based statistical test. An empirical evaluation reveals the need for uncertainty measures that are appropriate to upper-bound heavy-tailed empirical errors. Alongside, we transfer the variational U-Net classification architecture to standard supervised image-to-image tasks. We adopt it to the automotive domain and show that it significantly improves uncertainty realism compared to a plain encoder-decoder model.
The use of deep neural networks (DNNs) in safety-critical applications like mobile health and autonomous driving is challenging due to numerous model-inherent shortcomings. These shortcomings are diverse and range from a lack of generalization over insufficient interpretability to problems with malicious inputs. Cyber-physical systems employing DNNs are therefore likely to suffer from safety concerns. In recent years, a zoo of state-of-the-art techniques aiming to address these safety concerns has emerged. This work provides a structured and broad overview of them. We first identify categories of insufficiencies to then describe research activities aiming at their detection, quantification, or mitigation. Our paper addresses both machine learning experts and safety engineers: The former ones might profit from the broad range of machine learning topics covered and discussions on limitations of recent methods. The latter ones might gain insights into the specifics of modern ML methods. We moreover hope that our contribution fuels discussions on desiderata for ML systems and strategies on how to propel existing approaches accordingly.
The performance of modern relation extraction systems is to a great degree dependent on the size and quality of the underlying training corpus and in particular on the labels. Since generating these labels by human annotators is expensive, Distant Supervision has been proposed to automatically align entities in a knowledge base with a text corpus to generate annotations. However, this approach suffers from introducing noise, which negatively affects the performance of relation extraction systems. To tackle this problem, we propose a probabilistic graphical model which simultaneously incorporates different sources of knowledge such as domain experts knowledge about the context and linguistic knowledge about the sentence structure in a principled way. The model is defined using the declarative language provided by Probabilistic Soft Logic. Experimental results show that the proposed approach, compared to the original distantly supervised set, not only improves the quality of such generated training data sets, but also the performance of the final relation extraction model.
In addition to objective indicators (e.g. laboratory values), clinical data often contain subjective evaluations by experts (e.g. disease severity assessments). While objective indicators are more transparent and robust, the subjective evaluation contains a wealth of expert knowledge and intuition. In this work, we demonstrate the potential of pairwise ranking methods to align the subjective evaluation with objective indicators, creating a new score that combines their advantages and facilitates diagnosis. In a case study on patients at risk for developing Psoriatic Arthritis, we illustrate that the resulting score (1) increases classification accuracy when detecting disease presence/absence, (2) is sparse and (3) provides a nuanced assessment of severity for subsequent analysis.
Ageing is associated with a decline in physical activity and a decrease in the ability to perform activities of daily living, affecting physical and mental health. Elderly people or patients could be supported by a human activity recognition (HAR) system that monitors their activity patterns and intervenes in case of change in behavior or a critical event has occurred. A HAR system could enable these people to have a more independent life. In our approach, we apply machine learning methods from the field of human activity recognition (HAR) to detect human activities. These algorithmic methods need a large database with structured datasets that contain human activities. Compared to existing data recording procedures for creating HAR datasets, we present a novel approach, since our target group comprises of elderly and diseased people, who do not possess the same physical condition as young and healthy persons. Since our targeted HAR system aims at supporting elderly and diseased people, we focus on daily activities, especially those to which clinical relevance in attributed, like hygiene activities, nutritional activities or lying positions. Therefore, we propose a methodology for capturing data with elderly and diseased people within a hospital under realistic conditions using wearable and ambient sensors. We describe how this approach is first tested with healthy people in a laboratory environment and then transferred to elderly people and patients in a hospital environment. We also describe the implementation of an activity recognition chain (ARC) that is commonly used to analyse human activity data by means of machine learning methods and aims to detect activity patterns. Finally, the results obtained so far are presented and discussed as well as remaining problems that should be addressed in future research.
Word-based embedding approaches such as Word2Vec capture the meaning of words and relations between them, particularly well when trained with large text collections; however, they fail to do so with small datasets. Extensions such as fastText reduce the amount of data needed slightly, however, the joint task of learning meaningful morphology, syntactic and semantic representations still requires a lot of data. In this paper, we introduce a new approach to warm-start embedding models with morphological information, in order to reduce training time and enhance their performance. We use word embeddings generated using both word2vec and fastText models and enrich them with morphological information of words, derived from kernel principal component analysis (KPCA) of word similarity matrices. This can be seen as explicitly feeding the network morphological similarities and letting it learn semantic and syntactic similarities. Evaluating our models on word similarity and analogy tasks in English and German, we find that they not only achieve higher accuracies than the original skip-gram and fastText models but also require significantly less training data and time. Another benefit of our approach is that it is capable of generating a high-quality representation of infrequent words as, for example, found in very recent news articles with rapidly changing vocabularies. Lastly, we evaluate the different models on a downstream sentence classification task in which a CNN model is initialized with our embeddings and find promising results.
Claus Weihs合作论文数Department of Statistics|University of Dortmund3