A bstract Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI’s ChatGPT, Google’s Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.
Conditional Independence (CI) tests are the statistical engine of constraint-based causal discovery: in algorithms such as PC (Peter-Clark) and FCI (Fast Causal Inference), skeleton pruning and key orientations follow directly from CI decisions. This survey reviews CI testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type settings common in biomedical domains. The survey organizes widely used CI methods into six families: partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning-based. Special emphasis is provided on the robustness layers that address the limitations of these families. For each family, the survey examines when CI decisions reflect the data-generating distribution and when they fail. By this, we link test-level properties, including power decay with conditioning set size and asymmetric type I/II error consequences, to graph-level errors in skeleton recovery and v-structure orientation. The survey also compares adoption across major R and Python libraries and summarizes open challenges, including mixed-type CI testing without discretization, small-sample error control, and strategies for improving scalability of CI-testing.
This study examines the use of subjective and sentimentally charged language in crowd-sourced articles by focusing on time and how articles in environments like Wikipedia tend to evolve as edits are made by multiple contributors. More specifically, we measure linguistic subjectivity (the systematic, asymmetrical use of language) using Mean Abstraction Level, an established subjectivity measure, and polarity (the use of positively or negatively sentimentally charged words) through time. For the latest case, we introduce a new measure called Polarity Density. We focus on Wikipedia biographies and their evolution over time and we perform a detailed analysis per gender and per personality category. Our empirical evaluation provides evidence of increased subjectivity in female biographies as per personality category and per gender, while the same also occurs when considering sentimental charge over time.
Regulations and laws, such as the EU GDPR, require service providers to inform the users about their data collection and processing practices. The existing method used for the portrayal of the rights and responsibilities of both the user and the service provider in terms of data collection, processing and sharing, are the privacy policies, that depict the practices that an organization or company follows when handling the personal data of its users. In this work, we introduce an automated approach, i-Right that classifies the text of privacy policies from the domains of fitness trackers and smart homes, extracting information regarding the eight GDPR user rights present (e.g. Right to Object). Our results show that i-Right achieves classification of the text with high accuracy. The proposed approach could provide a valuable tool for users to understand how their personal data is handled by service providers and to comprehend the possible risks from using their devices. A side contribution of our work is the creation of a labelled dataset of 133 privacy policies to assist the above process.
Aspect terms extraction (ATE), a key subtask for aspect-based sentiment analysis, opinion summarization, and topic modeling aims at extracting grammatical elements (nouns, phrases, and adjectives) from user reviews that reveal the discussed features of the entity under review. These aspect terms are usually the targets of the opinions expressed. Identifying them requires tackling substantial linguistic challenges but, due to the multiple commercial and social applications, significant research effort has been invested in efficiently mining aspects. Recent advances in ATE address methods that exploit a sentence or a word-level encoding of a user review as a solution. This article proposes a novel and effective word- and sentence-level encoding framework, which utilizes a neural network architecture that learns to extract aspect terms. The main advantage of our approach is that it can extract explicit and implicit aspects (i.e., aspects that are not directly mentioned in the user-generated text). We evaluate our method on four widely used datasets where we prove its efficiency against state-of-the-art alternative approaches.
Many patients readily share experiences about their medical conditions and treatments on online social media, which makes these platforms a potentially valuable source of information on adverse drug reactions (ADRs). In this work, the detection of mentions of ADRs in Reddit posts is approached as a multi-label classification problem. A dataset of 537 annotated posts was created by supplementing a publicly available dataset with freshly collected and annotated posts. The labels were mapped to the Medical Dictionary for Regulatory Activities (MedDRA) and their distribution within each MedDRA level guided the creation of 12 data subsets. On each data subset, we applied 4 different multi-label learning methods - Binary Relevance (BR), Classifier Chains (CC), Label Powerset (LP) and random k-labelsets (RAkEL), each associated with 4 different base classifiers: Decision Trees (DT), Naive Bayes (NB), Random Forest (RF) and Support Vector Machine (SVM). The best F-scores were with DT on the data subset based on the 20 most frequent labels at MedDRA Preferred Term (PT) level. The best hamming loss was with the data subset based on all labels at PT level. The type of multi-label learning method did not appear to influence performance significantly. Our results show a promising direction in the use of multi-label classification of ADRs from social media posts for pharmacovigilance purposes.
Central nervous system diseases (CNSDs) lead to significant disability worldwide. Mobile app interventions have recently shown the potential to facilitate monitoring and medical management of patients with CNSDs. In this direction, the characteristics of the mobile apps used in research studies and their level of clinical effectiveness need to be explored in order to advance the multidisciplinary research required in the field of mobile app interventions for CNSDs. A systematic review of mobile app interventions for three major CNSDs, i.e., Parkinson's disease (PD), multiple sclerosis (MS), and stroke, which impose significant burden on people and health care systems around the globe, is presented. A literature search in the bibliographic databases of PubMed and Scopus was performed. Identified studies were assessed in terms of quality, and synthesized according to target disease, mobile app characteristics, study design and outcomes. Overall, 21 studies were included in the review. A total of 3 studies targeted PD (14%), 4 studies targeted MS (19%), and 14 studies targeted stroke (67%). Most studies presented a weak-to-moderate methodological quality. Study samples were small, with 15 studies (71%) including less than 50 participants, and only 4 studies (19%) reporting a study duration of 6 months or more. The majority of the mobile apps focused on exercise and physical rehabilitation. In total, 16 studies (76%) reported positive outcomes related to physical activity and motor function, cognition, quality of life, and education, whereas 5 studies (24%) clearly reported no difference compared to usual care. Mobile app interventions are promising to improve outcomes concerning patient's physical activity, motor ability, cognition, quality of life and education for patients with PD, MS, and Stroke. However, rigorous studies are required to demonstrate robust evidence of their clinical effectiveness.
In the IoT era, sensitive and non-sensitive data are recorded and transmitted to multiple service providers and IoT platforms, aiming to improve the quality of our lives through the provision of high-quality services. However, in some cases these data may become available to interested third parties, who can analyse them with the intention to derive further knowledge and generate new insights about the users, that they can ultimately use for their own benefit. This predicament raises a crucial issue regarding the privacy of the users and their awareness on how their personal data are shared and potentially used. The immense increase in fitness trackers use has further increased the amount of user data generated, processed and possibly shared or sold to third parties, enabling the extraction of further insights about the users. In this work, we investigate if the analysis and exploitation of the data collected by fitness trackers can lead to the extraction of inferences about the owners routines, health status or other sensitive information. Based on the results, we utilise the PrivacyEnhAction privacy tool, a web application we implemented in a previous work through which the users can analyse data collected from their IoT devices, to educate the users about the possible risks and to enable them to set their user privacy preferences on their fitness trackers accordingly, contributing to the personalisation of the provided services, in respect of their personal data.
This paper presents Zenon, an affective, multi-modal conversational agent (chatbot) specifically designed for treatment of brain diseases like multiple sclerosis and stroke. Zenon collects information from patients in a non-intrusive way and records user sentiment using two different modalities: text and video. A user-friendly interface is designed to meet users' needs and achieve an efficient conversation flow. What makes Zenon unique is the support of multiple languages, the combination of two information sources for tracking sentiment, and the deployment of a semantic knowledge graph that ensures machine-interpretable information exchange.
Sentiment analysis is a fast-accelerating discipline that develops algorithms for knowledge discovery from opinionated content. The challenges however, when it comes to analyzing user reviews are plenty. Bad-quality, informal use of language and lack of labels, are only a few obstacles. Most importantly, users, consciously or subconsciously, use different approaches for expressing their opinion about a product or a service. Some of them go sentence by sentence mentioning some positive and negative aspects whereas others provide a mixed piece of text where the reader is supposed to see the big picture to understand the message. In this work, we propose a novel neural network that deals with both situations. Our method, by combining convolutional, recurrent and attention neural networks can extract rich linguistic patterns that reveal the user’s sentiment towards the entity under review. We evaluate our method in nine datasets that represent both binary and multi-class classification tasks. Experimental evaluation indicates that our method outperforms well-established deep learning approaches. Our approach outperformed the competitive methods in 8 out of 9 cases.
Semantic Web technologies are increasingly being deployed in various e-health scenarios, prominently due to their inherent capacity to harmonize heterogeneous information from diverse sources and devices, as well as their capability to provide meaningful interpretations and higher-level insights. This paper reports on ongoing work in the recently started EU-funded project ALAMEDA towards a semantic toolkit for bridging the gap between early diagnosis and treatment in a variety of brain diseases. The toolkit comprises (a) a semantic model serving as the underlying knowledge base for the toolkit; (b) a flexible semantic data integration framework; (c) a semantics-enabled conversational agent for interacting with human users and other components of the ALAMEDA system.
Machine Learning (ML) is now becoming a key driver empowering the next generation of drone technology and extending its reach to applications never envisioned before. Examples include precision agriculture, crowd detection, and even aerial supply transportation. Testing drone projects before actual deployment is usually performed via robotic simulators. However, extending testing to include the assessment of on-board ML algorithms is a daunting task. ML practitioners are now required to dedicate vast amounts of time for the development and configuration of the benchmarking infrastructure through a mixture of use-cases coded over the simulator to evaluate various key performance indicators. These indicators extend well beyond the accuracy of the ML algorithm and must capture drone-relevant data including flight performance, resource utilization, communication overhead and energy consumption. As most ML practitioners are not accustomed with all these demanding requirements, the evaluation of ML-driven drone applications can lead to sub-optimal, costly, and error-prone deployments. In this article we introduce FlockAI, an open and modular by design framework supporting ML practitioners with the rapid deployment and repeatable testing of ML-driven drone applications over the Webots simulator. To show the wide applicability of rapid testing with FlockAI, we introduce a proof-of-concept use-case encompassing different scenarios, ML algorithms and KPIs for pinpointing crowded areas in an urban environment.
In the smart home, vast amounts of data are being collected via various interconnected devices. Although this assists in improving the quality of life at home, often the user is not aware of the details concerning data collection apart from the information available on the provider privacy policy. It is however important to put the user inside this loop of information, so that she is well informed on possible uses of the data and the potential risks that this may entail. Previous works have identified user activity inside the smart home and have pointed out privacy threats. In this work, we go one step further by offering data inference techniques and giving this information back to the user. We use a number of machine learning techniques to draw conclusions about the user routines or activities and we inform the user about our findings concerning data inferences through a dedicated web application. Our aim is toward user-centred privacy and is a proof of concept that can be reused by smart home and Internet of Things service providers in general in order to improve the services offered to the end-users. Our results indicate that a large number of data inferences are possible by using a combination of techniques.
As forecasting becomes more and more appreciated in situations and activities of everyday life that involve prediction and risk assessment, more methods and solutions make their appearance in this exciting arena of uncertainty. However, less is known about what makes a promising or a poor forecast. In this article, we provide a multi-factor analysis on the forecasting methods that participated and stood out in the M4 competition, by focusing on Error (predictive performance), Correlation (among different methods), and Complexity (computational performance). The main goal of this study is to recognize the key elements of the contemporary forecasting methods, reveal what made them excel in the M4 competition, and eventually provide insights towards better understanding the forecasting task.
This commentary introduces a correlation analysis of the top-10 ranked forecasting methods that participated in the M4 forecasting competition. The “M” competitions attempt to promote and advance research in the field of forecasting by inviting both industry and academia to submit forecasting algorithms for evaluation over a large corpus of real-world datasets. After performing the initial analysis to derive the errors of each method, we proceed to investigate the pairwise correlations among them in order to understand the extent to which they produce errors in similar ways. Based on our results, we conclude that there is indeed a certain degree of correlation among the top-10 ranked methods, largely due to the fact that many of them consist of a combination of well-known, statistical and machine learning techniques. This fact has a strong impact on the results of the correlation analysis, and therefore leads to similar forecasting error patterns.
The ability to accurately understand opinionated content is critical for a large set of applications. Models targeting at learning from such content should overcome the inherent difficulties of the data. We propose a novel hybrid neural network embedded in a deep learning framework that can be used for sentiment classification. Our method consists of an independent set of feed forward learning models that are able to identify rich linguistic patterns through recurrent semantic trees. We evaluate our method in four sentiment classification problems that include both binary and multi-class classification tasks. Moreover, we compare our model's prediction accuracy with state-of-the-art methods. We observe that our method outperforms the alternative approaches. The strengths of the proposed approach are due to i) a novel Convolutional Neural Network which can be employed autonomously or as part of a greater framework, ii) a hybrid framework which consists of a set of independent blocks that propagates information and improve the classification task.
We define and solve the problem of event detection and delineation as a task of identifying events and decomposing them to their major sub-events, with a description and a timeline. We propose DeLi, an algorithm that focuses on providing such an understanding of events and sub-events. DeLi, to the best of our knowledge, is the first method that addresses the problem in a generic stream of text, and in an online fashion. Extensive evaluation on social streaming data demonstrates that, by combining the structure of a social network with content attributes, our method outperforms the state-of-the-art techniques.
Sentiment analysis is a challenging task that attracted increasing interest during the last years. The availability of online data along with the business interest to keep up with consumer feedback generates a constant demand for online analysis of user-generated content. A key role to this task plays the utilization of domain-specific lexicons of opinion words that enables algorithms to classify short snippets of text into sentiment classes (positive, negative). This process is known as dictionary-based sentiment analysis. The related work tends to solve this lexicon identification problem by either exploiting a corpus and a thesaurus or by manually defining a set of patterns that will extract opinion words. In this work, we propose an unsupervised approach for discovering patterns that will extract domain-specific dictionary. Our approach (DidaxTo) utilizes opinion modifiers, sentiment consistency theories, polarity assignment graphs and pattern similarity metrics. The outcome is compared against lexicons extracted by the state-of-the-art approaches on a sentiment analysis task. Experiments on user reviews coming from a diverse set of products demonstrate the utility of the proposed method. An implementation of the proposed approach in an easy to use application for extracting opinion words from any domain and evaluate their quality is also presented.
Dimitrios Gunopulos合作论文数Department of Informatics and Telecommunications, National and Kapodistrian University of Athens18