This paper investigates the potential of LLMs for automatically annotating the usefulness, supportiveness, and credibility of search results. These aspects, while essential to the construction of misinformation benchmarks, are expensive and difficult to obtain at scale. Our comparative study suggests that, under certain conditions, LLMs can provide reasonable estimates of usefulness and supportiveness. In contrast, credibility judgments generated by LLMs show almost no agreement with human assessments. This raises concerns for the exploitation of LLMs to assist in the construction of collections that require annotations that go beyond relevance.
Topic modeling has emerged as a crucial tool in the field of natural language processing, enabling the automatic discovery of latent structures in large textual corpora. However, determining the quality of the topics remains a significant challenge, particularly in measuring the coherence of the top words of the extracted topics. Early efforts relied on human judgments, but these approaches are resource-intensive. Automated coherence metrics have since been developed. For example, some measures exploit word co-occurrence, while other methods are grounded in distributional semantics (e.g., employing word embeddings). In this study, we thoroughly explore the application of embedded representations to evaluate the quality of topics. While a number of isolated studies have analyzed the role of specific word representation techniques for measuring topic coherence, a complete picture of their effectiveness is still lacking. This work brings together different embedding-based approaches, including Word2Vec, FastText, GloVe, and BERT, which had been studied separately, and extends prior research by incorporating additional models, such as RoBERTa, ALBERT and MPNET. Topic coherence is measured by computing similarity scores between word embeddings, thus obtaining rich semantic associations that traditional measures may overlook. Our analysis demonstrates that these methods are as effective as, and often surpass, classical coherence measures. Our results contribute to a growing body of research advocating for advanced semantic representations as robust alternatives to traditional approaches in evaluating topic model coherence.
Automatic depression detection from social media has been widely explored as a complementary approach to mental health assessment; however, most existing work has focused on binary user-level classification, paying limited attention to how individual depressive symptoms are linguistically manifested in online discourse. This study addresses this limitation through symptom-level depression detection grounded in the 21 items of the BDI-II. The evaluation considers general-purpose transformer models, architectures adapted to the mental health domain, and approaches based on embeddings leveraging large language models across multiple editions of the eRisk benchmark. Results indicate that embedding-based approaches such as GPT-4+SVM and LLaMA+SVM, provide competitive performance with lower computational cost, while domain-adapted models consistently outperform general-purpose transformers. Symptom-level analysis reveals that affective symptoms are more reliably detected, whereas cognitively complex, behavioral, or sensitive symptoms remain underrepresented in social media text.
Since the release of the first ChatGPT model in 2022, large language models (LLMs) have evolved significantly, and an increasing number of users now turn to these generative information systems for inquiries as sensitive and consequential as those related to health. The primary objective is to identify the main strengths and weaknesses of generative AI systems when responding to information needs as critical as those arising in the health domain. The study was structured using a question–answer format, in which each question corresponded to a user query and each answer represented the output generated by a model in response. The study employed a human evaluation framework involving two distinct panels of clinical experts from different specialties. The evaluation criteria encompassed three dimensions: adherence to medical consensus; presence or absence of inappropriate or incorrect information; and the potential to cause harm to users. GPT-4o mini, Llama 3, and MedLlama 3 were selected as three representative systems for the experiments. This study presents a detailed analysis of the performance of widely used contemporary large language models in addressing common health-related queries posed by online users. The results reinforce the potential of LLMs as tools for online health information seeking among non-expert users. However, the performance limitations identified underscore the need for further studies to monitor the future development of these models. Among them, performance issues have been identified in areas where users may be more vulnerable, leading to the retrieval of clinically incorrect information, particularly in matters relating to rare diseases. Furthermore, it has been noted that these models can become trapped in obsolete medical knowledge due to continuous scientific progress. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement Yes ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study did not involve human subjects in the sense of patient participation, clinical intervention, or the use of identifiable personal health data. The research was based exclusively on publicly available, anonymized search queries, and the outputs generated by large language models (LLMs). As such, no institutional review board (IRB) approval or informed consent was required in accordance with applicable regulations and institutional guidelines. The study was designed with a strong commitment to patient safety and public health, emphasizing that LLMs should not be used as substitutes for professional medical advice, diagnosis, or treatment. The selection of health-related queries was conducted using aggregated and non-identifiable data reflecting common online information-seeking behavior. Care was taken to ensure that no queries contained personally identifiable information or sensitive individual-level data. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The online version contains supplementary material available at https://doi.org/10.5281/zenodo.19135840 The code is publicly available at https://github.com/MarcosFP97/From-Dr.-Google-to-Dr.-ChatGPT
Online health information seeking (OHIS) plays a vital role in individuals’ self-management of health. These include understanding the symptomatology of a potential illness, improving health-related habits, assessing possible risks, and determining whether to seek medical care. Since the release of the first ChatGPT model in 2022, large language models (LLMs) have evolved significantly, and an increasing number of users now turn to these generative information systems for inquiries as sensitive and consequential as those related to health. This study presents a detailed analysis of the performance of widely used contemporary large language models in addressing common health-related queries posed by online users. The primary objective is to identify the main strengths and weaknesses of generative AI systems when responding to information needs as critical as those arising in the health domain. The study was structured using a question–answer format, in which each question corresponded to a user query and each answer represented the output generated by a model in response. The set of queries was derived from the most frequently searched terms on a major web search engine, reflecting real users’ health-related information needs. The study employed a human evaluation framework involving two distinct panels of clinical experts from different specialties. The first panel selected the queries deemed most relevant and clinically significant. The selected queries were then submitted to different LLMs, after which the second panel of experts evaluated the responses generated. The evaluation criteria encompassed three dimensions: adherence to medical consensus; presence or absence of inappropriate or incorrect information; and the potential to cause harm to users. ChatGPT-4, Llama 3, and MedLlama 3 were selected as three representative systems for the experiments. The findings indicate that the models performed reasonably well across the three evaluated dimensions. Based on aggregated statistics of the three models analyzed, 80.4% of the responses adhered to medical consensus, 85.0% provided clinically accurate information, and 100.0% posed no potential harm to users. GPT-4 and MedLlama 3 demonstrated superior performance compared to the base Llama 3 model, primarily due to Llama 3’s higher proportion of clinically incorrect responses and tendency to generate ambiguous answers in which medical consensus was not clearly reflected. Despite these relatively strong performance metrics, the healthcare domain requires particularly high standards; therefore, effectiveness levels of 80–90% remain insufficient for deployment in clinical environments. This study reinforces the potential of LLMs as tools for online health information seeking among non-expert users. However, the performance limitations identified underscore the need for further studies to monitor the future development of these models. Moreover, the use of generative AI systems by individuals without medical expertise should remain limited to supportive or preliminary information-gathering purposes and should never replace consultation with a healthcare professional.
This paper introduces Query Harmfulness Prediction (QHP), a novel extension of Query Performance Prediction that focuses on predicting the potential harmfulness of search results. While traditional QPP methods predict standard retrieval metrics, QHP addresses the growing need for safer information retrieval by anticipating when queries might return harmful but topically relevant results. We investigate three families of predictors: classical pre-retrieval QPP methods, LLM-based strategies leveraging signals such as controversy and misinformation, and a query quality classifier adapted from prior work. Using datasets from TREC and CLEF campaigns, we evaluate these approaches with compatibility harmful as the target measure. Our results show that while traditional QPP predictors capture limited signals of harmfulness, LLM-based methods consistently provide stronger correlations, especially on high-risk queries. These findings establish QHP as a timely research direction for developing safer retrieval systems that balance relevance with user safety.
Social media provides valuable insights into users' thoughts, behaviors, and emotions, offering opportunities for mental health research. In this work, we explore how personality traits and demographic attributes manifest in the online behavior of individuals suffering from mental disorders. Focusing on the Big-5 personality dimensions, we analyze social media users associated with four mental health disorders (Anorexia, Depression, Gambling, and Self-harm), investigating how these traits differ across groups. Using the PANDORA dataset -which provides annotations for personality traits, age, and gender-, we train models for personality prediction and author profiling. These models are subsequently transferred to various eRisk collections. Besides confirming known trends (e.g., high association between anorexia and certain young female groups, or between gambling and young males), our analysis reveals intriguing personality traits. For example, we found high neuroticism and agreeableness, and low extraversion and conscientiousness shared across most disorders. These trends underscore the relevance of these personality traits for these mental health problems. Finally, we conclude by analyzing demographic biases in risk detection systems and show that alert accuracy differs significantly across demographic groups.
Misinformation on the Internet poses significant risks to users seeking health information. This paper addresses the challenge of generating effective health-related queries to promote reliable search results. We propose a method leveraging Large Language Models to generate synthetic narratives that guide the creation of alternative queries. These queries are designed to retrieve more helpful and fewer harmful documents compared to those retrieved by the original user queries. We evaluate the effectiveness of these queries using classic and neural retrieval models across multiple datasets, demonstrating promising improvements in retrieving reputable content.
People frequently experience difficulties when seeking information to complete tasks. To overcome these difficulties, people require help. Regarding struggles with information needs, past research focuses on unclear information requests, such as ambiguous, under‐specified, and ill‐defined queries, and repairing these by user‐led strategies (e.g., clarification). In an exploratory qualitative study where information clerks were interviewed, we, however, found that well‐formed and seemingly reasonable requests can conceal misconceptions inquirers have (e.g., about what information is required for their current task) and, therefore, interfere with information seeking and task completion, too. Besides being more difficult to identify than unclear requests, such hidden misconceptions also undermine current user‐led repair strategies as they cause inquirers to believe they are making appropriate requests. Understanding misconceptions in information seeking and requests concealing these is, therefore, essential to building more effective information systems. Our study contributes to addressing this task: It is the first to provide empirical insights into how misconceptions can negatively influence information requests, information‐seeking conversations, and task completion. Ultimately, our findings highlight that inquirers' perceived information needs can present an unreliable and even counterproductive basis for task support, implying that researchers and professionals should rethink the prevailing focus on user requests in designing information systems.
Conversational agents struggle to answer questions during complex tasks such as do-it-yourself (DIY) projects and cooking due to difficulties in understanding task context and user information needs. This study examines the efficacy of integrating conversational and task context in query and document representations to enhance question answering (QA) performance in cooking tasks. We evaluated three document representations with increasing granularity on two task-based QA datasets with a total sample size of 6217 question–answer pairs: full recipe documents (document-based), segmented recipes by cooking steps (step-based), and detailed task structures (task-based). The results show step- and task-based representations outperform traditional document-based approaches by 10% on average (p<0.05). Task-based representations provide superior performance for fact-based needs (e.g., ingredients, time, equipment) in most cases, while step-based representations better address competence needs (e.g., preparation, cooking techniques). Simple conversational history prepending of two to three turns yielded the best performance, improving results by up to 24% over no context. These results emphasise the importance of selecting a representation that matches the structure of the surrounding task in order to enhance QA performance.
Due to their increasing popularity, researchers and health professionals are actively utilizing social media networks as valuable tools to recognize linguistic patterns associated with mental health. In this research, our aim was to better understand to what extent the Beck Depression Inventory (BDI) could undergo automated screening based on users’ social media feeds. To this end, we conducted different experiments to analyze the prevalence of BDI items on social media. We present an approach to categorizing and ranking BDI items considering the quantity of information that can be obtained from social media posts. Given publications written by people who have personally reported being diagnosed with depression, we run different search methods and, based on the number of elements retrieved, we study the prevalence of BDI symptoms at two levels of coverage. Finally, we investigate the impact of prevalence and various characteristics on the efficacy of automated assessment tools. Our analysis indicates that specific elements occur consistently across various search methods and social media platforms, implying a higher prevalence of related symptoms in the data sets analyzed. Interestingly, some items with low incidence in the data sets are those of the BDI questionnaire, whose responses are more accurately estimated using automated methods.
Search engines (SEs) have traditionally been primary tools for information seeking, but the new large language models (LLMs) are emerging as powerful alternatives, particularly for question-answering tasks. This study compares the performance of four popular SEs, seven LLMs, and retrieval-augmented (RAG) variants in answering 150 health-related questions from the TREC Health Misinformation (HM) Track. Results reveal SEs correctly answer 50-70% of questions, often hindered by many retrieval results not responding to the health question. LLMs deliver higher accuracy, correctly answering about 80% of questions, though their performance is sensitive to input prompts. RAG methods significantly enhance smaller LLMs' effectiveness, improving accuracy by up to 30% by integrating retrieval evidence.
Computational methods for depression detection aim to mine traces of depression from online publications posted by Internet users. However, solutions trained on existing collections exhibit limited generalisation and interpretability. To tackle these issues, recent studies have shown that identifying specific depressive symptoms can lead to more robust and effective models. The eRisk initiative fosters research on this area and has recently proposed a new ranking task focused on developing search methods to find sentences related to depressive symptoms. This search challenge relies on the symptoms specified by the Beck Depression Inventory-II (BDI-II), a questionnaire widely used in clinical practice. It includes symptoms such as sadness, irritability or lack of sleep. Given the input submitted by systems participating in eRisk, we first apply top-k pooling over the systems’ relevance rankings, obtaining a diverse set of sentences. These sentences are judged for relevance, leading to DepreSym, a dataset consisting of 21,580 sentences annotated according to their relevance to the 21 BDI-II symptoms. This dataset serves as a valuable resource for advancing the development of models that monitor depression markers. Due to the complex nature of this relevance annotation, we designed a robust assessment methodology carried out by three expert assessors, including a trained psychologist. As part of this study, we explore the potential of recent Large Language Models (ChatGPT, GPT4 and Vicuna) as assessors in this complex task. We undertake a comprehensive examination of the LLMs’ performance, studying their main limitations and analysing their role as a complement or replacement for human annotators. Finally, we incorporate our dataset into the Benchmarking Information Retrieval (BEIR) framework for a thorough search evaluation. We use state-of-the-art retrieval systems, including lexical, sparse, dense and re-ranking architectures, to gain insights about the dataset’s complexity and identify potential avenues for improvement.
While it is often assumed that searching for information to evaluate misinformation will help identify false claims, recent work suggests that search behaviours can instead reinforce belief in misleading news, particularly when users generate queries using vocabulary from the source articles. Our research explores how different query generation strategies affect news verification and whether the way people search influences the accuracy of their information evaluation. A mixed-methods approach was used, consisting of three parts: (1) an analysis of existing data to understand how search behaviour influences trust in fake news, (2) a simulation of query generation strategies using a Large Language Model (LLM) to assess the impact of different query formulations on search result quality, and (3) a user study to examine how 'Boost' interventions in interface design can guide users to adopt more effective query strategies. The results show that search behaviour significantly affects trust in news, with successful searches involving multiple queries and yielding higher-quality results. Queries inspired by different parts of a news article produced search results of varying quality, and weak initial queries improved when reformulated using full SERP information. Although 'Boost' interventions had limited impact, the study suggests that interface design encouraging users to thoroughly review search results can enhance query formulation. This study highlights the importance of query strategies in evaluating news and proposes that interface design can play a key role in promoting more effective search practices, serving as one component of a broader set of interventions to combat misinformation.
In recent years, there has been a growing research interest focused on identifying traces of mental disorders through social media analysis. These disorders significantly impair millions of individuals' cognitive and behavioral functions worldwide. Our study aims to advance the understanding of four prevalent mental disorders: Anorexia, Depression, Gambling, and Self-harm. We present a comprehensive framework designed for the domain adaptation of models to analyze and identify signs of these conditions on social media posts. The language models' adapting strategy consisted of three key stages. First, we gathered and enriched substantial data on the four psychological disorders. Second, we adapted the different models to the language used to discuss mental health concerns on social media. Finally, we employed an adapter to fine-tune the models for multiple classification tasks (specific to each mental health condition). The intuitive idea is to adapt a language model smoothly to each domain. Our work includes a comparative study of different language models under in- and cross-domain conditions. This allows us to, for example, assess the ability of a depression-based language model to detect signs of disorders such as anorexia or self-harm. We show that the resulting mental health models perform well in early risk detection tasks. Additionally, we thoroughly analyze the linguistic qualities of these models by testing their predictive abilities using conventional clinical tools, such as specialized questionnaires. We rigorously examine the models across multiple predictive tasks to provide evidence of the adaptation approach's robustness and effectiveness. Our evaluation results are promising. They demonstrate that our framework enhances classification performance and competes favorably with state-of-the-art models.
To understand more about cognitive processes in credibility judgements, an electroencephalogram (EEG) study was conducted with 20 participants viewing health-related website screenshots. EEG data, segmented from -100 to 900 ms, underwent pairwise classification using a support vector machine, with decoding accuracy as a dissimilarity measure for Representational Similarity Analysis (RSA). Averaging across participants yielded a mean decoding accuracy (DA) at every time-point t. Various website features, including nine low-level visual features, a visual representation extracted from convolutional neural network model pre-trained with ImageNet (VGG16), and three credibility measures representing judgements of web searchers and experts, were investigated for their role in decoding. The DA-curve revealed early decoding during visual and cognitive processing. RSA showed significant correlations with different representations, various visual features and web users' credibility judgements, however not with the credibility ratings of experts. This research not only sheds light on the cognitive processes underlying information assessment but also contributes to our understanding of the vulnerabilities that influence individuals’ online information-seeking behaviours.
The exploitation of Motivational Interviewing concepts for text analysiscontributes to gaining valuable insights into individuals' perspectives and attitudestowards behaviour change. The scarcity of labelled user data poses a persistentchallenge and impedes technical advances in research under non-English languagescenarios. To address the limitations of manual data labelling, we propose a semi-supervised learning method as a means to augment an existing training corpus. Ourapproach leverages machine-translated user-generated data sourced from social me-dia communities and employs self-training techniques for annotation. To that end,we consider various source contexts and conduct an evaluation of multiple classifierstrained on various augmented datasets. The results indicate that this weak labellingapproach does not yield improvements in the overall classification capabilities of themodels. However, notable enhancements were observed for the minority classes.We conclude that several factors, including the quality of machine translation, canpotentially bias the pseudo-labelling models and that the imbalanced nature of thedata and the impact of a strict pre-filtering threshold need to be taken into accountas inhibiting factors.
In 2017, we launched eRisk as a CLEF Lab to encourage research on early risk detection on the Internet. Since then, thanks to the participants' work, we have developed detection models and datasets for depression, anorexia, pathological gambling and self-harm. In 2024, it will be the eighth edition of the lab, where we will present a revision of the sentence ranking for depression symptoms, the third edition of tasks on early alert of anorexia and eating disorder severity estimation. This paper outlines the work that we have done to date, discusses key lessons learned in previous editions, and presents our plans for eRisk 2024.
In this research, we investigate the effectiveness of Large Language Models (LLMs) in answering health-related questions. The rapid growth and adoption of LLMs, such as ChatGPT, have raised concerns about their accuracy and robustness in critical domains such as Health Care and Medicine. We conduct a comprehensive study comparing multiple LLMs, including recent models like GPT-4 or Llama2, on a range of binary health-related questions. Our evaluation considers various context and prompt conditions, with the objective of determining the impact of these factors on the quality of the responses. Additionally, we explore the effect of in-context examples in the performance of top models. To further validate the obtained results, we also conduct contamination experiments that estimate the possibility that the models have ingested the benchmarks during their massive training process. Finally, we also analyse the main classes of errors made by these models when prompted with health questions. Our findings contribute to understanding the capabilities and limitations of LLMs for health information seeking.
Álvaro Barreiro合作论文数IRLab, Computer Science Department, University of A Coruna, Spain20