Large Language Models (LLMs) display notable variation in multilingual behavior, yet the role of genealogical language structure in shaping this variation remains underexplored. In this paper, we investigate whether LLMs exhibit sensitivity to linguistic genera by extending prior analyses on the MultiQ dataset. We first check if models prefer to switch to genealogically related languages when prompt language fidelity is not maintained. Next, we investigate whether knowledge consistency is better preserved within than across genera. We show that genus-level effects are present but strongly conditioned by training resource availability. We further observe distinct multilingual strategies across LLMs families. Our findings suggest that LLMs encode aspects of genus-level structure, but training data imbalances remain the primary factor shaping their multilingual performance.
We introduce BUST, a comprehensive benchmark designed to evaluate detectors of texts generated by instruction-tuned large language models (LLMs). Unlike previous benchmarks, our focus lies on evaluating the performance of detector systems, acknowledging the inevitable influence of the underlying tasks and different LLM generators. Our benchmark dataset consists of 25K texts from humans and 7 LLMs responding to instructions across 10 tasks from 3 diverse sources. Using the benchmark, we evaluated 5 detectors and found substantial performance variance across tasks. A meta-analysis of the dataset characteristics was conducted to guide the examination of detector performance. The dataset was analyzed using diverse metrics assessing linguistic features like fluency and coherence, readability scores, and writer attitudes, such as emotions, convincingness, and persuasiveness. Features impacting detector performance were investigated with surrogate models, revealing emotional content in texts enhanced some detectors, yet the most effective detector demonstrated consistent performance, irrespective of writer's attitudes and text styles. Our approach focused on investigating relationships between the detectors' performance and two key factors: text characteristics and LLM generators. We believe BUST will provide valuable insights into selecting detectors tailored to specific text styles and tasks and facilitate a more practical and in-depth investigation of detection systems for LLM-generated text.
Abstract Introduction: Inflammatory and nutritional biomarkers, such as neutrophil to lymphocyte ratio (NLR), C-Reactive Protein (CRP), lactate dehydrogenase (LDH) and albumin, are valuable prognostic indicators, refining stratification strategies in oncological patients. This study aims to delineate the impact of these markers on overall survival in patients with pancreatic adenocarcinoma, in order to aid in clinical risk stratification. Methods: We conducted a retrospective analysis on patients with adenocarcinoma of the pancreas, treated at the Cantonal Hospital Baselland (January 2018 - February 2022). Clinical data, including diagnosis, treatment, and blood markers, were extracted from electronic records. The analysis incorporated the computation of median values for each blood marker, followed by survival analysis and determination of Cox proportional hazard ratios with a 95% confidence interval. Results: Our cohort included 43 patients, 53% women, 44% men, median age 73 years (IQR 14.75) with adenocarcinoma of the pancreas. Most patients (56%) had advanced distant disease and 25% of these did not receive any oncological therapy. At least one cycle of chemotherapy was administered in 85% of patients. NLR<4 was reported in 63% of cases, normal albumin levels in 72%, normal LDH levels in 79%, and normal CRP levels in 47%. In the multivariate Cox regression analysis, there was a significant higher risk for death with metastatic disease (HR 6.216 CI 2.00-19.28, p=0.002), age ≥75 (HR 4.88, CI 1.65-14.39, p=0.004) and CRP>10 (HR 3.38, CI 1.18-9.66, p=0.023) and lower risk with addition of chemotherapy (HR 0.076, CI 0.02-0.29, p<0.001). NLR, LDH and albumin levels did not significantly correlate with higher HR for death. At a median follow-up period of 8.5 months (IQR 14), the median overall survival was 9 months (IQR 14.5) with a worsening prognostic for increased CRP (p=0.00027), NLR>4 (p=0.051), albumin <35g/l (p=0.0024), LDH >250 U/l (p=0.037). Better survival rates were seen in patients receiving chemotherapy 75% vs 20 for those without chemotherapy (p<0.0001) and in those with local disease 80% vs 60% for metastatic disease p=0.034). Age ≥75 years carried a significant poorer prognostic (p=0.0017), while gender was not relevant. However, due to small sample size, age adjusted survival was not carried out and age was considered not relevant for the survival analysis. Conclusion: Beyond the expected factors like stage of disease and therapy, an increased CRP appears to carry significant correlation with hazard ratio for death among patients with adenocarcinoma of the pancreas. In our cohort, abnormal CRP, NLR, albumin and LDH levels demonstrated significant association with shorter survival period, underscoring the potential utility of these biomarkers in prognostic stratification and their incorporation into routine oncological assessments. Citation Format: Elena D. Chiru, Marie Knufike, Martina Sonderegger-Stalder, Raphael Mosimann, Antonia Sgries, Robert Rosenberg, Emanuel Burri, Sandra Mitrovic, Michèle Voegeli, Kristen Mertz, Alessandra Angelini, Marcus Vetter. Prognostic value of serum biomarkers in pancreatic adenocarcinoma: Results from a Swiss single center cohort analysis [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 3680.
Background Air pollution has emerged as an unexpected risk factor for diabetes. However, the mechanism behind remains ill-defined. So far, the lung has been considered as the main target organ of air pollution. In contrast, the gut has received little scientific attention. Since air pollution particles can reach the gut after mucociliary clearance from the lungs and through contaminated food, our aim was to assess whether exposure deposition of air pollution particles in the lung or the gut drive metabolic dysfunction in mice. Methods To study the effects of gut versus lung exposure, we exposed mice on standard diet to diesel exhaust particles (DEP; NIST 1650b), particulate matter (PM; NIST 1649b) or phosphate-buffered saline by either intratracheal instillation (30 µg 2 days/week) or gavage (12 µg 5 days/week) over at least 3 months (total dose of 60 µg/week for both administration routes, equivalent to a daily inhalation exposure in humans of 160 µg/m 3 PM 2.5 ) and monitored metabolic parameters and tissue changes. Additionally, we tested the impact of the exposure route in a “prestressed” condition (high-fat diet (HFD) and streptozotocin (STZ)). Results Mice on standard diet exposed to particulate air pollutants by intratracheal instillation developed lung inflammation. While both lung and gut exposure resulted in increased liver lipids, glucose intolerance and impaired insulin secretion was only observed in mice exposed to particles by gavage. Gavage with DEP created an inflammatory milieu in the gut as shown by up-regulated gene expression of pro-inflammatory cytokines and monocyte/macrophage markers. In contrast, liver and adipose inflammation markers were not increased. Beta-cell secretory capacity was impaired on a functional level, most likely induced by the inflammatory milieu in the gut, and not due to beta-cell loss. The differential metabolic effects of lung and gut exposures were confirmed in a “prestressed” HFD/STZ model. Conclusions We conclude that separate lung and gut exposures to air pollution particles lead to distinct metabolic outcomes in mice. Both exposure routes elevate liver lipids, while gut exposure to particulate air pollutants specifically impairs beta-cell secretory capacity, potentially instigated by an inflammatory milieu in the gut.
Nowadays, researchers unanimously agree on the undeniable importance of mental health. However, the literature related to tracking mental disorders in textual content from social media platforms is heavily inclined towards specific problems. In particular, panic disorder/panic attacks are heavily understudied in the current literature and the relevant resources are missing. Therefore, in this work we focus on collecting an annotated dataset. To this end, in order to mitigate the annotation effort by selectively annotating unlabeled data, we propose an active-learning based approach with uncertainty sampling supported by contextualized (Transformer-based) representations, symptomatic and psychometric features and domain expertise. Our evaluation demonstrates the efficiency of the proposed approach both in terms of classification accuracy and predictions confidence. Our contribution to the research community is an annotated dataset of 13,036 tweets that distinguishes between personal panicking experiences such as panic attacks, other panic-related content and completely panic-unrelated content hoping that it will foster research on the topic.
ChatGPT has the ability to generate grammatically flawless and seemingly-human replies to different types of questions from various domains. The number of its users and of its applications is growing at an unprecedented rate. Unfortunately, use and abuse come hand in hand. In this paper, we study whether a machine learning model can be effectively trained to accurately distinguish between original human and seemingly human (that is, ChatGPT-generated) text, especially when this text is short. Furthermore, we employ an explainable artificial intelligence framework to gain insight into the reasoning behind the model trained to differentiate between ChatGPT-generated and human-generated text. The goal is to analyze model's decisions and determine if any specific patterns or characteristics can be identified. Our study focuses on short online reviews, conducting two experiments comparing human-generated and ChatGPT-generated text. The first experiment involves ChatGPT text generated from custom queries, while the second experiment involves text generated by rephrasing original human-generated reviews. We fine-tune a Transformer-based model and use it to make predictions, which are then explained using SHAP. We compare our model with a perplexity score-based approach and find that disambiguation between human and ChatGPT-generated reviews is more challenging for the ML model when using rephrased text. However, our proposed approach still achieves an accuracy of 79%. Using explainability, we observe that ChatGPT's writing is polite, without specific details, using fancy and atypical vocabulary, impersonal, and typically it does not express feelings.
Novel joint techniques capture both the microscopic context and the mesoscopic structure of networks by leveraging two previously separated fields of research: node representation learning (NRL) and community detection (CD). However, several limitations exist in the literature. First, a comprehensive comparison between these joint NRL-CD techniques is non-existent. Second, baseline techniques, datasets, evaluation metrics, and classification algorithms differ significantly between each method. Thirdly, the literature lacks a synchronized experimental approach, thus rendering comparison between these methods strenuous. To overcome these limitations, we present a uni-fied experimental setup mutually comparing six joint NRL-CD techniques and comparing them with corresponding NRL/CD baselines in three different settings: non-overlapping and over-lapping CD and node classification. Our results show that joint methods underperform on the node classification task but achieve relatively solid results for overlapping community detection. Our research contribution is two-fold: first, we show specific weaknesses of selected joint techniques in different tasks and data sets; and second, we suggest a more thorough experimental setup to benchmark joint techniques with simpler NRL and CD techniques.
The scientific publication output grows exponentially. Therefore, it is increasingly challenging to keep track of trends and changes. Understanding scientific documents is an important step in downstream tasks such as knowledge graph building, text mining, and discipline classification. In this workshop, we provide a better understanding of keyword and keyphrase extraction from the abstract of scientific publications.
Panic phenomenon is one of the main challenges in the current pandemic time. In this work, we aim to explore the approaches to detect the panic-related COVID-19 tweets. Aligned to this, we propose an unsupervised clustering approach considering negation cues as an extracted feature input to the pre-trained model. This task cannot be done by simply applying state-of-the-art transformer models, since we observed that they occasionally fail in handling negations. Hence, we propose to utilize features based on Contextual Valence Shifters (CVS) along with the pre-trained BERT embeddings. We evaluate and compare the approaches in an unsupervised setup, using standard clustering metrics on a large set of COVID-19 tweets. The obtained results show that CVS effectively facilitates negation handling (positive/negative tweet discrimination).
Recently, machine learning is benefiting from advantages of quantum computing which has resulted in a new stream of algorithms known as quantum machine learning algorithms. This paper presents the literature describing implementations of quantum machine learning algorithms in the various quantum machine learning frameworks. In addition, for each of the observed algorithms and frameworks, the literature in which they are described is stated. To the best of our knowledge, this is so far the most comprehensive overview of the existing QML algorithms with their corresponding implementation frameworks.
Background A triple storage (TS) set allows for pathogen inactivation (PI) treatment of triple-dose apheresis platelet products with amotosalen + UVA. We evaluated the quality and metabolic parameters of platelet concentrates (PCs) pathogen inactivated and stored for 7 days. Materials and methods Twelve triple-dose products collected with two different apheresis platforms were treated with amotosalen+UVA. Products were split into three single-dose units. Testing was made pretreatment, after splitting, at days 5 and 7 of storage. Results Single-dose PI PCs had a mean platelet content of 2.89 +/- 0.35 x 10(11). From baseline to day 7, pH remained stable (7.1 +/- 0.1 vs. 7.0 +/- 0.1), pO(2) increased (11.3 +/- 2.4 vs. 18.3 +/- 3.5 kPa) as did LDH (201 +/- 119 vs. 324 +/- 203 U/L) and lactate (3.6 +/- 1.7 vs. 12.1 +/- 1.5 mmol/L) (all p < 0.01); pCO(2) decreased (4.1 +/- 0.8 vs. 1.5 +/- 0.7 mmHg; p < 0.01) and so did bicarbonate (6.6 +/- 1.1 vs. 2.5 +/- 1.4 mmol/L), glucose (5.6 +/- 1.2 vs. 0.4 +/- 0.4 mmol/L) and ATP (3.4 +/- 0.9 vs. 2.5 +/- 1.4 nmol/10(8) platelets) (all p < 0.05). Conclusion Triple-dose PCs processed with the TS sets fulfilled the quality requirements and displayed metabolic changes of expected extent during 7-day storage.
Applying social network analytics for telco churn prediction has become indispensable for almost a decade. However, in the current literature, the uptake does not reflect in a significantly increased leverage of the available information that these networks convey. First, network featurization in general is a very cumbersome process due to the complex nature of networks and the lack of a respective methodology. This results in ad hoc approaches and hand-crafted features. Second, deriving certain structural features in very large graphs is computationally expensive and, as a consequence, often neglected. Third, call networks are mostly treated as static in spite of their inherently dynamic nature. In this study, we propose tcc2vec, a panoptic approach aiming at devising representation learning (to address the first problem) on enriched call networks that integrate interaction and structural information (to overcome the second problem), which are being sliced in different time periods in order to account for different temporal granularities (hence addressing the third problem). In an extensive experimental analysis, insights are provided regarding an optimal choice of interaction and temporal granularities, as well as representation learning parameters. (C) 2019 Elsevier Inc. All rights reserved.
Procedural knowledge is generally dispersed across many experts within or across organizations which might lead to inefficiencies and redundancy. Historically, computers have been well suited to store procedural knowledge but they have lacked the capability to produce natural language text. Nonetheless, recent advances in machine learning permit a higher linguistic coherence which benefits applications with longer text outputs such as procedures. This work closes the gap between human experts and computers by proposing a framework for automatic, computer generation of procedures based on neural machine translation and the BART model. Furthermore, we define two benchmark problems for procedure generation and establish a set of evaluation metrics that can be used as a reference in further work. We demonstrate the potential of this solution on the task of generating cooking recipes based on available ingredients. The evaluation results on the Recipe1M dataset showcase the method’s superiority over other, fairly novel, neural architectures.