Generative AI can turn scientific articles into narratives for diverse audiences, but evaluating these stories remains challenging. Storytelling demands abstraction, simplification, and pedagogical creativity-qualities that are not often well-captured by standard summarization metrics. Meanwhile, factual hallucinations are critical in scientific contexts, yet, detectors often misclassify legitimate narrative reformulations or prove unstable when creativity is involved. In this work, we propose StoryScore, a composite metric for evaluating AI-generated scientific stories. StoryScore integrates semantic alignment, lexical grounding, narrative control, structural fidelity, redundancy avoidance, and entity-level hallucination detection into a unified framework. Our analysis also reveals why many hallucination detection methods fail to distinguish pedagogical creativity from factual errors, highlighting a key limitation: while automatic metrics can effectively assess semantic similarity with original content, they struggle to evaluate how it is narrated and controlled.
The rapid advancement of AI-generated content has made deepfakes increasingly realistic, posing serious risks to identity security, social trust, and public and democratic institutions. Existing detection systems, typically focused on single modalities such as video or audio, often fail to generalize to new manipulation techniques and cannot effectively detect hybrid or low-effort deepfakes. In this perspective letter, we advocate for a new paradigm in deepfake detection that emphasizes the integration of audio, video, and textual content. We examine the limitations of current systems, including their over-reliance on outdated datasets and limited adversarial robustness. We outline the technical motivations for integrating these modalities and highlight emerging research directions. By aligning detection strategies with the multimodal nature of AI-driven manipulation, we call for a new generation of systems that are more generalizable and trustworthy.
In telecommunications and computer networks, effective incident management highly depends on handling data heterogeneity and providing detailed event context. While knowledge graphs can assist with data integration and AI techniques with contextualization, these aspects are often treated separately, limiting progress toward detailed, explainable, shareable network behavior understanding. This article offers a structured overview, through three perspectives, of how integrating semantic knowledge representations and AI can address this gap. First, we analyze current Network Monitoring Systems (NMS) and Security Information and Event Management (SIEM) systems with respect to the needs of NetOps and SecOps experts. We identify key limitations and discuss enhancements through the incorporation of network topology and operational data, the use of semantic models, and the integration of multiple analytical techniques working together. Next, we review semantic models aligned with NetOps and SecOps, assessing their coverage and expressivity to inform decision support system designers about their potential for reuse, combination, and their capabilities in representing and reasoning about network and system state changes. Finally, we categorize AI techniques by approach, level of determinism, the knowledge representations used, and the incident management steps they address. We identify families of techniques and how each serves operational needs. Additionally, we highlight system design patterns that could maximize, either within families of techniques or through their combination, a detailed understanding of the interplay between network architecture and operational dynamics. We synthesize these three perspectives into a high-level design proposal for a next-generation NMS/SIEM combining logic-based and probabilistic reasoning within a semantic ecosystem, aiming to further automate context-aware incident management in complex Information and Communications Technology (ICT) environments.
Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. Their knowledge is deeply connected but scattered across text, tables, and knowledge graphs. This raises a practical question: when these modalities disagree, how can we detect and explain the conflict? We study this problem as modality-level inconsistency detection. We first introduce a taxonomy of cross-modal knowledge inconsistencies, covering information granularity differences, direct conflicts, temporal changes, and KG incompleteness. We then present Kontrast, an automatic framework that uses Text-to-SPARQL and LLM reasoning to compare table-based answers with KG evidence and categorize the resulting inconsistencies. Experiments on various Table-QA datasets show that cross-modal inconsistencies are common and informative. They reveal not only true knowledge conflicts, but also missing KG structure and temporal mismatches while being limited by Text-to-SPARQL errors and noise. Our analysis shows that text, tables, and KGs can complement and correct one another through systematic comparison. Kontrast provides a practical tool for large-scale knowledge auditing and establishes a benchmark for future work on cross-modal knowledge consistency. Code and data are available at https://github.com/ECLADATTA/KONTRAST.
One of the key challenges to predict odor from the molecular structure is unarguably our limited understanding of the odor space and the complexity of the underlying structure-odor relationships. Here, we introduce an expert-curated taxonomy (ET) that captures the hierarchical relations between odor descriptors for molecular datasets. To quantify the usefulness and relevance of this expert taxonomy, we provide a systematic validation that leverages the predictive performance of machine learning models for structure-based odor predictions, as well as known structure-odor relationships. As a control next to the ET based on semantic and perceptual similarities, we provide a data-driven taxonomy (DT) based on clustering co-occurrence patterns of odor descriptors from our expert-curated molecular dataset. The latter is derived from available datasets in the Pyrfume repository. Both taxonomies (ET and DT) add value in the semantic organization of odor descriptors and provide an avenue for novel insights in molecular structure-odor prediction. Together with in-depth validation steps that highlight the value of odor taxonomies, the quality of the ET is quantitatively assessed. The DT further allows critical evaluation of the expert taxonomy, identification of potential inconsistencies, and a better understanding of the molecular odor space. Finally, we highlight the results of the expert-curated taxonomy by showcasing odor predictions for the case of pear odorants used in perfumery. Both taxonomies as well as a full molecular dataset are made available to the community, providing a stepping stone for a future community-driven exploration of the molecular basis of smell.
In fact-checking applications, a common reason to reject a claim is to detect the presence of erroneous cause-effect relationships between the events at play. However, current automated fact-checking methods lack dedicated causal-based reasoning, potentially missing a valuable opportunity for semantically rich explainability. To address this gap, we propose a methodology that combines event relation extraction, semantic similarity computation, and rule-based reasoning to detect logical inconsistencies between chains of events mentioned in a claim and in an evidence. Evaluated on two fact-checking datasets, this method establishes the first baseline for integrating fine-grained causal event relationships into fact-checking and enhance explainability of verdict prediction.
One of the key challenges to predict odor from molecular structure is unarguably our limited understanding of the odor space and the complexity of the underlying structure-odor relationships. Here, we show that the predictive performance of machine learning models for structure-based odor predictions can be improved using both, an expert and a data-driven odor taxonomy. The expert taxonomy is based on semantic and perceptual similarities, while the data-driven taxonomy is based on clustering co-occurrence patterns of odor descriptors directly from the prepared dataset. Both taxonomies improve the predictions of different machine learning models and outperform random groupings of descriptors that do not reflect existing relations between odor descriptors. We assess the quality of both taxonomies through their predictive performance across different odor classes and perform an in-depth error analysis highlighting the complexity of odor-structure relationships and identifying potential inconsistencies within the taxonomies by showcasing pear odorants used in perfumery. The data-driven taxonomy allows us to critically evaluate our expert taxonomy and better understand the molecular odor space. Both taxonomies as well as a full dataset are made available to the community, providing a stepping stone for a future community-driven exploration of the molecular basis of smell. In addition, we provide a detailed multi-layer expert taxonomy including a total of 777 different descriptors from the Pyrfume repository.
This demo presents an interactive playlist recommendation system that relies exclusively on playlist titles. By fine-tuning a transformer-based language model on clustered playlists, we enable real-time playlist generation for a given title, relying on the semantic meaning of known playlists’ and tracks’ titles. The playlist title provided in input is freely expressed in natural language in a user-friendly web interface. The system is lightweight, fast, and fully accessible through a simple web page.
Large Language Models have shown high performances in a large number of tasks, being recently applied also to support Knowledge Graphs construction. An important step for data modeling consists in the definition of a set of competency questions, which are often used as a guide for the development of an ontology and as a mean to evaluate the resulting schema. In this work, we investigate the suitability of LLMs for the automatic generation of competency questions given an existing ontology. We compare different large language models under various settings in order to give a comprehensive overview of what LLMs can do to support the knowledge engineer.
In recent years, continuous integration and deployment (CI/CD) has become increasingly popular in both the opensource community and industry. Evaluating CI/CD performance is a critical aspect of software development, as it not only helps minimize execution costs but also ensures faster feedback for developers. Despite its importance, there is limited fine-grained knowledge about the performance of CI/CD processes, while this knowledge is essential for identifying bottlenecks and optimization opportunities. Moreover, the availability of large-scale, publicly accessible datasets of CI/CD logs remains scarce. The few datasets that do exist are often outdated and lack comprehensive coverage. To address this gap, we introduce GHALogs, a new dataset comprising 116k CI/CD workflows executed using GitHub Actions (GHA) across 25k public code projects spanning 20 different programming languages. This dataset includes 513k workflow runs encompassing 2.3 million individual steps. For each workflow run, we provide detailed metadata along with complete run logs. To the best of our knowledge, this is the largest dataset of CI/CD runs that includes full log data. The inclusion of these logs enables more in-depth analysis of CI/CD pipelines, offering insights that cannot be gleaned solely from code repositories. We postulate that this dataset will facilitate future CI/CD pipeline behavior research through log-based analysis. Potential applications include performance evaluation (e.g., measuring task execution times) and root cause analysis (e.g., identifying reasons for pipeline failures).
The title of a playlist often reflects an intended mood or theme, allowing creators to easily locate their content and enabling other users to discover music that matches specific situations and needs. This work presents a novel approach to playlist generation using language models to leverage the thematic coherence between a playlist title and its tracks. Our method consists in creating semantic clusters from text embeddings, followed by fine-tuning a transformer model on these thematic clusters. Playlists are then generated considering the cosine similarity scores between known and unknown titles and applying a voting mechanism. Performance evaluation, combining quantitative and qualitative metrics, demonstrates that using the playlist title as a seed provides useful recommendations, even in a zero-shot scenario.
Large-scale Information and Communications Technology (ICT) systems give rise to difficult situations such as handling cascading failures and detecting complex malicious activities occurring on multiple services and network layers. For network supervision, managing these situations while ensuring the high-standard quality of service and security requires a comprehensive view on how communication devices are interconnected and are performing. However, the information is spread across heterogeneous data sources which triggers information integration challenges. Existing data models enable to represent computing resources and how they are allocated. However, to date, there is no model to describe the inter-dependencies between the structural, dynamic, and functional aspects of a network infrastructure. In this paper, we propose the NORIA ontology that has been developed together with network and cybersecurity experts in order to describe an infrastructure, its events, diagnosis and repair actions performed during incident management. A use case describing a fictitious failure shows how this ontology can model complex situations and serve as a basis for anomaly detection and root cause analysis. The ontology is available at https://w3id.org/noria and empowers the largest telco operator in France.
Jacco Van Ossenbruggen合作论文数Centrum voor Wiskunde en Informatica ( CWI )7