
We present EmbER, a new neuro-symbolic method for embedding-based entity comparison and ranking in text-rich knowledge graphs (KGs) that have frequent textual fields. Traditional symbolic and homogeneous embedding approaches struggle with distinguishing the semantic nuances of diverse attribute data types in text-rich KGs. EmbER addresses this problem by partitioning KG attributes into categories - numerical values, short terms, and long textual descriptions - guided by the KG schema, and applying type-specific embedding strategies to each data type. The resulting embeddings are then combined into a unified entity representation for similarity computation and relevance ranking of entities. To evaluate EmbER, we construct a curated subset of DBpedia entities, focusing on Irish populated places, e.g. towns and counties. We use SMATCH, a graph-based similarity metric commonly used to evaluate semantic structures, along with human evaluation, to build a ground truth of entity similarity and ranking. Results show that EmbER outperforms a symbolic query-driven method and general KG embedding models, and is competitive with strong textual baselines. These results demonstrate the effectiveness of schema-guided data type-specific embeddings for entity representation, which offer the potential to enhance applications such as semantic search, entity linking and recommendation systems.
Property graphs and knowledge graphs are two dominant paradigms for graph-based data modeling, yet they remain structurally and semantically incompatible. Property graphs support richly attributed relationships but lack formal semantics and reasoning. Knowledge graphs, grounded in RDF, RDFS, and OWL, enable inference and interoperability but cannot natively represent edge metadata. This paper introduces a categorical framework that bridges this gap through a functorial transformation from typed, attributed property graphs to RDF-star knowledge graphs enriched with OWL/RDFS semantics. The transformation preserves edge attributes using quoted triples, consolidates repeated relationships, and ensures semantic fidelity through alignment functions. This formalization provides a foundation for integrating the structural expressiveness of property graphs with the logical rigor of knowledge graphs, paving the way for interoperable, ontology-aware graph systems that combine flexible data modeling with semantic reasoning.
Epistolary letter collections are stored in distributed local archives as letters are sent from one place to another. To find and study letters of a particular person or group on a global level, data from different local sources can be aggregated and harmonized into a global knowledge graph (KG). This paper argues that it is important to understand possible quality issues of the global KG that may arise due to the heterogeneity of the local datasets, aggregation process, and mutual linkedness of the local data. For example: In what ways do the local collections enrich each other? How complete is the aggregated dataset? Are there duplicates or misaligned entities and concepts in the aggregate? This paper presents a set of data-analytic tools to address such issues in order to support data literacy in Digital Humanities (DH) research. As a case study, the LetterSampo Finland Linked Open Data (LOD) KG is considered and the results are reported. It aggregates data about almost 1.3 million historical letters sent in the Grand Duchy of Finland (1809–1917) and 118 000 related actors harvested from 18 different archival data sources and 1670 fonds, enriched by data from 12 external databases.
Large Language Models (LLMs) demonstrate impressive capabilities in Question Answering tasks, yet previous research suggests current approaches fail to accurately reflect their real SPARQL query generation capabilities due to memorization effects caused by integrating benchmark datasets into the training data. These effects artificially inflate perceived performance, i.e., the quality of the showing by LLMs cannot actually be achieved on previously unseen queries. This paper presents DynBench – a novel approach to creating high-quality benchmarking datasets that address this challenge in the field of Knowledge Graph Question Answering (KGQA). We develop new datasets based on two datasets – QALD-9-Plus and LC-QuAD – by systematically replacing entities in SPARQL queries with alternatives retrieved from Wikidata and within the corresponding natural-language questions. Our findings confirm that the proposed approach successfully creates new benchmark datasets that can be used for evaluating KGQA systems. Therefore, our approach drastically reduces the risk of memorization effects, thereby increasing the trustworthiness of benchmark results for LLM-based KGQA approaches.
With the growth in the use of knowledge graphs (KGs) in a variety of applications, it is more important than ever to understand how such KGs are constructed. A core component of many construction pipelines is entity alignment, the identification of corresponding entities across KGs. A number of embedding-based approaches have been developed in the literature to address this problem. However, these algorithms are generally based on a few benchmark datasets. These benchmark datasets are not reflective of reality in regard to data quality, given that real-world KG data will often include false triples and incorrect alignments. This paper presents an investigation into how embedding-based approaches can be affected by low data quality, and what this might mean for real-world application. We find that the effects of noise can be both dataset and algorithm specific, but that relative performance remains unchanged. We also show that data cleaning can in some cases be far more effective in improving performance than collecting more data.
Many tasks related to the enactment, interpretation, and application of laws require examining not only the current version of a statute but also its previous versions. For example, existing Finnish online legal databases provide little support for user-friendly comparison of different versions of the same statute. This paper presents a solution for modelling and visualising statute change histories, enabling easier tracking and interpretation of legislative changes over time. A new prototype application Statute History is introduced to use statutory history data in legal decision making. This application is based on extending an existing Knowledge Graph (KG) (Semantic Finlex) and embedding the prototype into an existing open source User Interface (UI) of a legal web service (LawSampo). Although Finnish legislation is used as a case study, the methods and software presented can arguably also be applied in other countries with a similar legislation system. Statute History enables comparison of different versions of statutes and easy access to related preparatory works. To evaluate the application, a user survey with legal experts was conducted. The respondents found the prototype a clear improvement over existing services, although also identified needs for finer-grained change tracking and more intuitive visual design. User-friendly tracking of statute change history and effortless retrieval of associated preparatory works was deemed essential for legal professionals.
Author name disambiguation in scientific publications faces critical challenges in the era of big data, mainly due to homonymy (different authors with the same name) and synonymy (multiple variants for one author). This study addresses the problem using the Hybrid Framework for Author Name Disambiguation (HFAND), which combines a co-authorship ontology to model semantic relationships with deep neural networks, evaluated on the LAGOS-AND benchmark. Three models (MLP, LSTM, GRU) were implemented using Python3/TensorFlow 2 with optimizations in Rust for scalability. Parallelization in Rust reduced training time versus standard implementations that were not able to run due to the high dimensionality of LAGOS-AND, highlighting computational advantages. The main contribution includes the optimization strategy of AND Ontology adaptable to multidisciplinary domains and a reproducible pipeline that unites the efficiency of Rust with the flexibility of Python for deep learning. These innovations improve scholarly metadata management, with practical applications in automated recommender systems and dynamic researcher profiling. Results show that the MLP achieves an F1-score of 90.51
In recent years, the concept of Explainable AI (XAI) has gained renewed interest as a significant research topic that dates back to the 1970s, largely due to the rapid progress in machine learning and the emergence of deep learning models. The main purpose of XAI is to explain the internal mechanism of models, often referred to as “Black Box”, and to give explanation regarding the reasons behind their predictions. A variety of methods, such as LIME, SHAP, and TCAV, have been proposed to explain either specific predictions or the overall behavior of these models. However, even with their impressive effectiveness and widespread adoption across different fields, these techniques struggle to propose a user-centered explanation. To handle this issue, researchers have considered ontology as a promoting solution to enrich XAI techniques due to its ability to provide more contextual and structured information about a given domain. Therefore, the main aim of this paper is to review and analyze these studies while presenting our proposition of this amalgamation in the area of Information Extraction (IE).
Ontology matching (OM) methods have a wide spectrum of applications, but most matchers find equivalence only correspondences. This limits their applicability in many real-world scenarios, where relations such as subclass, superclass, and overlap are needed. Such correspondences challenge not just the matchers. Alignments with correspondences beyond equivalence can have many equivalent forms and they can be imprecise or partially wrong, which challenges fair comparison. Furthermore, availability of real-world benchmark datasets is rare. Alignments between product classifications typically comprise correspondences beyond equivalence and reference mappings are available for standards like ETIM, eClass, GPC and UNSPSC. We introduce the notion of Product Master Data Model that captures classification systems and describe a construction method for benchmark datasets. The method is based on a class-based alignment representation, named isAmong alignment, which supports inference of correspondences (relation typing) and a standardized representation of alignments. It also supports a fair evaluation method, named isAmong evaluation, that assesses an alignment based on the degree of overlap with a reference alignment, measured using leaf-level coverage rather than weights. We present results that highlight current capabilities and gaps. The real-world benchmark datasets and the isAmong evaluation method are parts of a new OAEI track: Beyond Equivalence.
Over the past decade, the exponential growth of the Electric Vehicle (EV) industry has experienced unprecedented surge and diversification. This multifaceted field results in building comprehensive, cross-domain, extensible, sophisticated knowledge management systems that incorporate future needs and address battery-related information integration challenges. It needs sophisticated knowledge representation techniques such as ontologies and knowledge graphs (KGs) leveraging federated approaches to integrate diverse, disparate data from distributed sources, resolve data interoperability challenges, and also present in a standardized format to build various battery-related services on top of virtual data integration layer, without the need for data materialization. We propose the state-of-the-art federated virtual knowledge graph (FVKG) framework embedded with the virtualized knowledge graph (VKG) methodology to handle the auspicious challenges effectively across distributed environments. The suggested FVKG framework offers a unified view of scattered data sources and different models to create a virtual data federation leveraging Ontop, resolving data bottlenecks efficiently. The FVKG assists in automated data mapping from diverse, relational sources, enabling intuitive queries based on domain-centric federated ontology and loads into the VKG intelligently. The FVKG utilizes a virtualized technique to reduce data migration, guarantees low latency and freshness, and facilitates real-time access while upholding integrity and coherence throughout the federation system. The FVKG incorporates ontology-based data access (OBDA) to build a monolithic ontological model, integrating ontology-driven artifacts and ensuring semantic alignment using schema mapping techniques. As a result, the FVKG targets enabling more efficient battery performance analysis, predictive maintenance, and strategic decision-making in the EV ecosystem.
Computer-aided research techniques for accelerating scientific discovery in polymer (and materials) science has continued to grow in both utilization and access. There remain limitations, however, especially in the curation of data. Currently, data is primarily extracted and compiled from publications and other forms of text manually. This process can be time-consuming; and many existing forms of representation are rigid, unable to account for the evolution of data. Knowledge graphs – and ontology – provide a representation that allows for the complex nature of polymer data but still need to be populated with data from literature. Given the recent successes of large language models in interpreting massive corpora, we propose a pipeline for populating a modular knowledge graph that captures state-of-the-art polymer characterizations in combination with experimental metadata and methodology. In this work, we present different variations of this pipeline, demonstrating which configurations yield desirable and undesirable results.
Building a geo-historical reference dataset of geographical entities enables a wide range of applications, such as the study of urban dynamics. Several approaches in the literature demonstrate the feasibility of constructing such references, but they typically require structured and homogeneous datasets at multiple points in time, also known as snapshots. However, the increasing availability of digitized archival sources, combined with advances in information extraction methods, now makes it possible to produce large volumes of heterogeneous, fragmented, and incomplete data about past geographic entities and their evolution. Existing approaches to building geo-historical reference datasets are not yet well suited to integrating such data in a satisfactory way. In this paper, we propose a method to reconstruct the spatio-temporal evolution of geographic entities from heterogeneous and fragmented data originating from various sources. We also explain how the consistency of the resulting data graph is ensured. Finally, we evaluate our method by applying it on an area in the east of Paris, using sources spanning from the 18th century to the present day.
This research addresses a knowledge gap on the effect of trauma center organization on patient outcomes. To allow trauma care stakeholders to explore data from their center and identify potential opportunities to improve patient outcomes by adjusting organizational patterns, we are creating knowledge graphs from data on relevant parameters. These include data on organizational parameters of trauma centers, currently an understudied topic, and patient outcome data. We have recruited 42 trauma centers in the US and have collected data from those sites. Our knowledge graph uses an ontology to aid knowledge collection and data integration across medical institutions. We are presenting a novel tool, the Knowledge Path Explorer, that allows users to query and expand the knowledge graph at their own pace and based on their needs. This paper presents the basics of the Knowledge Path Explorer, the extension of the underlying ontology to capture trauma patient outcomes, and the data curation process underlying the knowledge graph.
Agricultural decision-making involves complex interdependencies between goals, risks, resources, and environmental factors. Traditional support systems often lack the contextual awareness and adaptability needed to assist farmers in navigating dynamic conditions. This paper introduces FarmerLikeMe, a novel framework to support goal- and risk-aware decision-making in agriculture. FarmerLikeMe combines three key components: (i) a goal-oriented model using the i* model and RiskML to explicitly capture learning experiences ( ℒE ) as farming intentions, farm operations, environmental data, agronomic practices, and ecological risks; thereby enabling farmers to interact through a controlled interface; (ii) a causal knowledge graph, which serves as a collective knowledge base used to seek and share ℒE of farming practices, fostering experience sharing and collective responses to climate challenges; and (iii) An explaining module leveraging LLM-enhanced knowledge graphs and Graph-based Retrieval-Augmented Generation (Graph RAG) to produce decisions that align with individual farmer goals while accounting for potential risks such as climate variability, crop diseases, and resource constraints. A mobile proof-of-concept shows real-world applicability, bridging knowledge representation, and natural language interaction to support intelligent, explainable, and sustainable farming.
Modern applications increasingly demand real-time analysis of live, dynamic graphs at scale, yet community detection remains a performance bottleneck—particularly for Louvain-based methods, which are inherently batch-oriented and unsuitable for high-velocity updates. In this paper, we present a novel real-time and incremental Louvain algorithm, integrated into a unified graph database framework that leverages Storage-Compute Clustering (SCC) and High-Density Computing (HDC). Our approach enables sub-second to sub-minute runtimes on billion-scale graphs by dynamically pruning change scopes, reusing modularity deltas, and incorporating deep neighborhood expansion. We analyze trade-offs of prior incremental Louvain variants and position our system as an architecturally grounded, accuracy-preserving alternative. Extensive evaluations demonstrate that our framework not only delivers superior runtime performance but also maintains high modularity, making it suitable for real-world use cases such as financial fraud detection, social network monitoring, and infrastructure anomaly detection.
Knowledge Graph completion techniques offer a powerful framework for enriching Protein Protein Interaction (PPI) networks with missing or context specific links. This work presents a Knowledge Graph completion framework that integrates literature derived semantic information into PPI network modeling. Biological networks, particularly PPIs, are inherently dynamic. It changes with physiological conditions such as disease states, tissue types, or environmental stressors. Obtaining experimental measurements like gene expression data for every possible scenario is often impractical, especially in unforeseen situations like sudden pandemic outbreaks, underscoring the need for assigning context specific confidence scores to interactions. Using the STRING database as the foundational interaction network, we augment protein pairs with context specific confident scores. To generalize beyond explicitly annotated interactions, we train a Graph Convolutional Network (GCN) that treats edge prediction based on the similarity score to the context. This enables the model to predict novel, relevant interactions and refine network structures in a biologically meaningful manner. The proposed method provides a scalable approach for generating dynamic, literature-informed PPI networks, facilitating more accurate modeling of condition dependent molecular interactions.
Violence against women constitutes a pervasive global human rights violation, demanding innovative approaches to strengthen legal responses. This paper presents the PREJUST4WOMEN project’s contribution: the construction of the first Legal Knowledge Graph (KG) specifically focused on cases of gender-based violence, as adjudicated by the European Court of Human Rights (ECHR). Developed using a bottom-up methodology, the KG is built from ECHR judgments and rigorously structured according to Linked Open Data (LOD) principles. It integrates established legal ontologies and is designed to answer a set of formal competency questions, ensuring domain relevance and practical utility. The resulting knowledge graph is publicly available as FAIR (Findable, Accessible, Interoperable, and Reusable) data, providing an open SPARQL endpoint for advanced querying and analysis. This work fills a significant gap in legal informatics by providing a semantically rich, interoperable resource that facilitates enhanced transparency, supports predictive justice applications, and offers a replicable framework for constructing specialized legal knowledge graphs in other domains.
Ontology matching (OM) is an important process for enabling the interoperability of heterogeneous systems in general and in the Semantic Web in particular. OM takes ontologies as input and determines alignments as output, that is, a set of correspondences between the semantically related entities of the input ontologies. However, several OM systems returns alignments that need to be validated. This validation remains a challenge because benchmarks only exist in certain specific domains and need to be updated in view of ontological changes. In this paper, we propose a systematic solution, grounded on the exploitation of the power of LLMs and KGs, for automatic the validation of ontology alignments. We add ontology concepts and information from a knowledge graph to the prompt to improve LLM reasoning and provide correct validation while minimizing errors. We experimentally proved that our approach reduces significantly the need for human intervention in the validation process. We demonstrate the effectiveness of our solution with experiments on OAEI (Ontology Alignment Evaluation Initiative) campaign and additional alignment benchmarks.
Organizations often use unstructured text to represent and exchange crucial information, such as product model descriptions, system requirements, and documentation. Even though the information within such unstructured text is precious, extracting it in more structured forms and exchanges is cumbersome, time-consuming, error-prone, and mostly reserved for experts. Combining neural and symbolic techniques, leveraging the integration of Large Language Models and Knowledge Graphs, can boost knowledge extraction and post-processing. While many approaches exist for transforming text into structured knowledge, they mostly lack automation, user integration, visual communication, and, consequently, trust and explainability. To address this gap, we present a user-driven hybrid neuro-symbolic approach that puts users in the loop to steer the transformation of unstructured text into highly semantically structured Knowledge Graphs. We evaluated our approach quantitatively and qualitatively to show that inexperienced users can create high-quality Knowledge Graphs with the same quality as experts, but in a fraction of the time. Our approach demonstrates that human interactions significantly enhance quality. Additionally, we confirmed our approach’s high usability, as measured by the System Usability Scale, and strengthened the trust-building and explainability-contributing aspects through expert interviews.
Explainability and interpretability are cornerstones of frontier and next-generation artificial intelligence (AI) systems. This is especially true in recent systems, such as large language models (LLMs), and more broadly, generative AI. On the other hand, adaptability to new domains, contexts, or scenarios is also an important aspect for a successful system. As such, we are particularly interested in how we can merge these two efforts, that is, investigating the design of transferable and interpretable neurosymbolic AI systems. Specifically, we focus on a class of systems referred to as ”Agentic Retrieval-Augmented Generation” systems, which actively select, interpret, and query knowledge sources in response to natural language prompts. In this paper, we systematically evaluate how different conceptualizations and representations of knowledge, particularly the structure and complexity, impact an AI agent (in this case, an LLM) in effectively querying a triplestore. We report our results, which show that there are impacts from both approaches, and we discuss their impact and implications.