
Knowledge Organization (KO) has historically been used to structure biological knowledge, from taxonomy to ontologies. This becomes increasingly challenging as life sciences evolve into a data-intensive domain. The advent of artificial intelligence (AI) has enabled knowledge organization systems (KOSs) to assume active roles in computational workflows rather than serve as passive repositories. This thematic review examines the evolution of KOSs in AI-augmented biological research by situating them within scientific paradigmatic and epistemological shifts. By synthesizing foundational theories from library and information science, philosophy of science, and biological systematics, we propose the Knowledge Organization Analysis Framework (KOAF) to capture bio-KOSs’ developments across functional sophistication, automation degree in system construction, and reasoning and inference capability. Representative empirical studies show that bio-KOSs enable semantic interoperability and data integration, while also contributing to hypothesis generation and reasoning. We argue that advanced bio-KOSs increasingly function as epistemic agents in scientific discovery. This transformation marks KOSs as theoretical frameworks shaping scientific inquiry through AI-KO convergence and highlights the need for future research on accountability, epistemic integrity, and scientific trustworthiness in AI-driven knowledge discovery.
Automatic categorization of fine art paintings across multiple semantic facets, such as artist, style, and genre, is fundamental for large-scale digital archiving, semantic indexing, and knowledge organization of cultural heritage collections. In this paper, we propose convolutional neural network (CNN)-Transformer Hybrid Attention model for art paintings categorization (CTHArt), a CNN-Transformer Hybrid Attention network for multitask art painting categorization. The model employs a dual-branch hybrid backbone that combines a CNN stream for fine-grained local texture modeling and a Transformer stream for global compositional and stylistic context learning. To further exploit inter-facet semantic dependencies, we introduce a Cross-Task Attention Head, which enables task-specific classifiers to exchange information through learnable cross-attention interactions. This design supports coordinated facet prediction consistent with knowledge organization principles. We evaluate the proposed framework on three benchmark datasets. Experimental results demonstrate that CTHArt consistently achieves state-of-the-art performance. The proposed approach provides an effective and scalable solution for artificial intelligence (AI)-assisted knowledge organization of art collections.
This paper investigates the application of machine learning (ML) to the automatic classification of records and archives, framing it as a critical challenge in Knowledge Organization (KO). As digitization creates massive volumes of uncategorized data, the following research question arises: how can fundamental archival principles-such as provenance, original order, and hierarchical description-be translated into this new computational paradigm? This study first synthesizes, based on a multidisciplinary review of archival science, classification theory, KO, computer science, and information science, a proposal of six fundamental guidelines for the responsible application of artificial intelligence (AI) in records and archives. These guidelines connect traditional archival theory with the modern imperatives of trustworthy and explainable AI. Second, we conduct a comparative analysis of 24 published ML experiments, assessing their adherence to these guidelines. Our analysis reveals a significant and troubling disconnect. While most experiments acknowledge the principle of provenance (75.0%), they demonstrate profound neglect of guidelines related to diverse perspectives (25.0%), explainability (16.7%), and, most critically, algorithmic accountability (0.0%). The results indicate that current practices often succeed in basic content categorization but fail in the more sophisticated archival task of preserving archives' evidentiary and relational integrity by treating records as decontextualized data. The study calls for urgently developing an Archival AI Lifecycle-a framework that weaves archival principles, classification theory, and knowledge organization into AI development, safeguarding archival practice's intellectual and ethical integrity in the digital age.
Knowledge Organization Ecosystems (KOE) constitute an analytical metaphor for understanding the growing complexity of contemporary information environments, characterized by interdependence, automation, continuous updating, and algorithmic mediation. Traditionally, Knowledge Organization has relied on relatively stable structures, such as documentary languages and ontologies. However, the incorporation of artificial intelligence and digital infrastructures has expanded this scope, requiring conceptual frameworks capable of integrating epistemological, sociotechnical, and ethical dimensions. This article critically examines KOE as ecosystems composed of people, technologies, and organizations, proposing an interpretative perspective that articulates epistemic plurality, distributed governance, and sociotechnical responsibility. Methodologically, the study is based on an analytical and conceptual review of 26 texts selected for thematic relevance, of which 10 form the theoretical core used to characterize KOE and compare them with Knowledge Organization Systems (KOS). The results indicate that KOS provide semantic stability and terminological control, whereas KOE incorporate collaborative dynamics, algorithmic mediation, and ethical governance, enabling a more comprehensive understanding of knowledge organization in artificial intelligence contexts. The article concludes that the ecosystem paradigm enhances the field's explanatory capacity by recognizing the distributed, ethical, and adaptive nature of contemporary knowledge mediation systems, thereby establishing KOE as a necessary methodological evolution within Information Science.
The rapid growth of digital social information service platforms has fundamentally transformed information diffusion in financial markets, creating an urgent need for knowledge organization systems that can integrate structured financial knowledge with dynamic social interactions. Traditional approaches, which rely on either ontological precision or social network analysis alone, face significant limitations in capturing the complex, multimodal associations that drive modern market behavior. This study proposes a novel Social-Knowledge Big Graph (SKBG) framework to address this gap. The SKBG is designed from a knowledge association perspective and systematically models five core patterns through a three-layer fusion model. The framework semantically grounds social entities in ontologies, enables consistent reasoning and knowledge discovery, and captures temporal dynamics and influence propagation. A case study of China's A-share market and associated investor communities demonstrates the SKBG's practical value. It identifies latent "cross-modal influencer" archetypes, traces semantically enriched misinformation propagation paths that reveal coordination mechanisms, and uncovers socially mediated stock price comovement that traditional models miss. These findings validate the SKBG as a knowledge organization system that bridges the divide between discrete social interactions and complex financial dynamics. This study contributes a robust methodological framework for constructing large-scale socio-financial knowledge organization systems, with important implications for intelligent financial services, risk surveillance, and market regulation in the digital age.
Libraries are widely recognized for providing relevant and reliable information through digital catalogs. However, the traditional formats used in cataloging do not support effective data sharing on the web. To address current informational demands and expand access, the modernization of catalogs is essential. This article proposes a theoretical alignment between the International Federation of Library Associations and Institutions Library Reference Model (IFLA LRM) and Schema.org. Specifically, it establishes a crosswalk mapping attributes from IFLA LRM entities (Work, Expression, Manifestation, and Item) to Schema.org's CreativeWork properties, aiming to support the semantic enrichment of library catalogs. The research adopts a qualitative, exploratory approach, using the crosswalk technique to map attributes of IFLA LRM entities (Work, Expression, Manifestation, and Item) to Schema.org CreativeWork properties. The Library Reference Model provides a strong foundation for bibliographic data modeling, enabling alignment with metadata standards and semantic vocabularies. However, the entities and attributes of IFLA LRM are not fully compatible with the types and properties defined in Schema.org, as the models serve different purposes. Even so, both IFLA LRM and Schema.org offer flexible data structures that allow for adaptation. This study addresses a gap in the literature on interoperability between IFLA LRM and Schema.org and offers insights into their potential alignment for the semantic enrichment of library catalogs. Aligning these models may promote greater openness and interoperability with web search engines, thereby enhancing user experience and increasing interaction with and exploration of library catalogs.
Traditional residences embody deep cultural and academic value. To objectively present their historical significance and stimulate public engagement, this study explores an objective approach to digital narrative design within a "story and discourse" narratology framework. Using the traditional residential buildings of Shanghai's Nanyin Hall as a case study, we collected relevant geospatial data and generated narrative content through visual analysis of data characteristics. We then identified common narrative structures within Shanghai's traditional residential stories using event extraction and feature vector clustering methods. Based on these findings, we designed a story script and presented it through an interactive digital storytelling website. The digital narrative primarily focuses on how environmental changes have influenced site selection, residential construction, and economic development. The story follows the progressive arc characteristic of Shanghai's traditional residential narratives. Throughout the narrative design process, full consideration was given to ensuring the credibility and persuasiveness of the narrative content, providing an operable solution for the design of objective and effective digital narratives.
This work introduces OntoBio, a biographical ontology grounded in Linked Data principles. The study reviews existing ontologies to identify limitations in their ability to represent the diversity and complexity of individual lives. To address these gaps, we developed OntoBio using a tripartite framework encompassing Personality, Environment and Milieu, and Achievements/Milestones. This framework is supported by an expressive vocabulary designed to capture the multifaceted nature of biographical knowledge. OntoBio provides a structural foundation for the semantic integration of existing ontologies that only partially represent biographical data. We present a set of use-case scenarios that motivated OntoBio's development and highlight its key modeling features. To ensure robustness and interoperability, OntoBio was developed following the Yet Another Methodology for large-scale faceted Ontology Construction (YAMO) methodology and in compliance with World Wide Web Consortium (W3C) standards, including Resource Description Framework (RDF), Web Ontology Language (OWL), and SPARQL Protocol and RDF Query Language (SPARQL). Finally, we evaluate the ontology by constructing a biographical knowledge graph illustrating the life of a prominent individual.
This study investigates the interoperability between Schema.org and Resource Description and Access (RDA) as complementary standards for describing bibliographic entities. The research problem arises from Schema.org's limitation in supporting context-dependent value assignment. RDA, as a content standard, is a promising candidate for bridging this gap, especially within the bibliographic context. Integrating these standards presents an opportunity to enhance bibliographic descriptions by combining the web discoverability of Schema.org with the contextual depth offered by RDA. However, realizing this opportunity requires investigating the alignment between specific components of these standards. Adopting an applied approach, this study employed qualitative content analysis to examine potential alignments of Schema.org properties with RDA guidelines, instructions, and vocabulary encoding schemes (VESs) for describing bibliographic entities. The analytical procedures primarily involved mapping samples of these components based on their predefined semantic specifications. The results delineate specific points of interface between the standards, offering key insights into their interoperability in the context of bibliographic entity description. Additionally, prototype implementations are presented to demonstrate practical methods for integrating these standards within real-world metadata management workflows.
This study constructs an ontology that supports both morphological analysis and historical contextualization of bronze weapons from the Shang and Zhou Dynasties, providing semantic support for the development of ancient Chinese military knowledge bases and advancing the structured organization of cultural heritage knowledge. The ontology is developed using a "term-concept-characteristic" methodology and integrates a "weapon-actor-event" semantic chain, enabling the representation of both structural characteristics and contextual relations. To ensure semantic interoperability and scalability, we reused standard ontology vocabularies from International Committee for Documentation Conceptual Reference Model (CIDOC CRM) and Simple Knowledge Organization System (SKOS), and formally represented the ontology in Web Ontology Language 2 Description Logic (OWL 2 DL). The resulting Bronze Weapon Ontology encompasses physical characteristics, functional attributes, manufacturing processes, and historical contexts of bronze weapons, achieving fine-grained semantic modeling across multiple dimensions. Evaluation through structural metrics and SPARQL Protocol and Resource Description Framework (RDF) Query Language (SPARQL)-based competency queries confirms the ontology's logical consistency, semantic expressiveness, and potential for supporting complex reasoning tasks. By providing a unified framework for weapon classification, morphological analysis, and contextual modeling, this ontology offers a robust methodological foundation for the semantic representation of cultural artifacts. It also contributes to broader applications in intelligent cultural heritage services, digital archaeology, and knowledge graph construction for ancient warfare studies.
This article argues that all classifications serve some purposes better than others, and therefore that the idea of an all-purpose classification is untenable. This should not be confused with the view that any classification is as good as any other, either for a specific purpose or for a wider range of purposes. It follows that the aim of a theory of classification is to develop criteria for constructing a classification that best serves a particular aim, and that unclear purposes pose difficulties for this aim. The article discusses a recent paper by Claudio Gnoli that defends the view that an all-purpose classification is possible. The problem of the purpose-laden nature of classification is related to issues such as classificatory pluralism versus monism, realism versus idealism, objective versus subjective classification, and natural versus artificial classification. The article briefly connects such different dichotomies to the issue of the purpose-laden nature of classifications.
Semantic web applications are witnessing a dramatic increase in complexity, data volume, and usage. Likewise, large language models (LLMs) are experiencing significant developments in performance and capabilities. Consequently, LLMs have been utilized in various fields and applications to support primary and secondary tasks. The proven ability of LLMs to process natural language (NL) has opened the door to integration into many tasks, including NL-related tasks such as Knowledge Graph Question Answering (KGQA), which involves translating NL questions into SPARQL queries to retrieve answers from Knowledge Graphs (KG). However, answering questions over domain-specific KGs is challenging due to complex schema structures, specialized vocabularies, and query complexity. Therefore, the development of domain-agnostic and user-friendly KG querying mechanisms has become necessary. Motivated by this need, this paper presents an LLM based approach for translating NL questions into SPARQL queries over domain-specific KG by investigating how various configurations of augmented KG data influence LLM responses. Our approach adopts a streamlined method for zero-shot SPARQL query generation by augmenting LLMs with different arrangements of previously extracted domain-specific KG information. Specifically, our experiments evaluate LLM generated SPARQL responses against twenty manually crafted questions of varying complexity using prompts augmented with different KG information: first, a reduced linearized KG, and second, discrete vocabulary information extracted from a reduced ontology KG. The results indicate that supplementing LLM prompts with discrete vocabulary information extracted from a reduced KG ontology yields competitive performance levels for the target LLM models compared to supplementing them with a reduced ontology. Ultimately, our approach reduces the augmented KG information size while preserving response accuracy, enables off-domain users to interact with domain-specific KG information and retrieve responses through a domain-agnostic interface, and facilitates benchmarking over a wide spectrum of LLM models.
In the field of Information Science, Knowledge Graphs (KG) have emerged as a prominent approach to Knowledge Organization (KO) and Knowledge Representation (KR), supported by Knowledge Organization Systems (KOS) such as taxonomies and thesauri. KGs provide scalable, ontology-enriched structures for managing large volumes of data. Despite criticisms concerning limited semantic expressiveness and modeling quality, KGs remain essential in big-data contexts, where exploratory analysis and visualization are fundamental to identifying patterns and generating insights. Scientific literature indicates that many visualization approaches are algorithm-centered or tool-oriented, often neglecting user-centered design and cognitive ergonomics. To address these limitations, this study aims to identify and evaluate visualization techniques that support user-centered design, focusing on cognitive and usability needs in exploratory and interactive analysis. The methodological approach is based on Design Science Research (DSR), combining a Systematic Literature Review (SLR) and semi-structured interviews to identify user requirements and inform the design of visualization solutions. As a result, the Deigmata system was developed-a user-centered KG visualization tool that facilitates collaborative analysis across multiple user profiles. The tool integrates interactive process triggers that reduce exploration barriers, particularly for non-expert users.
Artificial intelligence (AI) is profoundly reshaping production modes, lifestyles, and organizational paradigms. Enhancing individual adaptability and creativity in this intelligent era has emerged as an imperative research agenda within public literacy studies. This study investigates the evolution of themes in AI literacy research, with a focus on the evolutionary mechanisms through which the domain’s knowledge architecture transitions from techno-cognitive accumulation to synergistic innovation ecosystems. Employing text-mining methodology, we analyze academic literature spanning 2015–2024. Following standard processing steps—tokenization, stopword removal, and lemmatization—we apply Latent Dirichlet Allocation (LDA) topic modeling to extract and analyze latent topics, tracing paradigm shifts in knowledge production during socio-technical processes of innovation. Evolutionary pathways constructed through topic similarity metrics reveal a progression in AI literacy research from knowledge enlightenment to symbiotic knowledge development. LDA2vec hybrid modeling facilitates deep semantic exploration of stage-specific topics, while subsequent topic intensity evolution and semantic network analyses highlight breakthrough points in AI literacy-driven knowledge innovation. This dual-axis analytical framework—integrating topic evolution with intensity dynamics—provides empirical evidence for predicting disciplinary trajectories while establishing novel theoretical paradigms and methodological approaches to AI literacy. Aligned with knowledge organization principles, these frameworks emphasize systematic structuring, cross-domain classification, and semantic interrelations of AI literacy concepts to enhance knowledge accessibility and interoperability.
Researchers have been working on different aspects of representing scientific research papers as machine-processable knowledge graphs. This study extends that work to the domain of social science research papers-specifically aiming to construct scholarly knowledge graphs that represent research results together with their supporting arguments, showing how research claims are incrementally constructed following the discourse structure of the paper. The study analyzed 30 sociology research papers that aimed to find or confirm cause-effect relationships between concepts. It analyzed the causal structure of the research result statements in the papers and investigated what additional information is provided by other statements in the papers performing the argument/rhetorical functions of general statement, topic centrality, literature review, research gap, research objective, research method, research result, and research contribution. A cause-effect information frame that lists all possible roles in a causal information structure was used to guide the analysis. We show that a scholarly knowledge graph representing the research results in a social science research paper must include information extracted from multiple statements in the paper, to provide a comprehensive view of the constructed knowledge. Together, these statements, taken in sequence, form the argument flow of the paper, showing how knowledge is constructed. The study also implemented a prototype graph visualization application to help users examine the causal structure of the research result statements and how that structure is incrementally constructed when the information content of the other statements is viewed in sequence. This tool enables readers to explore conflicting or unexpected results, trace how cause-effect relationships are generalized or specialized, and understand the overall causal information structure accumulated across the paper.
Managing personal electronic records that individuals and households receive and must address in daily life such as bills, receipts, and warranties is often frustrating, and oversights can result in unnecessary costs, including fines for overdue bills or penalties for driving an unregistered vehicle. Important personal records sent as a hyperlink in an email rather than an attachment may not remain available. Systematically saving and sorting personal electronic records leads to higher levels of satisfaction, reduced oversights such as missed payments, and increased motivation to attend to the management of personal electronic records. Information in personal records is often summarized allowing users to view how much they are spending on various categories such as utilities or subscriptions. This paper discusses findings from a user trial of a prototype application designed to aid personal electronic records management and task management at home by downloading, reading, analyzing and summarizing the content of personal records. Findings suggest a personal records management application can assist with timely task completion, such as paying bills, simplifying personal records management, reducing oversights and improving records management. The prototype alerted users that they may need to retrieve records that they did not anticipate needing again. Easier management and tracking of expenses and identification of unnecessary spending may encourage users to enhance their records management. The ability to extract information from personal records can generate innovative ideas for making everyday life easier. This study adds to the body of knowledge in the development of personal records management applications and to the wider domain of personal information management.
Editorial records of a knowledge organization system (KOS) are useful for tracking the provenance of how and why a concept evolves over time. In this study, we explore the use of Large Language Models (LLMs) in obtaining structured provenance information from editorial records of KOSs. Specifically, this study focuses on one type of provenance, namely, "Warrant", which refers to external sources for decision-making on KOS changes. This study presents examples based on four Dewey Decimal Classification (DDC) editorial exhibits, each exhibit containing a discussion and proposed actions for change on a DDC topic. To explore whether LLMs can be used to extract Warrant from these DDC exhibits, we design experiments to test the models' consistency and factual accuracy. We use GPT-4o-mini with the Retrieval Augmented Generation (RAG) approach. For the system and user instructions, chain-of-thought and few-shot prompting strategies are used. For consistency, we test the impact of repeated prompting on ChatGPT's performance; for factual accuracy, we assess whether varying temperature settings yield divergent outputs. Our findings show that consistency and factual accuracy are maintained on most categories of warrant information (Document, Literature, and Concept Scheme), with an average F1-score greater than 70%. For extracting "Concept", the performance is low, with an average F1-score ranging from 30-40%. This demonstrates that ChatGPT is promising for extracting most warrant information from editorial records but still requires a human-in-the-loop verification step to fact-check the concepts extracted. Finally, a recommended process for KOS provenance documentation using LLMs is provided.
Chinese traditional paper-cutting, an important form of intangible cultural heritage (ICH), vividly reflects the richness of China's historical and cultural legacy. However, due to the unique characteristics of ICH, its preservation and transmission face significant challenges. This paper explores knowledge organization methods for traditional Chinese paper-cutting ICH proposes the construction of a knowledge ontology model to achieve systematization, standardization, and digitization. The goal is to provide robust support for the preservation, inheritance, and innovation of this cultural tradition.
Ancient Chinese guqin books hold significant historical and academic value within traditional musical literature, representing a key expression of ancient musical aesthetics and artistic accomplishment. However, current research faces challenges due to a lack of systematic organization and in-depth exploration, which limits the depth of their study and application. This study adopts digital humanities methodologies to explore feasible approaches for semantic modeling of ancient Chinese guqin books, aiming to achieve fine-grained organization, intelligent processing, and dynamic inheritance of these texts. This study utilizes digital humanities methodologies to organize, analyze, and revitalize ancient Chinese guqin books. The research employs a structured workflow that consists of three main steps: ontology construction, knowledge graph construction, and knowledge graph visualization and application. First, key concepts within ancient Chinese guqin books are defined through ontology, which includes the creation of class hierarchies and relational attributes. Using Prot & eacute;g & eacute;, this ontology is constructed and validated to ensure semantic accuracy. Next, the ontology is mapped to the Neo4j graph database to create a knowledge graph that represents multi-dimensional relationships between guqin compositions, related personas, and historical contexts. Finally, the knowledge graph is visualized and queried using Cypher to uncover hidden knowledge and facilitate deeper exploration of ancient Chinese guqin books. The semantic modeling approach proposed in this study enables the representation of the complex semantic knowledge embedded in the fragmented and diverse resources of ancient Chinese guqin books in a simplified and intuitive triplet format. This method facilitates the effective integration and clear presentation of multi-source knowledge. Additionally, intelligent querying and graph-based reasoning techniques are employed to uncover hidden knowledge associations, enabling the extraction of valuable insights and knowledge discovery. These findings not only enhance the understanding of ancient guqin texts but also provide perspective and methodological references for research in related fields of the humanities. This study introduces a novel, systematic approach for the organization and application of ancient Chinese guqin books in the digital age, advancing the preservation and modernization of historical knowledge. It also lays the theoretical foundation for constructing a knowledge resource system in the era of digital intelligence that integrates "knowledge organization, data mining, and interactive perception", as well as a technology-driven framework that connects knowledge, intelligence, and interconnectivity. The integration of ontology and knowledge graphs offers an actionable framework for knowledge discovery and serves as a valuable reference for future digital humanities applications in the preservation of ancient texts.