
The growing linked and multidimensional aspects of data on the web create challenges for analysts who need to balance their exploration and exploitation of such data sources. However, this balance is key to unlocking more value out of the data to meet organizational objectives. To address this challenge, the authors propose a framework designed to enable a more balanced web of data exploration and exploitation through a Personal Knowledge Graph Cube ontology. They evaluate the proposed approach through an online prototype developed to enable analysts to construct personalized knowledge graph cubes and reuse the cubes in external analytical tools. The prototype was used by 80 analysts who created 520 cubes from a repository of 482,000 RDF triples. Initial results of a mixed-methods evaluation revealed stable system performance and data manipulation times. Interviews and usability feedback from 31 analysts showed promising System Usability Scale score of 82.18, with qualitative insights of the system’s ability to support more personalized exploration-to-exploitation workflows on the web of data.
In complex traffic environments, conventional vehicle detection methods often show limited precision and robustness when facing distant, small-scale, and occluded vehicles. To address these issues, this research proposes a Multiscale, Multi knowledge, Attention enhanced You Only Look Once (YOLO) model. It is a knowledge-driven multiscale vehicle detection framework for intelligent transportation systems, built on the YOLO version 8 nano model. The framework introduces three modules. The first is a cross-stage partial mixed aggregation network module. This module enhances backbone representation through dynamic multiscale aggregation. The second is a multiscale downsampling module that combines dilated convolution and parallel pooling to preserve vehicle cues during downsampling. The third module is an aspect-ratio perception with cross-attention module that injects aspect-ratio-aware knowledge with cross-attention to adapt to vehicle shape variation and suppress background interference. Experiments on the Vehicle dataset on Kaggle.com and the Berkeley DeepDrive 100K dataset showed that the Multiscale, Multi knowledge, Attention enhanced-YOLO model outperformed the lightweight YOLO baseline models. On the Berkeley DeepDrive 100,000 dataset, it achieved a 48.17% mean average precision @0.5 and 26.70% mean average precision @0.5:0.95, improving YOLO version 8 nano by 3.34% and 2.19%, respectively. Visualization results further confirmed its effectiveness under occlusion and adverse-weather conditions while maintaining real-time efficiency.
Blockchain gaming has evolved from recreational play into a profit-oriented digital activity. However, the factors driving sustained player engagement in gamified finance remain insufficiently understood. This study proposes an integrated model combining the Theory of Planned Behavior and the Value-Intention Framework. Data were collected from 203 valid respondents recruited through Taiwanese online blockchain gaming communities. A two-stage analysis was employed, using structural equation modeling and artificial neural networks. The structural equation modeling results show that attitude, subjective norms, flow experience, perceived ease of use, and reward mechanisms significantly influence continuance intention, whereas perceived behavioral control and perceived risk do not show significant linear effects. Artificial neural network sensitivity analysis further identifies attitude as the strongest predictor and highlights the relative importance of perceived behavioral control, suggesting potential nonlinear mechanisms. The findings provide theoretical and practical insights into player continuance in gamified finance.
Recovering a complete 3D object from a single 2D image is important for semantic visual information systems, where geometry should support interpretation and interaction. Existing methods emphasize visual fidelity or geometric plausibility but insufficiently model semantics, spatial relations, and part-whole structures. This study proposed Hero-3D, an ontology-enhanced semantic-aware framework for single-image 3D reconstruction. It extracted semantic priors and organized them into object-, relation-, and part-level representations. These priors guided depth inference, geometric completion, and shape regularization. A Hero view provided a stable semantic-structural reference, while ViewToken-conditioned features were generated under shared semantic and geometric constraints. Semantic fusion and view attention integrated image priors, cross-view anchors, and latent 3D features, which were decoded into triplane and implicit neural field representations. Experiments showed that Hero-3D improved reconstruction quality, multiview consistency, and semantic structure preservation.
With the rapid growth of online short-term rental platforms, user reviews provide valuable insights into homestay experiences. However, existing studies mainly rely on sentiment analysis or satisfaction modeling, producing abstract evaluations that are difficult to translate into actionable spatial redesign solutions. To address this issue, this paper proposes EdiDiffusion, a diffusion-based framework for localized experience-driven redesign. The method takes user reviews and homestay design images as inputs to build an end-to-end pipeline from experience perception to visual redesign. An evidence-grounded edit plan aligns review semantics with spatial evidence to generate structured edit plans. Based on this representation, a constraint-guided local diffusion editor performs localized redesign through conditional diffusion constrained by spatial evidence, functional requirements, and style priors. Experimental results demonstrate improved semantic responsiveness, editing accuracy, and style consistency.
Unmanned aerial vehicle-based visual systems are essential for tasks such as reconnaissance, disaster monitoring, and traffic analysis; however, tracking remains challenging due to tiny targets, background clutter, and interference from similar objects. The authors propose multi-scale contrastive discrimination tracking, a transformer-based framework for tiny object tracking that integrates three modules: multi-scale detail enhancement, contrastive discrimination attention model for suppressing background interference via contrastive discrimination, and instance decoupling enhancement module for reducing feature coupling and identity drift among similar instances. Experiments on DTB70 and VisDrone2018 show that multi-scale contrastive discrimination tracking achieves 87.9% precision and a 67.4% success rate, surpassing state-of-the-art methods by 2.3% and 2.4%, respectively, while running at 185 frames per second on GPU with only 4.2 GMac and 11.2M parameters, demonstrating strong potential for resource-constrained unmanned arial vehicle deployment.
Agriculture is the backbone of national food security and sustainable economic development, which requires performance analysis of the correct crop choice based on soil nutrients and environmental conditions. However, traditional crop selection heavily depends on farmers’ experience and is not necessarily data driven. To tackle this issue, this study proposes an intelligent decision support system for crop recommendation based on deep learning techniques. Factors such as nitrogen, phosphorus, potassium, soil power of hydrogen, rainfall, and temperature were used to predict suitable crop selection. Several architectures—such as convolutional neural networks-1D, long short-term memory, residual deep multilayer perceptron, Wide and Deep, and ResNet-1D—were evaluated and compared. Among them, the highest performance was achieved by ResNet-1D, with 93.61% accuracy and very high precision, recall, and F1-score values. The system supports precision agriculture by leveraging local data to improve decision making and has been implemented as a Streamlit-based web application that allows real-time recommendations for crops.
As modern product design becomes increasingly complex and personalized, integrating heterogeneous data into automated optimization remains a critical challenge. Current generative methods frequently fail to balance high-dimensional semantic customization with strict performance requirements, resulting in structurally flawed designs. The Semantic-Aware Generative Multimodal Optimization System (GMOS-SGP) proposed herein addresses this dichotomy by fusing visual, textual, and structural inputs through a multimodal perception encoding module. This unified representation drives a conditional variational generator, while a policy-guided adaptive reinforcement agent, supported by a dual surrogate model, ensures reliability through closed-loop feedback. Results from the ScanNet dataset indicate superior geometric accuracy, with GMOS-SGP achieving 87.10% Intersection over Union and a 91.54% F-score, supporting its viability for real-time industrial workflows.
Technological advancements have significantly transformed the labor market, driving an increased demand for upskilling and reskilling. While online education offers opportunities to acquire the necessary qualifications, the abundance of options makes selecting appropriate courses challenging. Existing recommendation systems primarily focus on formal education, often neglecting the needs of non-formal education. To address this gap, this study proposes an ontology-based course recommendation system that integrates Thailand’s Professional Qualification Framework (TPQI) with an ontological framework to support personalized learning pathways. The system recommends courses aligned with the skills required for specific occupations, helping learners achieve their career goals. The proposed ontology is modeled using the Web Ontology Language (OWL) and is implemented within a web-based application specifically designed to integrate seamlessly with it. Instances were created to validate the ontology's ability to answer the defined competency questions (CQs), thereby demonstrating the completeness of the proposed model. The web-based application was further evaluated by 19 participants familiar with the data science profession to assess the relevance of the recommendations and the system’s usability. The results indicated a high level of user satisfaction among participants, suggesting the system's potential usability and relevance within the scope of the proof-of-concept study. Participant feedback also identified potential areas for enhancement, including the incorporation of pre-examination features for skill verification and the extension of the framework to support international occupational standards.
The study presents an intelligent Bayesian optimization framework with semantic-aware decision-making for designing coupled material–process energy conversion systems. A Gaussian process surrogate with heteroscedastic noise modeling encodes experimental uncertainty, while a composite acquisition function fusing expected improvement, probability of improvement, and upper confidence bound, together with an adaptive diversity penalty, dynamically balances exploration and exploitation via real-time model diagnostics. Applied to optimizing five fabrication variables of a near-infrared-to-visible upconversion layer in inverted perovskite solar cells, the closed-loop campaign identifies optimal conditions within six rounds starting from 16 initial experiments, yielding a champion device with 22.41% efficiency and 24.52 mA cm-2 short-circuit current density, gains of 13.18% and 8.4% over the control. The adaptive strategy outperforms a fixed-policy baseline and reduces trials by three orders of magnitude versus grid search, establishing a generalizable artificial intelligence–driven methodology for engineering optimization.
Under market-oriented electricity operations, pumped storage hydropower scheduling is influenced by multi-source market signals, system states, and risk disturbances. However, price volatility and temporal uncertainty make efficient operation difficult. Existing methods often separate market perception from strategy generation, resulting in limited interpretability and decision support. To address this issue, we propose an integrated framework for pumped storage revenue optimization. Specifically, an Anchor-driven Structural Modeling Framework is designed to semantically organize heterogeneous market and operational data and capture dynamic dependencies for robust market state perception. Based on structured state representations, a large language model is introduced to generate dynamic scheduling strategies. CVaR and operational constraints are further incorporated to ensure risk awareness and practical feasibility. Experimental results show improved forecasting accuracy, risk control, and overall revenue.
This study explores automated cyberbullying detection across major social networks and messaging platforms using state-of-the-art large language models in a zero-shot, multimodal setting. Models including LLaMA 4, Gemma 3, and GeminiAI were evaluated on images and videos without domain-specific fine-tuning. The system assigned a continuous score (0–10) to indicate the presence of cyberbullying across four categories: revenge porn, happy slapping, racism, and body shaming. Experiments on over 5,000 multimedia samples from Telegram, Reddit, and X (formerly Twitter) showed that large language model-based approaches achieve competitive performance, with Gemma 3-12B emerging as the most stable, accurate, and ethically compliant model. The results also highlighted the critical role of prompt engineering and multimodal context in detecting subtle or implicit online aggression.
As intelligent manufacturing advances, vision-driven assembly inspection increasingly requires modeling inter-part relations and semantic constraints beyond part-level defect recognition. To address this issue, the Assembly Relation Verification Network (ARV-Net), which reformulates assembly quality inspection as a relation consistency verification problem guided by assembly semantics, is proposed. ARV-Net introduces a relation proposal graph to construct sparse, semantically plausible assembly relations by jointly considering geometric feasibility and category compatibility, effectively reducing noisy relation assumptions. In addition, a constraint-aware verification fusion module is designed to jointly model local visual evidence and cross-relation semantic consistency, as well as to adaptively fuse them via constraint-type-driven gating. Experimental results demonstrate that ARV-Net achieves superior accuracy and robustness over existing methods on an assembly quality inspection benchmark.
Synthesizing compound Chinese characters is pivotal for modern typography and interaction design, yet intricate stroke coupling often leads to structural collapse in existing models. Current methods, which rely on loose implicit mappings, frequently produce disordered layouts and blurred artifacts. To rectify this, the author proposes latent structure-guided diffusion, which explicitly injects geometric priors into the diffusion process. By coordinating a latent cross-condition projector for semantic alignment and a latent deformable layout field for differentiable stroke manipulation, latent structure-guided diffusion ensures precise compositional control. Experiments confirmed its superior performance, noting that it achieved a structural similarity index of 0.95 and a mean edge gradient of 0.94, thus significantly outperforming state-of-the-art baselines in structural fidelity and stylistic coherence.
In this article, a comprehensive and intelligent framework (RL-OGSA-EEOSP) is presented, with the primary goal of reducing energy consumption and extending network lifetime in Mobile Wireless Sensor Networks (MWSNs). It employed an improved Voronoi approach with inspired-RL for targeted node distribution to maximize network coverage, the GSA discrete algorithm with opposition-based learning to enhance exploration for clustering and cluster head selection, and the EEOSP algorithm to determine the next point of the mobile sink’s movement. A multi-criteria re-clustering policy and routing based on sensor node trust were implemented to enhance the findings. The authors demonstrated the efficacy of the strategy in two ways. In the first phase, they showed that the suggested methodology is effective. They then demonstrated that it outperforms other methods and achieves considerable gains in energy efficiency, coverage area, and load balancing, thereby increasing network lifetime. The implemented code is available at https://github.com/R-askarizade/RL-OGSA-EEOSP.
As airport operations grow denser and more constrained, vehicle scheduling becomes critical for airside efficiency. Existing methods encode operational rules implicitly, leading to unstable decision spaces and limited generalization. The authors proposed a semantic-aware bi-context scheduling network that integrates semantic modeling with end-to-end learning by separating feasibility determination from priority ranking. A semantic constraint projection module represented airport operational rules as a semantic knowledge base and projected feasible vehicle-task links at each timestep, constraining the decision space before learning. Based on these links, a bi-context link scoring module jointly modeled local execution context and global system state to rank feasible assignments using a lightweight scoring mechanism. Experiments on multiple simulated airport scheduling scenarios showed that the semantic-aware bi-context scheduling network outperformed baseline methods in allocation rationality and operational efficiency, demonstrating the effectiveness of semantic constraint projection and bi-context modeling.
Voice activity detection (VAD) plays a crucial role in speech processing systems, determining the presence of speech in an audio signal. The accuracy of VAD significantly impacts downstream tasks such as automatic speech recognition and noise-suppressed real-time communication, especially under low signal-to-noise ratio conditions where speech boundaries become blurred and noise interference intensifies. However, in noisy environments, existing methods still face challenges in handling fine-grained local features and long-range temporal dependencies. Furthermore, due to coarticulation effects, gradual energy decay, and transient noise, decision boundaries in VAD often remain unclear. To overcome these limitations, the authors propose a robust and efficient self-supervised distillation framework for real-time VAD. Experiments on the PTDB-TUG (PTDB-TUG: Pitch Tracking Database from Graz University of Technology) corpus demonstrate that the method achieves over 95% accuracy and F1 score even at a signal-to-noise ratio of 0 dB, outperforming state-of-the-art models in both robustness and efficiency.
This paper introduces the Cultural Heritage Acquisition and Digitization Knowledge Graph (CHAD-KG), a knowledge graph (KG) for describing bibliographic metadata and digitization paradata of cultural heritage objects in exhibitions, museums, and collections. Built from two tabular datasets, the data were transformed into the Resource Description Framework (RDF) using CHAD-AP, a Web Ontology Language (OWL) application profile based on standards such as the International Committee for Documentation Conceptual Reference Model (CIDOC-CRM), the Object-Oriented Library Reference Model (LRMoo), the Conceptual Reference Model for provenance metadata (CRMdig), and the Getty Art & Architecture Thesaurus (AAT). Then, a Morph-Knowledge Graph Construction (Morph-KGC) extension allowed the graph’s materialization. Unlike many existing cultural heritage KGs, CHAD-KG placed digitization events and paradata at the core of its semantic model and offers an openly available, modular pipeline for KG generation. The current release (v1.0) contains 66,294 RDF triples, serves as the main metadata source for the digital twin of the exhibition The Other Renaissance – Ulisse Aldrovandi and The Wonders of the World, and supports related digitization work in the national Cultural Heritage Active Innovation for Next Gen Sustainable Society (CHANGES) project. It offers a SPARQL endpoint, a user interface, documentation, and is openly published on Zenodo under a CC0 waiver. The project enhanced semantic interoperability in cultural heritage data.
To enhance spatial detail in remote sensing images, super-resolution (SR) has become essential. Conventional single-image SR suffers from limited information, often producing overly smoothed results. Reference-based SR leverages high-resolution reference images to mitigate this issue, but real-world scenarios face cross-sensor discrepancies and temporal land-cover changes, causing concept omission and mismatch that hinder effective reference utilization. To address these challenges, the authors propose a diffusion-based SR framework with semantic reference matching, namely Semantic Reference Matching and Diffusion Learning for Intelligent Super-Resolution in Visual Information Systems (SRM-DL). It comprises two key modules: (a) Concept Activation, which uses diffusion priors to recover missing structures, and (b) Attribute Concentration, which used a local–global dual-branch alignment to robustly incorporate semantically consistent reference information while suppressing mismatches. Multiscale consistency constraints further align reference and target features across spatial and semantic domains. Extensive experiments validated its effectiveness and practical potential in multimedia visual information systems.
Multimodal sentiment analysis aims to identify emotional tendencies from text, audio, and visual data, but existing methods often struggle with weak temporal modeling within modalities and shallow cross-modal fusion. The proposed temporal modeling and synergistic attention–based multimodal sentiment analysis framework can address these issues. Word-level features are first extracted from all modalities, then modeled using a state-gated long short-term memory network combined with multi-head attention to capture temporal emotional dynamics while filtering noise. A hierarchical collaborative attention mechanism is further designed to enable deep, fine-grained cross-modal semantic interactions. Experiments on the Carnegie Mellon University multimodal corpus of sentiment intensity and multimodal opinion sentiment and emotion intensity datasets show that the modeling and synergistic attention-based multimodal sentiment analysis framework achieves an F1 score of 87.3% and an mean absolute error of 0.426, it achieves a 1.2–1.5% improvement while simultaneously reducing mean absolute error to its lowest value, outperforming existing state-of-the-art approaches and demonstrating its effectiveness in modeling complex multimodal emotions.