
Large language models have advanced automatic question generation, yet hallucinations continue to undermine correctness and instructional reliability. This paper introduces a unified framework that integrates causal-graph-guided chain-of-thought reasoning with a multi-agent hallucination-mitigation architecture to generate accurate and pedagogically sound question-answer pairs. Causal graphs provide structured domain knowledge, while specialized agents collaboratively detect and correct logical, factual, solvability, and computational errors through iterative refinement. A formal hallucination-scoring model guides optimization, enabling lightweight models to achieve high fidelity. Experiments on a large learning platform show up to a 90% reduction in hallucination and a 70% improvement in question quality over baseline systems, demonstrating a scalable foundation for trustworthy artificial intelligence-powered education.
Joint visual attention (JVA) provides important insight into how individuals coordinate attention during social interaction. Egocentric eye tracking enables the study of JVA in natural, multi-user settings. This work presents a multi-stage framework to identify and analyze JVA using egocentric video and gaze data. The approach consists of three steps: spatiotemporal tube-based visual similarity, gaze-guided object detection, and attention pattern analysis using the ambient-focal coefficient K. Results show that object-focused collaborative activities exhibit high JVA, with object detection capturing higher joint attention than visual similarity, whereas conversation-based or independent activities show lower and more fragmented joint attention. Analysis of K reveals convergence during shared object interaction and divergence during independent tasks. Overall, the study demonstrates the value of combining object-level semantics and attention dynamics to understand JVA in real-world settings, with implications for psychology, human-computer interaction, and social robotics.
Understanding how individuals focus and perform visual searches during collaborative tasks can help improve user engagement. Eye tracking measures provide informative cues for such understanding. This article presents A-DisETrac, an advanced analytic dashboard for distributed eye tracking. It uses off-the-shelf eye trackers to monitor multiple users in parallel, compute both traditional and advanced gaze measures in real-time, and display them on an interactive dashboard. Using two pilot studies, the system was evaluated in terms of user experience and utility, and compared with existing work. Moreover, the system was used to study how advanced gaze measures such as ambient-focal coefficient K and real-time index of pupillary activity relate to collaborative behavior. It was observed that the time a group takes to complete a puzzle is related to the ambient visual scanning behavior quantified and groups that spent more time had more scanning behavior. User experience questionnaire results suggest that their dashboard provides a comparatively good user experience.
The development of digital twin for smart city applications requires real-time monitoring and mapping of urban environments. This work develops a framework of real-time urban mapping using an airborne light detection and ranging (LIDAR) agent and game engine. In order to improve the accuracy and efficiency of data acquisition and utilization, the framework is focused on the following aspects: (1) an optimal navigation strategy using Deep Q-Network (DQN) reinforcement learning, (2) multi-streamed game engines employed in visualizing data of urban environment and training the deep-learning-enabled data acquisition platform, (3) dynamic mesh used to formulate and analyze the captured point-cloud, and (4) a quantitative error analysis for points generated with our experimental aerial mapping platform, and an accuracy analysis of post-processing. Experimental results show that the proposed DQN-enabled navigation strategy, rendering algorithm, and post-processing could enable a game engine to efficiently generate a highly accurate digital twin of an urban environment.
XAI requires artificial intelligence systems to provide explanations for their decisions and actions for review. Nevertheless, for big data systems where decisions are made frequently, it is technically impossible to have an expert monitor every decision. To solve this problem, the authors propose an explainability auditing method for image recognition whether the explanations are relevant for the decision made by a black box model, and involve an expert as needed when explanations are doubtful. The explainability auditing system classifies explanations as weak or satisfactory using a local explainability model by analyzing the image segments that impacted the decision. This version of the proposed method uses LIME to generate the local explanations as superpixels. Then a bag of image patches is extracted from the superpixels to determine their texture and evaluate the local explanations. Using a rooftop image dataset, the authors show that 95.7% of the cases to be audited can be detected by the proposed method.
Most existing wearable displays for augmented reality (AR) have only one fixed focal plane and hence can easily suffer from vergence-accommodation conflict (VAC). In contrast, light field displays allow users to focus at any depth free of VAC. This paper presents a series of text-based visual search tasks to systematically and quantitatively compare a near-eye light field AR display with a conventional AR display, specifically in regards to how participants wearing such displays would perform on a virtual-real integration task. Task performance is evaluated by task completion rate and accuracy. The results show that the light field AR glasses lead to significantly higher user performance than the conventional AR glasses. In addition, 80% of the participants prefer the light field AR glasses over the conventional AR glasses for visual comfort.
Motion vector approximation is an integral part of every video coding standard to reduce temporal correlation. Estimating motion necessitates a lot of computation. Several attempts were made to reduce the computation cost in exhaustive search of motion estimation. Test zone search was accepted as benchmark algorithm for fast motion estimation by the most recent video coding standard, versatile video coding. Quality and speed of test zone search completely depends on two parameters (i.e., sub sampling frequency of search space during raster scan and dimension of the search space). Cuckoo search, one of the popular nature-inspired optimization algorithms, is used to optimize the operational parameters of test zone search. The proposed optimization enhanced the speed up to 50% while maintaining or improving the Bjontegaard rate (BD-Rate) and Bjontegaard PSNR (BDSNR).
Testing deep learning systems requires expensive labeled data. In recent years, researchers began to leverage metamorphic testing to address this issue. However, metamorphic relations on image data remain poorly understood. To gain a deeper understanding of these metamorphic relations, we survey common image operations modeling covariate shift, manually classify and categorize the underlying metamorphic relations, and conduct experiments to validate our classifications. In our experiments, we train three popular convolutional neural network architectures on an image classification task. Next, we apply metamorphic operations on input test images and measure the change in classification accuracy and cross-entropy loss. A hierarchical clustering algorithm cluster these results and plots a dendrogram. We compare the groups from manual classification and the clusters from the algorithm to provide key insights. We find that Affine and Noise relations are consistent. Furthermore, we recommend metamorphic relationships to save time and better test deep learning systems in the future.
Augmented Reality (AR) allows users to interact with the virtual world in the real-world environment. This paper proposes a child avatar simulation framework using multimodal data integration and user interaction in the AR environment. This framework generates the child avatar that can interact with the user and respond to his/her behaviors and consists of three subsystems: (1) the avatar interaction system scrutinizes user behaviors based on the user data, (2) the avatar action control system generates naturalistic avatar activities (actions and voices) according to the avatar internal status, and (3) the avatar display system renders the avatar through the AR interface. In addition, a child tantrum management training application has been built based on the proposed framework. And a light machine learning model has been integrated to enable efficient and effective speech emotion recognition. A sufficiently realistic child tantrum management training based on the evaluation of clinical child psychologists is enabled, which helps users get familiar with child tantrum management.
Automatic analysis tools are ubiquitously applied on wireless embedded cameras to extract high-level information from raw data. The quality of images may be degraded by factors such as noise and blur introduced during the sensing process, which could affect the performance of automatic analysis. Object detection is the first and the most fundamental step for the automatic analysis of visual information. This paper introduces a quality adjustment framework to provide satisfactory object detection performance on wireless embedded cameras. Key components of the framework include a blind regression model for predicting the performance of object detection and two distortion type classifiers for determining the presence of noise and blur in an image. Experimental results show that the proposed framework achieves accurate estimations of image distortion types, and it can be easily applied on embedded cameras with low computational complexity to improve the quality of captured images.
In the last decade, we have seen an increase in the need for interpretable recommendations. Explaining why a product is recommended to a user increases user trust and makes the recommendations more acceptable. The authors propose a personalized explanation generation system, PEREXGEN (personalized explanation generation) that generates personalized explanations for recommender systems using a model-agnostic approach. The proposed model consists of a recommender and an explanation module. Since they implement a model-agnostic approach to generate personalized explanations, they focus more on the explanation module. The explanation module consists of a task-specialized item knowledge graph (TSI-KG) generation from a knowledge base and an explanation generation component. They employ the MovieLens and Wikidata datasets and evaluate the proposed system's model-agnostic properties using conventional and state-of-the-art recommender systems. The user study shows that PEREXGEN generates more persuasive and natural explanations.
The aspect-based sentiment analysis (ABSA) task consists of two closely related subtasks: aspect extraction and sentiment classification. However, the majority of previous studies looked into each task separately, limiting their effectiveness. In contrast, the integration of aspect extraction and sentiment classification into a single model improves results. The main focus in this work is to manage these two tasks into a new collapsed model. The proposed model relies upon the bidirectional long short-term memory (Bi-LSTM) architecture. On the one hand, it combines a multi-channel convolution layer with an optimization method for handling the aspect extraction task. On the other hand, it includes an attention mechanism based on the residual block and aspect position information for predicting the appropriate opinion orientation of an aspect. The experimental results demonstrate that the model achieved the best performance.
The 3D end-to-end video system (i.e., 3D acquisition, processing, streaming, error concealment, virtual/augmented reality handling, content retrieval, rendering, and displaying) still needs improvements. This paper scrutinizes the motion compensation/motion estimation (MCME) impact in the 3D video (3DV) from the end-to-end users' point of view deeply. The concepts of motion vectors (MVs) and disparities are very close, and they help to ameliorate all the stages of the end-to-end 3DV system. The high-efficiency video coding (HEVC) video codec standard is taken into consideration to evaluate the emergent trend towards computational treatment throughout the cloud whenever possible. The tight bond between movement and depth affects 3D information recovery from these cues and optimizes the performance of algorithms and standards from several parts of the 3D system. Still, 3DV lacks support for engaging interactive 3DV services. Better bit allocation strategies also ameliorate all 3D pipeline stages while being attentive to cloud-based deployments for 3D streaming.
Lack of standard learning infrastructures in secondary and tertiary institutions in most developing countries have made learning cumbersome for the disabled, as most available learning environments were designed to cater mainly for normal learners with little or no consideration for learners with disabilities, thereby resulting to poor academic performance among these affected groups. This paper implements an ANFIS (adaptive neuro-fuzzy inference system)-based ubiquitous learning middleware to support disabled learners by providing suitable learning content for them. The system evaluation showed the effectiveness of the developed ANFIS-based u-learning system in proffering solutions to some of the challenges faced by disabled learners as the system attained 94% correctness, 88% satisfaction, 78% validation, 78% system simplicity, 88% system feedback, and 94% efficiency for the learners.
Landmark recognition aims to detect popular natural and manmade structures within an image. It is challenging with one of the reasons being the lack of large annotated datasets. Existing work mainly focuses on landmarks located in Europe and North America due to regional and language bias. In this study, the authors build a comprehensive Chinese landmark dataset to complement the current data and to benefit research for landmark recognition. It is done by leveraging the vast amount of multimedia data on the web and utilizing image clustering and retrieval techniques in data preparation and analysis. This results in a Chinese landmark dataset with a total of 42,548 images for 987 unique landmarks. In addition, a landmark recognition model is developed based on advanced deep learning techniques and integrated into a mobile application that allows users to do landmark prediction without the need of internet access or cellular data coverage.
Recent studies have discovered that deep neural networks (DNNs) are vulnerable to adversarial examples. So far, most of the adversarial researches have focused on image models. Whilst several attacks have been proposed for video models, their crafted perturbation are mainly per-instance and totally polluted ways. Thus, universal sparse video attacks are still unexplored. In this article, the authors propose a new method to explore universal sparse adversarial perturbation for video recognition system and study the robustness of a 3D-ResNet-based video action recognition model. A large number of experiments on UCF101 and HMDB51 show that this attack method can reduce the success rate of recognition model to 5% or less while only changing 1% of pixels in the video. On this basis, by changing the selection method of sparse pixels and the pollution mode in the algorithm, the patch attack algorithm with temporal sparsity and the one-pixel attack algorithm are proposed.
Augmentative and alternative messaging systems will allow people to speak to improve sports skills and increase engagement and participation in everyday activities. It is an effective tool that can provide people with more control over contact and minimize dissatisfaction. Augmentative and alternative communication's challenging characteristics include lack of sports learning, social interaction, and lack of participation in sports events. In this paper, video-based augmentative communication design (V-BACD) has been proposed to enhance sports learning, facilitate social interaction, and increase society's participation in communication partners. The communication matrix coding technique is implemented to improve sports knowledge and encourage social contact between partners. Community temporal analysis is introduced to strengthen society engagement in the communication partner in an event. The simulation analysis is performed based on accuracy, performance, and efficiency and proves the proposed framework's reliability.
Cryptography is one of the most used techniques to secure data since antiquity. It has been largely improved by introducing several mathematical concepts. This paper proposes a new asymmetric cryptography approach using combined Arnold's cat map with hyperbolic function and Chebyshev chaotic map for audio and image encryption. The proposed scheme uses Chebyshev map for public and secrete keys generation and the same equation with Arnold's cat map for encryption and decryption. Hyperbolic functions are also introduced replacing regular integer values in Arnold's map. The results show a good and promising efficiency as well as the theoretical discussion. Several future possible improvements are presented in the conclusion.
In the rapidly changing air combat environment, it is quite difficult for pilots to make speedy and reasonable decisions in a very short period due to lack of experience and the uncertainty of perception situation. Hence, the authors propose an intelligent cognitive tactical strategy framework of air combat on multi-source information in uncertain air combat situations for decision support. A fuzzy inferring tree method is proposed to simulate human intellection. Then, to further improve the accuracy of the reasoning results, a genetic algorithm is introduced to optimize the structure and parameters of fuzzy rules. The simulation results show that the proposed model is reasonable, fast, accurate, repeatable, and fatigue-free, which lays a good foundation for future high-end unmanned combat explorations.
This paper introduces an improved HMM (hidden Markov model) for low altitude acoustic target recognition. To overcome the limitation of the classical CDHMM (continuous density hidden Markov model) training algorithm and the generalization ability deficiency of existing discriminative learning methods, a new discriminative training method for estimating the CDHMM in acoustic target recognition is proposed based on the principle of maximizing the minimum relative separation margin. According to the definition of the relative margin, the new training criterion can be equation as a standard constrained minimax optimization problem. Then, the optimization problem can be solved by a GPD (generalized probabilistic descent) algorithm. The experimental results show that the performance of the algorithm is significantly improved compared with the former training method, which can effectively improve the recognition ability of the acoustic target recognition system.