
Hand gestures play a crucial role in human-computer interaction, aiding applications like sign language interpretation, autonomous vehicles, and virtual reality. However, designing effective hand gesture recognition (HGR) systems remains challenging due to issues such as cluttered backgrounds and occlusion. This paper proposes a multi-scale oriented hierarchical vision transformer (HiT-HGR) framework, which targets extracting local and global receptive information for HGR from raw images. The HiT-HGR framework includes multi-scale feature assimilator (MSFA) to extract fine to coarse exquisite features with its multi-scale assimilation and a three-fold attention network (TFAN) to capture local to global dependencies. Cohesively, MSFA and TFAN allow the HiT-HGR framework to compile refined low-level to high-level features by considering local to global dependencies. Thus, the resultant features focus on only gesture features, ignore the background complexities, and enhance the performance of HGR. Extensive experiments demonstrate that the HiT-HGR framework outperforms conventional vision transformer models, and state-of-the-art HGR methods.
Pedestrian identification becomes difficult when facial cues are missing due to back-facing views, low light, or occlusions. This paper presents Deep-PPP, a pose-driven identification framework that relies solely on skeletal keypoints rather than facial or appearance features. Using OpenPifPaf for keypoint extraction, the system generates normalized pose-patched representations and embeds them into a discriminative pose-biometric descriptor integrating limb ratios, joint-angle invariants, and structural distances. Evaluated across multiple benchmarks, Deep-PPP achieves 89.5% Rank-1 accuracy and 78.3% mAP on Market-1501, outperforming prior pose-guided methods such as PSE+GFL. Cross-dataset experiments further show a moderate generalization drop to CrowdPose due to increased occlusion, consistent with the challenging conditions of dense scenes. These results demonstrate that pose-based biometrics provide a practical, privacy-preserving alternative for pedestrian identification when facial information is unreliable or unavailable.
Effectively leveraging lengthy evidence for explainability in multimodal rumor detection is a critical challenge, due to difficulty in localizing clues as well as the gap between claim and evidence. To this end, we propose Dempster-Shafer Theory-based Multi-Agent Evidence Localization for Explainable Multimodal Rumor Detection (DMEL), which consists of Evidential Consensus Tribunal (ECT) and Claim Salience Assessment (CSA). More specifically, ECT involves the proposal of candidate clues and explanations by Analyzing Agents, followed by their cross-validation of Verifying Agents, culminating in a Dempster-Shafer Theory-based selection of the most reliable clue. Furthermore, CSA identifies vulnerable tokens in the claim by adding Gaussian noise to their features and measuring the change in their co-attention with the clues from ECT. Finally, the salient claim, image, and evidence are fused via cross-modal attention and jointly trained for rumor detection. Extensive experiments conducted on two public datasets show that our method outperforms the state-of-the-art approaches.
JPEG Universal Metadata Box Format (JUMBF) provides a carrier format for embedding arbitrary data in image file formats and beyond. Through its extensibility and native referencing scheme, it enables the construction of rich data models to accommodate emerging use cases and applications. The proposed implementation framework enables JUMBF adoption in end applications as well as the extension of the standardized data model to cover complex requirements. Through standardized applications built on top of JUMBF, this work demonstrates the modularity and extensibility of the JUMBF data model. It further highlights the practical impact of the developed software by showcasing real-world applications and introducing a comprehensive conformance suite for JUMBF implementers. Endorsed by the JPEG Standardization Committee as a reference software, this library represents the first complete reference implementation of the JPEG Systems suite of standards.
Processing 360◦ images for machine vision tasks can significantly enhance various applications in augmented and virtual reality, autonomous driving, and drone surveillance. However, two main challenges arise in this area: the large feature space and angular distortions. To address this, we propose QML-360◦, a hybrid quantum–classical framework for 360◦ image classification that learns on spherical graphs. It uses a permutation-equivariant quantum head to preserve symmetry under node permutations. To address memory and runtime constraints on near-term quantum devices (and in classical simulation), we adopt a distributed training strategy. Gradient evaluations are parallelized across workers, and split-shot estimation reduces per-device load while maintaining estimator quality. Across datasets, our method achieves competitive performance relative to the baselines under large rotation perturbations while reducing per-device compute and memory. In addition, we establish theoretical formulations for equivariance, gradient variance, and cost scaling, and present empirical results that demonstrate their effectiveness.
Although existing stereo image super-resolutions enhance reconstruction performances by integrating intra-view feature extraction with cross-view interaction, they often overlook the problem of information redundancy between left view and right view during interaction. To address this limitation, we propose a lightweight stereo image super-resolution network called Multi-Scale Refined Interactive Network (MSRINet) by two new components: Multi-Scale Information Aggregation Block (MSIAB) and Cross-View Refined Interactive Module (CVRIM). The designed MSIAB can extract multi-level feature information hiding in views and refined interaction of cross-views by combining large separable kernel attention with local enhancement strategy for capturing multi-scale global context and fine local textures. The designed CVRIM can fully utilize complementary information of left view and right view by employing an efficient cross-view interaction mechanism for reducing redundant information while optimizing feature interaction and fusion across views. Extensive experiments illustrates the effectiveness of our MSRINet by achieving superior performances while with fewer parameters.
In video captioning, current methods often struggle with generating accurate descriptions due to challenges in modeling object interactions, especially predicates that rely on both object dynamics and motion patterns. To address this, we introduce the Behavior Modeling-Aware Video Captioning Network (BMVCap), which enhances video captioning by capturing fine-grained details of interactions, co-occurring objects, and contextual background using a transformer-based encoder for each modality stream. The model integrates these streams through an adaptive fusion mechanism in the transformer decoder, allowing it to generate more precise captions. Additionally, BMVCap employs a caption length control mechanism and optimized reinforcement learning, maximizing rewards from multiple evaluation metrics. Extensive experiments on the MSVD, MSR-VTT, and VATEX datasets show that BMVCap significantly improves captioning accuracy by better modeling complex interactions and activity attributes. The results demonstrate that emphasizing complex interactions and activity attributes leads to substantial improvements in the accuracy and reliability of video captioning.
Intelligence is the ability to solve problems in the context of specific individual needs and capabilities. This reveals why current AI falls short: language processing, however sophisticated, cannot capture multimodal reality or enable personalized action. We’re at a convergence moment where large language models can merge with deep multimodal sensing to create multimodal assistive intelligence (MAI). MAI integrates information across modalities, combines general knowledge with personal context, and enables appropriate action in dynamic situations. The technology components exist. This column introduces MAI and predicts that by 2030, all effective AI will be multimodal, personal, and assistive.
This paper evaluates the effectiveness of interactive, AI-powered virtual instructors in an extended reality (XR) training environment for electrical technicians. A user study was conducted in a 3D reconstruction of an electrical substation, where participants completed a four-step technical training procedure using voice and gesture interactions with the AI instructor. The study examines the instructor’s ability to guide users through a complex, multi-stage workflow and assesses performance outcomes, cognitive load, user experience, and overall training effectiveness. The results demonstrate that AI-guided XR instruction can enhance user engagement and knowledge acquisition while maintaining a relatively low cognitive load in high-risk technical training scenarios. These findings highlight the potential of interactive AI instructors to improve the design of future XR training systems across industrial domains.
The advancement of autonomous vehicle technology has heightened interest in modeling occupant states and behaviors. This paper models occupant circadian state using four modalities: physiological, thermal, linguistic, and acoustic modalities. Moreover, it advances research by developing a fully non-contact pipeline to derive physiological signals from thermal imagery. Linguistic and acoustic features derived from speech are incorporated with the extracted physiological and thermal features to enhance circadian rhythm models and detect enervation states. The proposed approach leverages a multimodal dataset of 36 subjects captured across multiple channels. Comparative analysis between non-contact and contactbased channels reveals that physiological signals from thermal imagery, combined with linguistic and acoustic data, achieve accuracy comparable to or exceeding those of contact-based methods. This non-contact approach achieves accurate identification of energized versus enervated states, reaching 78.6% accuracy. These findings offer a novel framework for integrating unobtrusive sensing technologies in future automotive designs, facilitating next-generation occupant monitoring.
Gait recognition has emerged as a crucial biometric technique because it enables unobtrusive and remote identification. However, silhouette-based methods are often affected by variations in clothing and carrying conditions. To address these challenges, we propose a novel multi-modal approach that integrates Gait Energy Images (GEI) and Skeletal Gait Energy Images (SkeGEI) to enhance recognition performance. GEIs capture appearance-based holistic motion patterns, while SkeGEIs, generated from normalized skeletal representations centred on the neck joint, provide robust structural features. Skeletal coordinates are extracted using HRNet for the CASIA-B dataset and AlphaPose for the OU-MVLP dataset, which are then used to generate SkeGEIs. Our composite fusion model effectively combines GEI and SkeGEI features through Cross-Alignment blocks (CAB), attention mechanisms, and Progressive Fusion Blocks (PFB). Extensive evaluations on the CASIA-B and OU-MVLP datasets demonstrate that the proposed approach achieves superior performance compared to state-ofthe- art methods across diverse conditions. However, the model struggles with crossdataset generalization, particularly in varied clothing conditions. Additionally, an ablation study highlights the effectiveness of each component, validating the superiority of our integrated model.
This paper presents a novel halftoning algorithm based on continuous relaxation and gradient-based optimization. By reformulating the original binary quadratic programming (BQP) problem into a differentiable non-convex objective, the proposed method integrates a perceptual loss with a binarization regularizer to guide the optimization toward high-quality binary outputs. Unlike deep learning-based approaches, our method is model-free, memory-efficient, and fully interpretable, enabling fast convergence without complex training or large memory overhead. Experimental results demonstrate that the proposed method outperforms both classical and neural halftoning techniques in terms of image quality and convergence speed. This work provides a practical and scalable solution for highresolution halftone image generation, particularly suitable for industrial and real-time applications under resource constraints.
Transformer-based U-Net architectures have demonstrated remarkable advantages in low-light image enhancement, where the long-range dependency modeling improves enhancement performance, yet the substantial computational complexity hinders real-time applications. To address this, we propose HaarReformer, a single-stage low-light enhancement network that integrates Haar wavelets with a lightweight Transformer architecture. To improve computational efficiency, an illumination-guided cross-channel context aggregation module is introduced, which fuses cross-channel context with expanded receptive fields to enhance self-attention, striking a balance between global modeling and computational overhead while achieving linear complexity. For high-frequency information preservation, we construct a Haar wavelet-based downsampling scheme with multi-scale invertible decomposition, enabling representative feature learning. Furthermore, an ultralightweight dynamic upsampling module is introduced for efficient feature reconstruction, leveraging a learnable point sampling mechanism within a plug-and-play framework. Experimental results validate that HaarReformer effectively enhances low-light images with low computational overhead, while demonstrating superior performance in preserving color fidelity, structural integrity, and textural details.
Blind face restoration (BFR) aims to recover high-quality (HQ) facial images from degraded inputs with unknown distortions, while preserving identity consistency. Existing methods either rely solely on visual priors or incorporate generic textual cues, struggle with noise sensitivity or insufficient semantic guidance. In this paper, we present a vision-language guided BFR framework that consists of three key components: (1) a Facial Key Attribute Estimation module that leverages structured text to compactly describe facial attributes; (2) a Dual-modal Based Latent Prediction module built on a shared-attention Transformer, which fuses visual and textual embeddings to predict dictionary elements representing the target face; (3) a Text-based Adaptive Feature Filter that dynamically suppresses noise in skip-connected features by aligning them with textual semantics, enhancing the generator’s ability to leverage multi-scale spatial information. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches on multiple benchmarks, particularly under severe degradation.
Out-of-distribution (OOD) generalization is a major challenge in federated graph learning, where client data distributions often differ significantly due to shifts in node features and structural patterns. This non-IID nature leads to poor performance of Graph Neural Networks (GNNs) trained solely on federated task losses. While domain-alignment techniques—such as adversarial training, invariant risk minimization, and kernel-based methods like Maximum Mean Discrepancy (MMD)—offer promising solutions for mitigating distribution shifts, they typically require access to shared embeddings, which compromises user privacy. We propose FedMMD-P, a framework for privacy-preserving alignment in federated out-of-distribution graph learning. Each client computes a compact, noisy RFF-based sketch of its node-embedding distribution, enabling a DP-MMD penalty with only O( nD) local cost and D-dimensional communication per round. Across five benchmarks (Reddit, OGB-Arxiv, Ethereum, Tox21, PubMed), FedMMD-P delivers up to 6 percentage-point accuracy gains over non-private baselines under severe domain shifts while incurring <0.1% communication overhead and only 5–10% runtime overhead. We derive end-to-end (ε, δ)-DP bounds via Moments Accountant and Rényi composition, demonstrating practical privacy guarantees.
Person re-identification can be viewed as a cross-camera retrieval problem, with the assumption that retrieved pedestrians exist in the gallery. However, in actual surveillance scenarios, pedestrians don’t always appear in the camera, and the pedestrians’ temporal and spatial trajectory need be recovered in order to track the pedestrian in the camera network. Motivated by this, we propose a new problem for person re-identification known as “trajectory recovery,” which reconstructs the trajectory of pedestrians in the camera network. Existing datasets lacked the necessary annotated data to address this issue at the same time, so we annotated two new datasets, Market-1501-TR and MSMT17-TR, with each pedestrian’s trajectory annotation. Light variation was also done on both datasets to investigate the effect of light variation on trajectory recovery. Then, we present a multipoint recovery benchmark method that takes into account the relationship between all the query pedestrians and results in an efficient trajectory recovery. To evaluate our method in a new setting, we introduce the degree of trajectory point recovery and the degree of trajectory segment recovery, two new evaluation metrics directly informed by practical surveillance system-level considerations. Experiments on two new datasets demonstrate that our method achieves promising performance and outperforms state-of-the-art methods.
Accurate detection of arbitrary-shaped text in natural scenes remains a challenging problem, particularly under real-world noise conditions, such as blur, occlusion, uneven lighting, and cluttered backgrounds. These factors degrade local features and boundary precision, hindering the performance of existing contour-based methods. To address this, we propose Ray Contour Net (RC-Net), a robust framework for boundary adaptation in noisy environments. RC-Net incorporates a noise-aware attention module to obtain noise-robust features and an active contour point initialization module that emits contour points in a polar coordinate system anchored at local text centers. These are refined by a deformation module using lightweight transformers to ensure shape-aware, noise-tolerant adaptation. Additionally, we introduce a tightness evaluation loss to enforce precise boundary predictions. Extensive experiments on six benchmark datasets, including Total-Text, CTW-1500, MSRA-TD500, ICDAR 2015, ICDAR2019 ArT, and ICDAR2019 MLT, demonstrate consistent F-measure improvements of 2% to 3% over strong baselines, validating the effectiveness and robustness of our method.