Forecasting the long-horizon evolution of mechanical systems from position-only observations is a pivotal yet difficult task, as hidden velocities and trajectory-specific physical properties must be inferred simultaneously. Although physics-guided neural networks like Lagrangian Neural Networks (LNNs) guarantee physical plausibility, they generally require complete state inputs and lack adaptability to changing system parameters. To break these limitations, we introduce History-informed Lagrangian Neural Networks (HiLNN). Grounded in the insight that temporal position sequences implicitly encode underlying dynamics, HiLNN employs a recurrent encoder to extract a latent context from history. This context not only reconstructs the unobserved initial velocity but also adaptively modulates the mass matrix, potential energy, and damping coefficients of a structured Lagrangian system. By leveraging a differentiable RK4 rollout scheme, the entire pipeline is optimized end-to-end under multi-step trajectory supervision and energy-consistency regularization. Empirical evaluations across conservative, dissipative, and heterogeneous variable-parameter systems show that HiLNN delivers superior long-term prediction accuracy and maintains precise energy profiles compared to state-of-the-art baselines. The source code is publicly available at https://github.com/yingtian22/History-informed-LNN.
3D Gaussian Splatting (3DGS) has achieved impressive results in 3D reconstruction, yet its performance in large-scale scenes is limited by geometric inconsistency, high storage cost, and low rendering efficiency. Existing methods often lack effective supervision for empty regions and fail to distinguish the importance of primitives, leading to floating artifacts and redundant Gaussians. To address these limitations, we propose PVGC-GSP, a unified framework that combines pseudo-view geometric constraints, global saliency pruning, and hierarchical representation for large-scale 3D Gaussian reconstruction. Pseudo-views are generated by perturbing training cameras, and an image warping loss is introduced to improve spatial consistency in a self-supervised manner. A global saliency scoring mechanism is further designed to measure primitive importance from multiview contribution and spatial volume, enabling accurate pruning of redundant Gaussians while maintaining visual quality. Finally, a BVH-based hierarchical Gaussian structure is built for visually continuous level-of-detail rendering. Experiments on multiple large-scale datasets show that our method achieves better reconstruction quality, lower storage overhead, and higher rendering efficiency.
Accurate footprint image segmentation is challenging in forensic applications because fine anatomical structures, weak boundaries, and background interference can degrade segmentation performance. This study presents a task-oriented generative adversarial network (GAN)-based framework for forensic footprint image segmentation. Channel Prior Convolutional Attention (CPCA) modules are integrated into the decoder stages of the generator to recalibrate fused encoder–decoder features and preserve fine details in the toe, arch, and heel regions. In addition, a dual-branch discriminator processes image–mask pairs at the original and downsampled scales, providing complementary constraints on local boundary details and global footprint morphology. The framework is trained with a least-squares adversarial loss and a binary cross-entropy (BCE)–Dice segmentation loss. Experiments on the self-collected aFoot_2025 dataset show that the proposed framework achieves an IoU of 0.9448 and a Dice coefficient of 0.9713, outperforming the evaluated baseline and attention-based alternatives. Under the evaluated synthetic Gaussian-noise settings, the proposed method retained relatively stable segmentation performance. Furthermore, an exploratory footprint-based height-prediction analysis showed modestly lower prediction errors than the baseline GAN. These findings indicate that, under the controlled acquisition conditions of the aFoot_2025 dataset, CPCA-based feature calibration and dual-scale discrimination may improve segmentation-mask quality and provide a possible benefit for subsequent anthropometric analysis.
In Intelligent Transportation Systems (ITS), tiny machine learning enables resource-efficient on-device learning and deployment of AI models for critical tasks such as traffic flow prediction and pedestrian safety. However, existing object counting models typically consist of a large number of parameters, which require high computational costs that are unsuited for on-device learning. Moreover, these models are usually trained in a centralized manner, which raises data privacy concerns and relies on the unrealistic assumption of independent and identically distributed (IID) data across all vehicles and environments. To address these challenges, we propose a lightweight network, termed Privacy-aware Efficient Generalization Network (PEGNet), that preserves user privacy and improves cross-scenario generalization. Specifically, this network incorporates a Lightweight Adaptive Enhancement (LAE) module with deformable convolutions to efficiently capture multi-scale features under strict resource constraints. We further employ a federated learning framework to collaboratively train the model across distributed IoV nodes with limited local data, thereby significantly reducing per-node computation while safeguarding user data privacy. To enhance the robustness of non-IID data, we introduce a Moment Distance loss (MDloss) that prompts the model to adapt to diverse traffic scenes. Experiments on multiple vehicle and crowd counting benchmarks demonstrate that the proposed method achieves superior efficiency and accuracy.
In the realm of Intelligent Transportation Systems, tiny machine learning facilitates the deployment and training of AI models directly on devices, enabling efficient execution of crucial tasks such as predicting traffic patterns and ensuring pedestrian safety. Nonetheless, most existing object counting approaches rely on parameter-heavy architectures, resulting in significant computational demands that hinder their suitability for on-device applications. To tackle this issue, we present a lightweight network, named Efficient Scale Recognition Network (ESRNet), which enhances counting precision. Our design integrates a Lightweight Multi-scale Enhancement (LME) module utilizing deformable convolutions, which allows effective extraction of multi-scale features within tight computational budgets. To further improve adaptability to non-i.i.d. data, we propose a Statistical Moment Alignment loss (SMAloss), which promotes generalization across various traffic conditions. Experimental results across multiple benchmarks for vehicle counting verify that the proposed method delivers superior performance in both efficiency and accuracy.
Neural Radiance Fields (NeRF) have shown strong potential in 3D reconstruction from multi-view images, achieving high visual fidelity and geometric accuracy. However, NeRF’s reliance on dense multi-view inputs limits its practical applicability in sparse scenarios. This paper introduces the Sparse Feature Enhancement Module (SFEM) to enhance NeRF performance under sparse-view conditions. SFEM leverages cross-view feature interactions to capture long-range dependencies and global context, addressing the limitations of local feature representations in sparse data. We validate our method on the LLFF dataset, demonstrating consistent improvements in both geometric accuracy and texture quality. Our results show that SFEM-NeRF reduces reconstruction errors and enhances texture details, even with only 6 input views. Future work will focus on refining geometric reconstruction strategies for extreme sparsity and optimizing feature aggregation efficiency to further reduce computational overhead.
Data steganography aims to conceal information within visual content, yet existing spatial- and frequency-domain approaches suffer from trade-offs between security, capacity, and perceptual quality. Recent advances in generative models, particularly diffusion models, offer new avenues for adaptive image synthesis, but integrating precise information embedding into the generative process remains challenging. We introduce Shackled Dancing Diffusion, or SD^2, a plug-and-play generative steganography method that combines bit-position locking with diffusion sampling injection to enable controllable information embedding within the generative trajectory. SD^2 leverages the expressive power of diffusion models to synthesize diverse carrier images while maintaining full message recovery with 100% accuracy. Our method achieves a favorable balance between randomness and constraint, enhancing robustness against steganalysis without compromising image fidelity. Extensive experiments show that SD^2 substantially outperforms prior methods in security, embedding capacity, and stability. This algorithm offers new insights into controllable generation and opens promising directions for secure visual communication.
Despite significant advancements in 2D human image generation, existing methodologies remain constrained by geometric ambiguities in single-view pose guidance and persistent challenges in maintaining cross-pose identity consistency. While 3D-aware approaches offer a promising solution to depth-related deficiencies inherent in 2D methods, they often necessitate critical trade-offs between resolution and computational efficiency. Alternative techniques employing the Skinned Multi-Person Linear (SMPL) model as a structural prior to reduce computational overhead and enhance pose flexibility; however, this generic parametric framework struggles to capture personalized features and motion dynamics. To address these limitations, we introduce a compositional triplane representation for animatable 3D human synthesis, which integrates SMPL as a foundational structural prior while incorporating style-based blending weight fields and pose-driven deformation fields to improve identity preservation and dynamic fidelity. Comprehensive experiments conducted across benchmark models—StyleSDF, EANRF-GAN, and EVA3D—demonstrate that our method achieves great performance, yielding the highest PCKh@0.5 score of 91.09 and the second lowest Fréchet Inception Distance (FID) of 68.50, thereby validating its efficacy in overcoming prior constraints for high-fidelity human synthesis.
Drowning is a leading cause of accidental fatalities worldwide, making timely intervention critical. This study analyzes underwater drowning detection using Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) with transfer learning to improve accuracy. We evaluate five CNN architectures—MobileNetV3, EfficientNet, ResNet, AlexNet, and DenseNet121—alongside ViTs to assess their effectiveness in detecting drowning incidents from underwater video footage. The models are trained and tested on a curated dataset with varying underwater conditions, including low visibility and lighting fluctuations. Performance is measured using accuracy, precision, recall, F1-score, and inference speed to identify the best model for real-time drowning detection. The results show that EfficientNet achieves the highest detection accuracy of 97% for swimming and 98% for drowning detection, outperforming other CNN models with a real-time inference time of 0.04 seconds. In contrast, ViTs demonstrate strong feature extraction capabilities but require higher computational resources with inference time of 0.12 seconds. Additionally, transfer learning significantly improves model generalization, reducing false alarms and enhancing response efficiency. This study highlights the potential of deep learning-based approaches for automated underwater drowning detection, providing a reliable solution for surveillance systems and rescue operations. Future work will focus on optimizing Vision Transformers (ViTs) for real-time deployment and integrating them with IoT-based alert systems to enhance the responsiveness and effectiveness of drowning detection solutions.
The prerequisite for reproducing a high-precision scene using Neural Radiation Fields (NeRF) is the availability of accurate pose of the input images. Existing models are mainly designed to optimize the camera poses of each input image independently. However, the relative poses between images are not sufficiently considered, and in the case of violent camera motion, the estimated pose may have a large drift. Aiming at the problems of using NeRF for pose estimation, we propose a NeRF pose estimation model based on depth supervision, which is provided for NeRF through monocular depth estimation based on deep learning. Our model constrains the shape radiation ambiguity so as to jointly optimize the poses and NeRF to achieve accurate estimation of the poses. Experimental results on the Tanks and Temples dataset and the LLFF dataset demonstrate that our model can handle challenging trajectories and outperforms other methods in terms of pose estimation accuracy.
The majority of existing counting models are designed to operate on a singular object category, such as crowds or vehicles. The emergence of multi-modal foundational models, e.g ., Contrastive Language-Image Pre-training (CLIP), has paved the way for class-agnostic counting. This approach facilitates the counting of objects across diverse classes within a single image based on textual indications. However, class-agnostic counting models based on CLIP confront two primary challenges. Firstly, the CLIP model exhibits limited sensitivity towards location information, which prioritizes global content over the precise localization of objects. Therefore, directly employing the CLIP model is regarded as suboptimal. Secondly, these models commonly employ frozen pre-trained vision and language encoders while disregarding potential misalignment within the constructed hypothesis space. In this paper, we propose a unified framework, named the Vision-Language Prior Guidance (VLPG) Network, to tackle these two challenges. The VLPG consists of three key components, namely the Grounding DINO module, Spatial Prior Calibration (SPC) module, and Object-Centric Alignment (OCA) module. The Grounding DINO module utilizes the spatial-awareness capability of extensive pre-trained object grounding models to incorporate the spatial position as an additional prior for a particular query class. This adaptation enables the network to concentrate more precisely on the exact location of the objects. Meanwhile, the SPC module is built to extract the long-range dependencies and local regions of the spatial position. Additionally, to align the feature space across different modalities, we design an OCA module that condenses textual information into an object query which serves as an instruction for cross-modality matching. Through the collaborative efforts of these three modules, multimodal representations are aligned while maintaining their discriminative nature. Comprehensive experiments conducted on various benchmarks validate the effectiveness of the proposed model.
This paper presents a 3D motion prediction framework based on multimodal sensing and deep learning to address the key challenge of dynamic pose modeling in skiing motion analysis. The multidimensional motion data of professional skiers are collected by deploying Noitom's highprecision inertial measurement units (IMUs), and accurate 2D pose estimation is achieved through the integration of AlphaPose. High-fidelity reconstruction from 2D joint coordinates to 3D kinematic parameters is accomplished using the MotionBERT algorithm, effectively capturing the trajectory features of athletes' movements in three-dimensional space. Furthermore, the HumanMAC architecture is introduced to construct a motion prediction model that forecasts the subsequent 2000 ms of motion patterns by analyzing a 500 ms historical motion sequence. For the first time, the synergistic effect of 3D pose estimation and spatiotemporal attention mechanisms in winter sports analysis is validated, providing anew technical paradigm for the development of intelligent skiing training Systems.
With the development of Artificial Intelligence Generated Content (AIGC) technology, deepfakes present a considerable threat to personal privacy and property security. Current deepfake detection methods are constrained by static feature extraction and limited multi-scale analysis and fail to localize subtle sequential artifacts. To detect the sequential deepfake, a dual-modality framework termed SATNet (Spatial Adaptive Transformer Network) is proposed. It consists of two main modules, namely Spatial Adaptation (SA) module and Global-Local Attention (GLA) module. The SA module is built to extract multi-scale feature and GLA module is designed to capture the hierarchical spatiotemporal dependencies. Experimental results demonstrate the effectiveness of SATNet in decoding sequences for adaptive multi-scale reasoning.
Existing gait recognition methods are capable of extracting rich spatial gait information but often overlook fine-grained temporal features within local regions and temporal contextual information across different sub-regions. Considering gait recognition as a fine-grained recognition task and each individual exhibits uniqueness in their movements across different temporal sequences, we propose a local multi-scale and global contextual spatio-temporal (LMGCS) network for gait recognition. It divides the whole gait sequence into sub-sequences with multiple spatio resolutions and extracts multi-scale temporal features. We extract the temporal context information of different sub-sequences with the transformer, and all sub-sequences are fused to form global features. Furthermore, the loss function that combines the triplet loss function and cross-entropy loss function is utilized to prompt the proposed model to fulfill the gait recognition. The proposed method achieved state-of-the-art results on two popular public datasets. It achieved rank-1 accuracy of 98.0%, 95.4%, and 85.0% on the three walk states of the CASIA-B dataset and 90.9% on the OU-MVLP dataset.
The majority of existing counting models are designed to operate on a singular object category, for instance, pedestrians or vehicles. The rise of multi-modality base networks, such as CLIP (Contrastive Language-Image Pre-training), has paved the way for object counting. This approach facilitates the counting of objects across diverse classeswithin one image by leveraging textual prompts. However, some counting models based on CLIP confront a primary challenge. Specially, the CLIP model exhibits limited sensitivity towards location information, which prioritizes global content over the precise localization of objects. Consequently, directly employing the CLIP model is regarded as suboptimal. To address the sensitivity issue of object position in multimodal models based on CLIP in object counting, we propose a spatial prior guidance (SPG) network in this paper. The spatial prior guidance module attempts to extract spatial positions for guiding the vision encoder. Thus, it concentrates on the specific object regions. Experimental results on FSC-147 and ShanghaiTech benchmark datasets demonstrate that our approach outperforms other methods.
We present SemiOccam, an image recognition network that leverages semi-supervised learning in a highly efficient manner. Existing works often rely on complex training techniques and architectures, requiring hundreds of GPU hours for training, while their generalization ability with extremely limited labeled data remains to be improved. To address these limitations, we construct a hierarchical mixture density classification mechanism by optimizing mutual information between feature representations and target classes, compressing redundant information while retaining crucial discriminative components. Experimental results demonstrate that our method achieves state-of-the-art performance on three commonly used datasets, with accuracy exceeding 95
Fluid-structure interaction is common in engineering and natural systems, where floating-body motion is governed by added mass, drag, and background flows. Modeling these dissipative dynamics is difficult: black-box neural models regress state derivatives with limited interpretability and unstable long-horizon predictions. We propose Floating-Body Hydrodynamic Neural Networks (FHNN), a physics-structured framework that predicts interpretable hydrodynamic parameters such as directional added masses, drag coefficients, and a streamfunction-based flow, and couples them with analytic equations of motion. This design constrains the hypothesis space, enhances interpretability, and stabilizes integration. On synthetic vortex datasets, FHNN achieves up to an order-of-magnitude lower error than Neural ODEs, recovers physically consistent flow fields. Compared with Hamiltonian and Lagrangian neural networks, FHNN more effectively handles dissipative dynamics while preserving interpretability, which bridges the gap between black-box learning and transparent system identification.
Aiming at the problem of gait recognition, a multi- view gait recognition method based on the human pose estimation model is proposed. This method can effectively solve the problem of low gait recognition rates caused by limited viewpoints. The VIBE method is used to extract human pose parameters of video frames, but there are errors in human pose parameters estimated by VIBE, and direct rotation with attitude parameters will inevitably cause greater errors. The attitude parameters were corrected in the design experiment. Finally, we used the Rodrigues rotation matrix to generate human pose parameters from other angles for VIBE to complete the conversion of character angles. The calibration network, including the attitude average and angle correction model, is designed. The attitude parameters under are input as the mean value as our final attitude parameters and then sent into the angle correction model to correct the root node. Finally, the corrected gait sequence is obtained. We can intuitively observe that the human model has obvious effects, and the accuracy of gait recognition is verified by Gaitset, which verifies that the model has good qualitative and quantitative effects. After correcting the correction network, we can generate gait sequences from other perspectives through the existing perspective. This method can well expand the perspective of the database and obtain more accurate gait sequence models from other perspectives.
Crowd counting has substantial practical applications in various consumer-oriented areas, particularly for safety assessments and marketing strategies. However, considering the complexities of the capturing conditions, the unavoidable background interference possesses the potential to disrupt the effectiveness of established counting methods, and it further poses degraded counting performance. To address this challenge, we propose a Region-Aware Quantum Network (RAQNet) by attentively learning from the crowd region. It consists of four key components, namely the feature extractor, the object region awareness module (ORA), the quantum-driven calibration (QDC) module, and the decoder module. The cascaded ORA modules are engineered for the extraction of local information, which addresses background interference. Additionally, two QDC modules are incorporated to capture global information, which utilizes quantum states to calibrate features. Extensive experimental results conducted on four crowd benchmark datasets and three cross-domain datasets prove that the RAQNet outperforms the state-of-the-art competitors, both subjectively and objectively.
Multi-ship tracking (MST) as a core technology has been proven to be applied to situational awareness at sea and the development of a navigational system for autonomous ships. Despite impressive tracking outcomes achieved by multi-object tracking (MOT) algorithms for pedestrian and vehicle datasets, these models and techniques exhibit poor performance when applied to ship datasets. Intersection of Union (IoU) is the most popular metric for computing similarity used in object tracking. The low frame rates and severe image shake caused by wave turbulence in ship datasets often result in minimal, or even zero, Intersection of Union (IoU) between the predicted and detected bounding boxes. This issue contributes to frequent identity switches of tracked objects, undermining the tracking performance. In this paper, we address the weaknesses of IoU by incorporating the smallest convex shapes that enclose both the predicted and detected bounding boxes. The calculation of the tracking version of IoU (TIoU) metric considers not only the size of the overlapping area between the detection bounding box and the prediction box, but also the similarity of their shapes. Through the integration of the TIoU into state-of-the-art object tracking frameworks, such as DeepSort and ByteTrack, we consistently achieve improvements in the tracking performance of these frameworks.