
Abstract Addressing the inherent low stiffness of flexible manipulators, existing control schemes often face an intrinsic contradiction where rapid convergence leads to increased vibration amplitudes, making it challenging to achieve high-precision trajectory tracking while effectively suppressing elastic vibrations. To tackle this issue, this paper proposes a neural network-based Fixed-Time learning control strategy. This strategy is capable of simultaneously handling output constraints, model uncertainties, and input dead-zone nonlinearities of the system. The designed controller effectively compensates for the adverse effects of the input deadzone, ensuring that all system states converge to a small neighborhood around the origin within a fixed time, thereby significantly improving the system’s convergence speed and transient performance. By introducing a logarithmic Barrier Lyapunov Function (BLF), the prescribed tracking error constraints are strictly guaranteed. Furthermore, high-frequency chattering is mitigated through a smooth approximation of the sign function. Experimental results demonstrate that, compared with the PSF controller, the proposed Fixed-Time control scheme reduces the steady-state tracking errors by 61.9% and 69.2%, respectively. In terms of vibration suppression, the steady-state values of elastic vibrations are reduced by 49.2% and 32.6%, respectively. These results fully validate the superiority and robustness of the proposed control strategy in balancing rapid convergence with vibration suppression.
Abstract Underwater small object detection is one of the emerging core technologies at the intersection of computer vision and underwater imaging detection, which aims to achieve efficient and accurate detection and identification of faint, small targets in underwater environments, and has been widely applied in fields of aquatic organism detection, mineral resource exploration and other fields. Among them, underwater object detection methods based on optical images, with their high resolution and good flexibility, have demonstrated significant advantages in underwater biological detection, recognition and other scenarios. However, there is currently a relative lack of comprehensive reviews on existing research regarding underwater small object detection. To address this gap, this paper presents a systematic review of small object detection in underwater optical images. Firstly, at the level of problems and data, the specific problems faced by underwater small detection are comprehensively sorted out and classified into two categories: the inherent characteristics of underwater small objects, and the degradation characteristics of underwater images. Then 16 existing typical datasets are sorted out and analyzed in detail. Subsequently, at the methodological level, existing underwater small object detection algorithms are systematically summarized from the perspectives of traditional machine learning and deep learning, and an in-depth analysis of improved deep learning-based strategies is conducted by covering five aspects: data augmentation and pre-processing, multi-scale feature fusion, attention mechanism introduction, loss function optimization, and lightweight design. Finally, the challenges faced by underwater small object detection at present are pointed out, and the potential directions for future development are explored, aiming to provide valuable references for subsequent related research.
Abstract Decentralized federated learning (DFL) allows a group of distributed users to train a shared model without relying on a central server. Most existing decentralized aggregation methods, however, still treat the communication network as a flat Euclidean space and overlook the geometric differences that may arise in the model parameter landscape. To better capture these differences, we develop a curvature-aware decentralized federated learning framework, termed CA-DFL, which introduces geometric structure into the aggregation process. The main idea is to view the communication graph from a discrete Riemannian perspective. According to the divergence between neighboring model parameters, each communication edge is associated with one of three constant-curvature geometries: the sphere S2, the Euclidean plane R2, or the pseudosphere P S2. This classification allows model updates to be transferred between tangent spaces through closed-form parallel transport operators, so that the aggregation step respects the local geometry rather than mixing parameters in a purely Euclidean manner. Building on this, we further design a tensor-field-based geometric weighting scheme to assign adaptive aggregation weights using manifold inner products. This weighting scheme only borrows the idea of adaptive neighbor weighting, while the core operation remains curvature-aware transport and aggregation on manifolds. A subtree partition scheme is also introduced to encourage parameter sharing while reducing communication overhead. We establish convergence guarantees for the proposed method under standard assumptions. Experiments on MNIST and CIFAR-10 under ring, exponential, and random communication topologies demonstrate that CA-DFL achieves accuracy competitive with standard decentralized methods while providing adaptive geometric aggregation. A curvature-threshold sensitivity analysis confirms the robustness of the framework to hyperparameter choice.
Electroencephalograph (EEG)-based vigilance estimation methods have achieved significant progress. In vigilance-associated EEG signals, non-adjacent electrodes exhibit strong coupling, and distant time points also demonstrate significant dependence, not limited the adjacent ones. Therefore, fully extracting such rich globalu2212local spatiotemporal characteristics is critical. In this paper, we propose the global enhanced LSTM Network (GE-LSTM-Net), in which synergizes the Transformeru2019s attention mechanism with LSTM to enhance the extraction of spatiotemporal features. Firstly, a specialized sample partitioning strategy along with the designed feature fusion module is adopted to reorganize raw EEG signals into structured 3D differential entropy (DE) feature representations, effectively preserving spatiotemporal and frequency dependencies across electrode channels and time points. Secondly, the attention mechanism and LSTM are encapsulated into a novel module (GE-LSTM module), serving as the core of the proposed GE-LSTM-Net to simultaneously extract spatiotemporal features from 3D representations. In this module, the attention mechanism will extract global information and integrate it into each unit of the LSTM, enabling LSTM to focus on more critical electrode channels and time points and extract richer globalu2212local features. Subsequently, the GE-LSTM-Net demonstrate competitive performance and achieved SOTA results compared to existing methods on two public vigilance datasets. The codes are available at: https://github.com/Lanhao23-nudt/GE-LSTM-Net.
Large language models (LLMs) have demonstrated remarkable capabilities, yet frequently encounter issues with hallucinations. Retrieval-augmented generation (RAG) mitigates this problem by integrating external knowledge. However, current RAG approaches face significant limitations, such as redundant tokens that dilute semantic focus of LLMs and suboptimal prompt ordering. To address these challenges, this paper introduce a fine-grained framework called FG-RAG, with two key components: (1) refined prompt generation. Standalone propositions are firstly extracted from raw documents and dynamically organized into a semantic graph to capture their interrelationships. After retrieving relevant subgraphs from this graph, a directional diffusion model (DDM) is adapted to iteratively refine these graph representations, which are subsequently transformed into soft prompts compatible with LLMs; (2) prompt ordering. With these soft prompts, this work formulates prompt ordering as a markov decision process (MDP) and optimize it through reinforcement learning (RL). By leveraging reward derived from prompts performance, the RL agent learns to order prompts to maximize the accuracy of the LLMsu2019 output, ensuring optimal utilization of the generated prompts. Extensive evaluations of question-answering benchmarks demonstrate that FG-RAG outperforms state-of-the-art RAG methods, with 1.3% EM and 1.5% F1 score improvements.
Complex prediction models widely adopted in the field of financial risk control, despite their superior performance in credit default prediction accuracy, have long suffered from issues regarding the explainability of their predictive results. This makes it difficult to satisfy stringent requirements for regulatory compliance and algorithmic fairness. Current mainstream explainability techniques primarily focus on feature attribution and generally lack structured modeling and deep reasoning capabilities for complex community correlations between financial entities, thereby failing to reveal the transmission mechanisms of community-based risks. To this end, this paper proposes an explainability-enhanced framework, namely GraphCredit. Specifically, GraphCredit first extracts and quantifies the community risks of borrowers to construct a borrower-centric knowledge graph. Subsequently, GraphCredit employs a GraphSAGE model combined with a gating mechanism to achieve dynamic feature weighting and credit default prediction, achieving an average performance improvement of 8.70% compared with the SOTA models. Finally, leveraging large language models (LLMs), the framework transforms the extracted complex risk evidence chains into logically clear natural language narrative reports that comply with regulatory standards. Experimental results demonstrate that the explainability score under both human and LLM evaluations increased by 17.98% compared with shapley additive explanations (SHAP). GraphCredit elevates the explainability of credit default prediction from traditional static u201Cfeature attributionu201D to a dynamic u201Crisk narrativeu201D dimension, providing a new paradigm that balances high precision with robust trustworthiness for humanu2212AI collaboration in high-risk financial decision-making.
Image field-of-view (FOV) enhancement is a promising technology that expands image visual scope and improves image visual effects, which has become a core research topic in the fields of image processing and computer vision. This technology has gained significant attention in both natural and underwater scenarios, demonstrating great application value across various domains. This article surveys a comprehensive overview of image FOV enhancement, and these technologies are primarily classified via two avenues: fisheye image unwarping and image stitching, covering relevant approaches, system architectures, and future development directions in both natural and underwater scenarios. We summarize these existing researches and conduct an in-depth analysis of various methodologies, providing comprehensive and valuable references for future researchers. To distinguish from previous reviews, we provide a unified cross-domain comparison between natural and underwater scenarios, introduce a consistent classification framework that integrates both traditional and learning-based methods, and summarize system-level design insights that are significantly ignored in previous surveys. In addition, we analyze the critical challenges in image FOV enhancement, and explore potential directions for its future development.
Nowadays, deep learning has demonstrated impressive performance in the area of computer vision and pattern recognition, such as objects recognition, videos classification and image segmentation. In particular, convolutional neural networks (CNNs) an advanced deep-learning technique have achieved strong performance in image and video analysis owing to their powerful feature-extraction capabilities.Based on that, current research also make a breakthrough in image and video recognition via aggregating features extracted from various layers of CNNs. Over the last decade, feature fusion strategies have developed from conventional schemes restricted to conditions we need to control over advanced ones which can be achieved with a large number of images or videos. However, due to the rapid development in this field, it is challenging to track and systematically analyze recent advancements. This has inspired us to offer a comprehensive survey of the significant steps taken towards feature aggregation strategies. To better organize and facilitate understanding, we first introduce preliminary knowledge on deep learning. Then this paper focuses on categorizing and reviewing the current strategies from two main aspects: feature aggregation on images and feature aggregation on videos. We further highlight the comparative strengths and limitations among these strategies. Finally, we point out the challenges in this area and motivate further work via proposing future directions on feature aggregation techniques for both images and videos.
This paper investigates a platform supply chain consisting of a manufacturer and a platform and explores the impact of the manufactureru2019s altruistic preference on the platformu2019s decision to introduce blockchain technology in reselling and agency models. The results indicate that, firstly, when the investment cost of blockchain technology is low, both the manufacturer and the platform prefer to use blockchain in the reselling model, and when the investment cost and the commission are high, both parties opt for an agency sales model. Then, the reselling model is a better option when consumer traceability demand is high, but for the additional unit information cost expended on products, the platform will not introduce blockchain technology in the agency model when the demand weakening effect from this unit cost exceeds 5.7% of the total demand. Finally, considering the manufactureru2019s altruistic preference behavior, it is not feasible to simultaneously boost the revenue of both parties under either sales model. However, in the agency model, the manufactureru2019s altruistic preference can incentivize the platform to introduce blockchain technology.
Efficient analysis and processing of dental images are crucial for dentists to achieve accurate diagnosis and optimal treatment planning. However, dental imaging inherently poses several challenges, such as low contrast, metallic artifacts, and variations in projection angles. Combined with the subjectivity arising from differences in clinicians' expertise, manual interpretation often proves time-consuming and prone to inconsistency. Artificial intelligence (AI)-based automated dental image analysis (DIA) offers a promising solution to these issues and has become an integral part of computer-aided dental diagnosis and treatment. Among various AI technologies, deep learning (DL) stands out as the most widely applied and influential approach due to its superior feature extraction and representation capabilities. To comprehensively summarize recent progress in this field, we focus on the two fundamental aspects of DL research-datasets and models. In this paper, we systematically review 260 studies on DL applications in DIA, including 49 papers on publicly available dental datasets and 211 papers on DL-based algorithms. We first introduce the basic concepts of dental imaging and summarize the characteristics and acquisition methods of existing datasets. Then, we present the foundational techniques of DL and categorize relevant models and algorithms according to different DIA tasks, analyzing their network architectures, optimization strategies, training methods, and performance. Furthermore, we summarize commonly used training and evaluation metrics in the DIA domain. Finally, we discuss the current challenges of existing research and outline potential future directions. We hope that this work provides a valuable and systematic reference for researchers in this field. All supplementary materials and detailed comparison tables will be made publicly available on GitHub.
Segment Anything Model (SAM) has demonstrated powerful zero-shot segmentation performance in natural scenes. The recently released Segment Anything Model 2 (SAM2) has further heightened researchers' expectations towards image segmentation capabilities. To evaluate the performance of SAM2 on class-agnostic instance-level segmentation tasks, we adopt different prompt strategies for SAM2 to cope with instance-level tasks for three relevant scenarios: Salient Instance Segmentation (SIS), Camouflaged Instance Segmentation (CIS), and Shadow Instance Detection (SID). In addition, to further explore the effectiveness of SAM2 in segmenting granular object structures, we also conduct detailed tests on the high-resolution Dichotomous Image Segmentation (DIS) benchmark to assess the fine-grained segmentation capability. Qualitative and quantitative experimental results indicate that the performance of SAM2 varies significantly across different scenarios. Besides, SAM2 is not particularly sensitive to segmenting high-resolution fine details. We hope this technique report can drive the emergence of SAM2-based adapters, aiming to enhance the performance ceiling of large vision models on class-agnostic instance segmentation tasks.
Recurrent neural networks (RNNs) have been employed extensively as intelligent control approaches across various industrial control fields. However, existing research often lacks sufficient focus on discrete-form time-variant problems and disturbance rejection capability. This paper proposes a novel discrete-form integral-reinforcing RNN (DF-IR-RNN) approach. This approach integrates an innovative integral-reinforcing RNN (IR-RNN) design thought into the RNN approach to enhance the disturbance rejection capability in controlling the Stewart platform under discrete-form time-variant environment. Compared to traditional approaches, the proposed approach overcomes their limitations of disturbance rejection. The experimental results demonstrate that the proposed approach is highly effective in disturbance rejection and accurate trajectory tracking.
Underwater light field (ULF) imaging has emerged as a promising technology for capturing a more comprehensive array of visual information from real-world underwater environments. Unlike conventional underwater photography, which captures a two-dimensional projection of light that integrates over the angular domain, ULF imaging collects radiance from multiple directions. This enables the recovery of angular information that is typically lost in traditional imaging methods. Although ULF presents high-dimensional challenges such as inaccurate modeling and complex feature extraction, its ability to represent underwater visual data enhances the understanding of marine scenes. This, in turn, significantly improves the performance of various underwater vision tasks. The field of ULF imaging has garnered increasing attention in both the computer vision and computer graphics communities. This paper presents a comprehensive overview of the research conducted in this area over the past two decades. We focus on various aspects of ULF imaging, including ULF models, theory progression, parameter calibration, external and internal influencing factors, underwater scattering and refraction removal, underwater image enhancement/restoration, expansion of underwater imaging distance, underwater object detection, and underwater 3D reconstruction. Additionally, we analyze the current challenges facing ULF imaging technology and explore potential directions for its future development.
Using computer vision technology to detect prohibited items in X-ray images is an effective method for realizing intelligent security checking. Dual-view security checking can capture the images from both vertical and horizontal perspectives of the same package at the same time, addressing issues such as unfavorable imaging angles and object occlusion at the image acquisition end. In this paper, we proposed a novel prohibited item detection model based on Transformer architecture in dual-view X-ray images. Two feature fusion module, named as feature selection module and corss-attention fusion module, are introduced to make interaction and enhancement. To improve the model inference efficiency, we use MobileViT as the backbone network to reduce the model size. Simulation results based on Dualray dataset has demonstrated the performance of the proposed model.
Person re-identification (ReID) is a desirable yet challenging issue in computer vision, which has numerous potential applications in the public security area. In this paper, the information bottleneck (IB) theory is introduced to filter out the undiscriminating information and noise, and keep the distinctive information for persons, to identify the assigned person from the candidates. To meet the requirement of the IB theory, a dual-stream IB-enhanced (DSIB) ReID framework is proposed to make unentangled estimations of the expectation and variance of the feature map. Moreover, the unconditional feature distribution is relaxed from standard normal distribution to Gaussian distribution, to enrich the information contained in each feature channel. Comprehensive experiments are conducted with several baseline models on datasets MARS, LS-VID, iLiDS-VID, and PRID-2011. The results reveal that our proposed DSIB framework has improved the performance of all baseline models, and it outperforms most state-of-the-art methods in terms of mean average precision (mAP) and Rank-1 precision.
Segmenting greenhouse gases from hyperspectral images can provide detailed information regarding their spatial distribution, which is significant for the monitoring of greenhouse gases. However, accurate segmentation of greenhouse gases is a challenging task due to two main reasons: (1) Diversity: greenhouse gases vary in concentration, size, and texture; (2) Camouflage: the boundaries between greenhouse gases and the surrounding background are blurred. Existing methods primarily focus on designing new modules to address the above challenges, often neglecting the design of the upsampling method within the model, which is crucial for achieving accurate segmentation. In this work, we propose Gas-Aware Upsampling (GasUpper), a novel and efficient upsampling method tailored for greenhouse gas segmentation. Specifically, we first generate a coarse segmentation mask during the upsampling process. Based on the roughly segmented gas and background, we then extract the global features of the gas and combine them with the original features to obtain de-camouflaged feature map that include both the global characteristics of the gas and the local details of the image. This de-camouflaged feature map serves as the foundation for subsequent point sampling. Finally, we utilize the de-camouflaged feature map to generate upsampling coordinate offsets, enabling the model to adaptively adjust the sampling regions based on the content during the sampling process. We conduct comprehensive evaluations by replacing the upsampling method in various segmentation approaches with GasUpper on two hyperspectral datasets. The results indicate that GasUpper consistently and significantly enhances the performance across all segmentation models (0.08%u20139.44% Intersection over Union (IoU), 0.47%u20136.26% Accuracy), outperforming other upsampling methods.
Camouflaged Object Detection (COD) aims to detect objects with camouflaged properties. Although previous studies have focused on natural (animals and insects) and unnatural (artistic and synthetic) camouflage detection, plant camouflage has been neglected. However, plant camouflage plays a vital role in natural camouflage. Therefore, this paper introduces a new challenging problem of Plant Camouflage Detection (PCD). To address this problem, we introduce the PlantCamo dataset, which comprises 1250 images with camouflaged plants representing 58 object categories in various natural scenes. To investigate the current status of plant camouflage detection, we conduct a large-scale benchmark study using 20+ cutting-edge COD models on the proposed dataset. Due to the unique characteristics of plant camouflage, including holes and irregular borders, we develope a new framework, PCNet, dedicated to PCD. Our PCNet surpasses performance thanks to its multi-scale global feature enhancement and refinement. Finally, we discuss the potential applications and insights, hoping this work fills the gap in fine-grained COD research and facilitates further intelligent ecology research. All resources will be available on https://github.com/yjybuaa/PlantCamo.
Quantization is a promising method that reduces memory usage and computational intensity of Deep Neural Networks (DNNs), but it often leads to significant output error that hinder model deployment. In this paper, we propose Bias Compensation (BC) to minimize the output error, thus realizing ultra-low-precision quantization without model fine-tuning. Instead of optimizing the non-convex quantization process as in most previous methods, the proposed BC bypasses the step to directly minimize the quantizing output error by identifying a bias vector for compensation. We have established that the minimization of output error through BC is a convex problem and provides an efficient strategy to procure optimal solutions associated with minimal output error, without the need for training or fine-tuning. We conduct extensive experiments on Vision Transformer Models (ViTs) and Large Language Models (LLMs), and the results show that our method notably reduces quantization output error, thereby permitting ultra-low-precision post-training quantization and enhancing the task performance of models. Especially, BC improves the accuracy of ViT-B* with 4-bit PTQ4ViT by 36.89% on the ImageNet-1K task, and decreases the perplexity of OPT-350M with 3-bit GPTQ by 5.97 on WikiText-2. Our codes are publicly available at https://github.com/GongCheng1919/bias-compensation.
The repeated nature of sponsored search auctions allows the seller to implement Myersonu2019s auction to maximize revenue using past data. But since these data are provided by strategic buyers in the auctions, they can be manipulated, which may hurt the selleru2019s revenue. We model this problem as a Private Data Manipulation (PDM) game: the seller first announces an auction (such as Myersonu2019s) whose allocation and payment rules depend on the value distributions of buyers; the buyers then submit fake value distributions to the seller to implement the auction. The selleru2019s expected revenue and the buyersu2019 expected utilities depend on the auction rule and the game played among the buyers in their choices of the submitted distributions. Under the PDM game, we show that Myersonu2019s auction is equivalent to the generalized first-price auction, and under further assumptions equivalent to the Vickreyu2013Clarkeu2013Groves (VCG) auction and the generalized second-price auction. Our results partially explain why Myersonu2019s auction is not as popular as the generalized second-price auction in the practice of sponsored search auctions, and provide new perspectives into data-driven decision making in mechanism design.
Object navigation, whose goal is to let the agent to reach some places (or objects), has been a popular topic in embodied Artificial Intelligence (AI) researches. However, in our real-world applications, it is more practical to find the targets with particular goals, raising the new requirements of finding the places to achieve the particular functions. In this paper, we define a new task of affordance navigation, whose goal is to find possible places to accomplish the required functions, achieving some particular effects. We first introduce a new dataset for affordance navigation, collected by the proposed affordance algorithm. In order to avoid the high cost of labor, the groundtruth of each episode which is annotated with the interaction data provided by the AI2-THOR simulator. In addition, we also propose an affordance navigation framework, where an Object-to-Manipulation Graph (OMG) is constructed and optimized to emphasize the corresponding nodes (including object nodes and manipulation nodes). Finally, a navigation policy is implemented (trained by reinforcement learning) to guide the navigation to the target places. Experimental results on AI2-THOR simulator illustrate the effectiveness of the proposed approach, which achieves significant gains of 14.0% and 11.7% (on success rate and Success weighted by Path Length (SPL), respectively) over the baseline model.