The proliferation of Internet of Things (IoT) devices has led to the generation of vast amounts of sensitive and decentralized data. Federated learning (FL) offers a promising solution by enabling collaborative model training without centralizing raw data. However, existing methods still face challenges such as reduced model accuracy, low robustness, and potential data privacy leakage. To address these issues, we propose a secure and robust FL framework (SRFL) under a dual-server architecture. In particular, we use two servers to separate the functions of model aggregation and verification, minimizing the threat of single points of failure. We further construct verification models based on perturbed or encrypted client updates to filter malicious updates and defend against poisoning attacks. Meanwhile, we integrate the Cheon-Kim-Kim-Song (CKKS) scheme with random perturbation and ciphertext transformation mechanisms, ensuring that servers cannot access plaintext client updates throughout the training process. Formal security analysis and extensive experiments demonstrate that our scheme consistently achieves high accuracy and robustness under adversarial conditions while preserving data privacy.
Spatiotemporal predictive learning aims to forecast future sequences by modeling autocorrelations within observed frames. Accurate prediction requires both effective modeling of long-term dependencies and preservation of fine spatial details. However, existing memory-based methods often lack explicit modeling of temporal regularities and struggle to effectively integrate hierarchical features. To address these issues, we propose a memory enhancement module augmented ConvLSTM (MEMA-ConvLSTM). The proposed framework introduces two key components. First, a multi-scale autocorrelation memory (AC-Memory) is designed to explicitly aggregate historical states across different temporal scales. By computing spatiotemporal autocorrelation coefficients, it identifies important temporal intervals and retrieves the most relevant historical memory states for current prediction. Second, a memory enhancement module (MEM) is introduced at the top layer of ConvLSTM with a dual-path structure. The global memory path integrates local memory with AC-Memory to enhance temporal modeling, while the detail preservation path fuses hierarchical hidden states through convolution and attention to preserve fine-grained spatial details. Experiments on standard benchmarks demonstrate that MEMA-ConvLSTM achieves competitive performance. In addition, ablation studies and feature map visualizations validate the effectiveness of the proposed method and provide further insight into its working mechanism.
Grayscale conversion plays a crucial role in image processing, particularly for edge detection and segmentation tasks, where decolorization quality directly impacts subsequent analysis. An ideal decolorization algorithm should be both efficient and robust while preserving color consistency and detail contrast. In this study, we revisit the RTCP (Real-time Contrast-Preserving Decolorization) algorithm and propose three key optimizations: a clustering-guided decolorization approach, a locally adaptive decolorization strategy, and a weight-optimized decolorization method. To enhance solution quality, we implement a constrained particle swarm optimization framework to systematically explore the parameter space. Experimental validation on two standard datasets (Cad & iacute;k and CSDD) demonstrates that our optimized methods handle diverse decolorization scenarios more effectively while maintaining competitive performance against existing approaches. Recognizing the limitations of current evaluation metrics in assessing detail contrast preservation, we introduce the D-C2G-SSIM metric for more accurate quantitative assessment. Comparative results show consistent improvements over the original RTCP algorithm, with the average D-C2G-SSIM score increasing from 0.8331 to 0.8442 on Cad & iacute;k dataset and from 0.8696 to 0.8847 on the CSDD dataset, confirming the effectiveness of our approach.
The generation of Chinese calligraphic characters presents a challenging task in computer vision. Under few-shot learning scenarios, the scarcity of calligraphy training data substantially increases the task difficulty. Current generative models commonly exhibit deficiencies, including missing strokes and structural misalignment when generating Chinese handwritten characters. Furthermore, the inherent stylistic diversity and brushwork complexity in Chinese calligraphy further constrain generative model performance. To overcome these challenges, we propose an Error Correction Diffusion Model for Multi-Style Chinese Calligraphy Generation, termed CalliECD. We constrain the generation process by integrating style reference information with stylistic features extracted via high-frequency filters, proposing an iterative skeleton extraction network based on a standard Chinese font library mapping. Besides, multi-level feature fusion enables repeated skeleton feature injection across a multi-scale UNet architecture, facilitating joint content-style learning. Significantly, we introduce a real-time error recognition and correction module that substantially enhances character generation accuracy and consistency. More crucially, we establish the first multi-class Chinese calligraphy dataset, termed CHC. This dataset encompasses 11 distinct script styles from master calligraphers, containing over 50,000 high-quality single-character images. The dataset annotations provide reliable data support for calligraphy generation research. Experimental results show that under few-shot conditions, the CalliECD method significantly outperforms baseline models in both glyph structure accuracy and generalization capability for unseen character generation. The dataset and code are available at https://github.com/LPDLG/CalliECD.
Computer vision techniques have revolutionized digital mural inpainting. However, single-stage networks often yield suboptimal results with blurred textures and structural distortion, while existing progressive strategies struggle to effectively balance local and global information. To address these limitations, we propose a novel generative adversarial model that progressively reconstructs mural details by adaptively integrating multi-scale local features and global context based on damage severity. We first obtain initial coarse results using an encoder-decoder network. Then, a mask-guided network adaptively extracts and fuses local features according to damage levels. Next, multi-level residual learning further refines details at different scales. Finally, a global network captures overall artistic characteristics using an optimized Transformer-UNet architecture. In this way, our method harmonizes detailed local restoration with the preservation of overall artistic integrity throughout the progressive inpainting process. Extensive experiments on multiple mural datasets demonstrate that our method achieves state-of-the-art performance in terms of texture clarity and structural coherence. We release the source code at https://github.com/Kk01Qq/Mural-Inpainting.
Digital inpainting of traditional Chinese murals is challenged by the difficulty of disentangling intricate structures from unique artistic styles, often leading to artifacts. To address this, we propose DCADif, a novel diffusion model for high-fidelity mural restoration. DCADif’s core innovation is a Decoupled Conditional Encoder that uses parallel pathways a pre-trained CLIP for structural line art and a new SwinStyle Encoder for stylistic features to achieve independent control. Furthermore, a Time-Adaptive Feature Fusion (TAFF) module dynamically adjusts the influence of these features during denoising, prioritizing structure in early stages and style in later ones, mimicking an expert’s coarse-to-fine workflow. Evaluated on our new large-scale MuralVerse-S dataset, DCADif significantly outperforms state-of-the-art methods across all degradation levels. It establishes a new benchmark for digital cultural heritage preservation by effectively balancing structural accuracy with artistic authenticity. The dataset and code are publicly available.The dataset and code are available at https://github.com/LPDLG/DCADif .
Point cloud registration is an essential task in 3D computer vision. However, for repetitive patterns and low-geometric region in indoor scenarios, most existing methods cannot distinguish real useful features, and filter out incorrect correspondences effectively. To tackle this challenge, we propose a geometric-image collaborative point cloud registration approch, so-called GIReg. We perform pixel-wise and patch-wise feature extraction (P2FE) of images, and design a superpoint feature representation module to enhance learned 3D point cloud through a two-stage fusion of geometric and image features. Besides, through a geometric and image feature correlation via attention guidance (GIFCAG) module, the global correlation of geometric and image within a point cloud, and the consistency of geometric and image between point cloud are captured. Finally, the distinctive superpoint features are used for registration. Extensive experiments demonstrate that our method achieves superior performance, especially in the geometrically challenging scenarios.
The Dunhuang Mogao Grottoes are among the most significant historical and cultural sites in China, with the Dunhuang Mural serving as key representative artifacts. Due to natural erosion, human activities, and other factors, it is crucial to protect these frescoes authentically and accurately to ensure their preservation. To address this, we propose a technological framework, 3DSynBrush, for high-quality 3D reconstruction of Chinese painting data. First, single elements are extracted from the mural through the Perspective-Driven Synthesis (PDS) Module, and a sparse perspective is generated. The sparse view is then passed through the Neural Rendering Synthesizer to obtain a continuous view image, which ultimately meets the input requirements of the Light Field Fusion Meshing Module to achieve 3D reconstruction. 3DSynBrush aims to reconstruct a realistic model of the scene depicted using only a single primary view of the fresco. Moreover, we develope a high-quality Chinese Mural Elements dataset, termed CME. Compared to current 3D reconstruction algorithms, 3DSynBrush produces visually coherent results while using only 40% of the vertices and triangles, thereby minimizing computational resources.
The preservation of cultural artifacts is vital for maintaining historical continuity, particularly for traditional Chinese paintings that often suffer from decay and damage over time. Existing inpainting methods struggle to simultaneously recover complex brushwork structures, maintain visual coherence, and preserve consistency across multiple resolutions. To address these challenges, we present the Twin Cascade Spatial Multi-scale Attention Filtering (TCSMAF) method, which adopts a symmetric multi-scale dual-branch architecture to capture complex structures and semantic details through parallel processing. A Spatial Kernel Module is proposed to enhance spatial perception by coordinating hierarchical features with spatial coordinate encoding. Moreover, a Multi-scale Spatial and Channel Attention module that adopts progressive convolution kernel sizes is introduced to improve texture reconstruction by leveraging features across different scales and channels. These technical innovations significantly advance digital inpainting methodologies, providing a robust framework specifically designed to handle the intricate textures and details of damaged paintings. The dataset and code are available at https://github.com/LPDLG/TCSMAF .
Generative models are widely used to create synthetic images with pixel-level annotations, reducing the need for extensive image collection and annotation in training semantic segmentation models. However, challenges such as limited image diversity, noisy pseudo labels, and domain gaps between synthetic and real images often undermine their effectiveness in semantic segmentation tasks. To this end, we propose a novel synthetic data-driven framework that integrates high-fidelity data generation with noise-aware iterative self-training, enhancing performance in semantic segmentation. Our framework consists of two parts: First, two synthetic data generation strategies are proposed: Class-Aware Text-to-Image Synthesis (CATS) and Caption-Driven Text-to-Image Synthesis (CDTS), which generate large-scale synthetic data with precise pixel-level annotations. Specifically, CATS leverages class-aware chain generation for diverse synthesis, unsupervised instance segmentation for pixel-level annotations, and feature similarity-based selection to ensure visual-semantic consistency. CDTS employs large language and vision-language models to directly generate semantically rich, pixel-annotated images, complemented by a dual-modality quality assessment framework to guarantee visual and label fidelity. Second, the generated data underpin our iterative self-training (IST), which progressively enhances semantic segmentation models and refines pseudo labels through self-training and our proposed label filtering strategy (LabFilt). LabFilt improves pseudo label quality by applying class-adaptive techniques at both the pixel and object levels. Extensive experiments demonstrate that our proposed frameworks, IST-CATS and IST-CDTS, achieve superior performance, consistently surpassing most existing approaches across multiple benchmarks.
Non-negative matrix factorization (NMF) has proven effective for multi-view clustering. However, most existing methods rely on a single low-dimensional representation, which weakens the modeling of cross-view consensus information and hinders the explicit characterization of view-specific complementary information. To overcome this limitation, we propose information-decoupled dual-branch non-negative matrix factorization (ID-DBNMF). This method introduces an additive factorization mechanism that decomposes the latent representation of each view into a consensus component and a complementary component, while promoting soft orthogonality between them through a disentanglement regularizer. Additionally, ID-DBNMF leverages a dual-branch manifold regularization framework with adaptive view-weight learning to preserve branch-wise geometric structure, while dynamically adjusting view contributions. Experimental results across six benchmark multi-view datasets demonstrate that ID-DBNMF consistently outperforms several representative NMF-based methods, highlighting the effectiveness of information decoupling in multi-view representation learning.
Sketches hold considerable research value for archaeologists as they convey the ancient culture, artistic techniques, and social contexts of murals. However, the widespread presence of deteriorated artifacts and the scarcity of artifact databases make it difficult to train a sketch extraction model. This study develops SemiSketch, a novel semi-supervised framework that extracts clean, coherent sketches from ancient murals. By leveraging a dual-branch learning paradigm, it effectively mitigates challenges such as deterioration artifacts, noise, and data scarcity. SemiSketch innovatively uses a pixel-level reference mechanism as an intermediary "pivot" between deteriorated murals and clean sketches. It decomposes the training process into two branches: an unsupervised branch that transforms murals into sketches, and a supervised branch that refines noiseless line styles through pixel-level correspondences. We introduce a shared CNN-hybrid Vision Transformer generator to integrate the two branches, combining CNN-based transpose self-attention and axial attention to capture local and global information, thereby enhancing the extraction of key lines in murals. Additionally, a gradient frequency compensation module is employed to effectively mitigate noise caused by deterioration artifacts, resulting in more complete and cleaner sketches. Empirical evaluations are conducted on various styles of datasets, including the Fengguo Temple Buddhist frescoes, Dunhuang murals, and Indian murals. Extensive experiments show that SemiSketch substantially outperforms a wide range of baselines and effectively extracts clear and coherent sketches. We release the source code at https://github.com/Alice77bai/SemiSketch.
To accurately estimate full-degree-of-freedom (DOF) model parameters of the 3D point cloud acquisition device, a calibration method by using a calibrator of a simple space ball is proposed. The method can achieve full-DOF (DOF) estimation without additional hardware and step-by-step calculations. Firstly, a measurement model is established according to the rotation characteristics of the 3D point cloud acquisition device. Secondly, with the spherical constraints of the sphere, a nonlinear optimization model of the 3D point cloud acquisition device is established by adopting the bidirectional kernel mean p-power error (BiK) loss function which is robust to the measurement noise and outliers. Finally, the successful history-based differential evolution parameter adaptation (SHADE) algorithm and the Levenberg-Marquardt (LM) algorithm are combined to solve the nonlinear optimization model, and the full-DOF model parameters of the 3D point cloud acquisition device can be estimated. Experimental results demonstrates that the proposed method can accurately estimate the full-DOF model parameters of the 3D point cloud acquisition device. More encouragingly, the effect of measurement noise and outliers can be significantly suppressed.
Automatic sketch colorization is a challenging task in both computer graphics and computer vision since all the color, texture, and shading generation must be created based on the abstract sketch. Moreover, historical degradation factors-including pigment fading and physical erosion-have severely compromised the original chromatic and structural integrity of traditional Chinese paintings, introducing unique challenges for digital restoration and recolorization. Previous methods often yielded unsatisfactory results when dealing with Chinese paintings. To address this issue, we propose a Style-Guided Colorization Generative Adversarial Network for Traditional Chinese landscape paintings, termed SGCGAN. Specifically, we propose a novel AdaCAttN network that can more accurately learn the content and style features of both Grayscale Sketch and Color Painting. Moreover, we conducted hybrid training using both paired and unpaired datasets, utilizing the unpaired data to increase the generalization ability of the model, and using the paired data to guide the learning process of the model. Then, we introduce a Color Histogram Structural Similarity Loss to preserve chromatic distribution patterns during the colorization process. The color style constraint is realized by measuring the difference in color distribution between the source image and the target image. More importantly, we innovatively extract the Grayscale Sketch of traditional Chinese landscape paintings, which not only retains part of the grayscale information but also visually matches the characteristics of white landscape painting. We create a Chinese Landscape Painting dataset with Grayscale Sketch named SketchCLP. Extensive experiments and ablation studies demonstrate the superiority of the proposed method over previous state-of-the-art methods. The code is available at https://github.com/LPDLG/SketchCLP.
Small targets in remote sensing images are typically characterized by small scale, dense distribution, and complex backgrounds. During continuous downsampling and multi-scale feature fusion, key information can be easily lost, posing a significant challenge for target detection. To address the limitations of existing methods in small-target feature representation and context modeling, this paper proposes a Multi-scale Context-Aware Feature Aggregation Network (MCFAD-YOLO) for remote sensing small-target detection. This method improves the discriminative feature representation of small targets by leveraging a multi-scale context-aware mechanism and constructing a dual-path feature aggregation structure. It effectively integrates local detailed features with global semantic information during multi-scale feature transfer, thus mitigating feature degradation caused by downsampling and cross-layer fusion. In addition, a fine-grained context fusion detection head is introduced to improve the localization and recognition capabilities of small targets in complex backgrounds and high-density scenes. Experimental results demonstrate that the proposed strategy achieves superior detection performance compared to mainstream YOLO series models on the SIMD, VisDrone2019, and RSOD datasets. Compared with the baseline model YOLOv11-S, the proposed method achieves improvements of 2.1%, 3.0%, and 0.4% in mAP50, and 1.4%, 2.2%, and 1.1% in mAP50: 95 on SIMD, VisDrone2019, and RSOD respectively, while maintaining a reasonable model size, validating its effectiveness and practical value in complex aerial photography scenarios. The code will be publicly available at: https://github.com/wbbjjj/MCFAD-yolo-main.
While epitaphs recording tomb occupants’ identities and biographies provide critical insights for archeological discoveries, their cinnabar inscriptions often suffer severe degradation during prolonged burial. To address this challenge, this study proposes a hyperspectral processing framework designed to enhance the readability of degraded cinnabar inscriptions on epitaphs. The framework quantifies sample spectral curves using Euclidean distance, enabling classification into high-contrast and low-contrast groups. For high-contrast groups (Branch 1), optimal spectral bands are selected through the optimum index factors, with subsequent integration of imaging differentials and edge detection to enhance text contrast and boundary definition. Branch 2 employs pseudo-color image synthesis, dark channel prior-based feature restoration, and bilateral filtering Retinex algorithms to achieve noise reduction and feature enhancement for low-contrast groups. Experimental results demonstrate the framework’s superiority over conventional methods in significantly improving the readability of cinnabar inscriptions, highlighting its potential as a valuable tool for archeological epigraphy.
Generalized Image Extrapolation is an image generation sub-task and a challenging ill-posed problem. This task intends to predict unknown regions based on the center area. Unfortunately, existing methods encounter the triple dilemma: (1) Convolutional Neural Networks (CNNs)-based methods can precisely extract local details but underperform in capturing global semantic information due to inductive bias, resulting in a lack of consistency in the image structure and layout. (2) Vision Transformer (ViT)-based methods, although superior to global information extraction, are not sufficiently fine-grained in detail and texture generation, and (3) ViTbased approaches rely on the self-attention mechanism, which leads to a tremendous computational burden in processing images and makes model training inefficient. We propose a novel model named Mamba-GIE, designed to effectively balance information of different granularities and address the unresolved challenges in GIE tasks. At the macro level, Mamba-GIE adopts a U-shaped encoder-decoder architecture, with its core basic block being the improved Hybrid State Space Models (Hybrid-SSMs). Specifically, within the basic blocks, the input feature map is processed via two parallel branches: (1) Extracting global information via the Mamba branch and (2) Handling local details using the CNNs branch. At the micro level, we introduce the dual-level adaptive feature fusion mechanism to achieve adaptive feature fusion in intra- and inter-HybridSSMs blocks. Extensive experiments on three public datasets demonstrate that our approach outperforms existing GIE methods inmost evaluation metrics and image generation quality. Comprehensive ablation studies and resource consumption assessments further reveal the efficiency and effectiveness of Mamba-GIE. Code: https://github.com/zrymsm/Mamba-GIE.
In uncertain battlefield environments, rapid and accurate detection, identification of hostile targets, and assessment of threat levels are crucial for supporting effective decision-making. Despite offering the advantage of structural transparency, traditional analytical methods rely on expert knowledge to construct models and often fail to comprehensively capture the non-linear causal relationships among complex threat factors. In contrast, data-driven methods excel at uncovering patterns in data but suffer from limited interpretability due to their black-box nature. Owing to probabilistic graphical modeling capabilities, Bayesian networks possess unique advantages in threat assessment. However, existing models are either constrained by the limitation of expert experience or suffer from excessively high complexity due to structure learning algorithms, making it difficult to meet the stringent real-time requirements of uncertain battlefield environments. To address these issues, this paper proposes a new method, the Tree-Hillclimb Search method-an efficient and interpretable threat assessment method specifically designed for uncertain battlefield environments. The core of the method is a structure learning algorithm constrained by expert knowledge-the initial network structure constructed from expert knowledge serves as a constraint, enabling the discovery of hidden causal dependencies among variables through structure learning. The model is then refined under these expert knowledge constraints and can effectively balance accuracy and complexity. Sensitivity analysis further validates the consistency between the model structure and the influence degree of threat factors, providing a theoretical basis for formulating hierarchical threat assessment strategies under resource-constrained conditions, which can effectively optimize sensor resource allocation. The Tree-Hillclimb Search method features (1) enhanced interpretability; (2) high predictive accuracy; (3) high efficiency and real-time performance; (4) actual impact on battlefield decision-making; and (5) good generality and broad applicability.