Transformer models have achieved powerful performance in various computer vision tasks. However, their black-box nature severely limits model interpretability and the reliability of real-world applications. Most existing interpretation methods generate explanation maps by perturbing masks from the last layer of the Transformer encoder, but they often overlook uncertain information in masks and detail loss during upsampling and downsampling, resulting in coarse localization, blurred boundaries, and significant background noise in explanations. To address these issues, this paper proposes a self-distillation object segmentation method based on sequential three-way mask and attention fusion (SAF-SD), targeting salient and camouflaged binary object segmentation tasks (sub-tasks of binary pixel-level segmentation). The method consists of two core modules: the sequential three-way mask (S3WM) module and the attention fusion (AF) module. The S3WM module performs strict threshold filtering on masks generated from the final-layer feature maps of the Transformer, aiming to accurately segment foreground objects from backgrounds via binary pixel-level prediction. The AF module aggregates attention matrices across all Transformer encoder layers to construct a cross-layer relation matrix, capturing global semantic dependencies among image patches (e.g., interactions between foreground, background, and edge regions). It then computes the importance score for each patch, refining details and suppressing noise in the initial explanation results. Extensive experimental results demonstrate that SAF-SD significantly outperforms existing baseline methods across key evaluation metrics.
The transformer architecture has demonstrated significant performance in image super-resolution (SR). However, for existing Transformer-based models, there are common drawbacks. They often fall short in local feature modeling and limited feature representation capabilities. When it comes to reconstructing high-resolution (HR) images, these deficiencies become more prominent, resulting in the poor restoration of fine details. In order to resolve the existing issues, we propose the Hierarchical Multiscale Transformer Architecture Based on Hybrid Attention (HMT). Notably, this architecture has the ability to effectively capture the fine-grained interactions between local image features and other regions, which in turn enables it to generate details that are clearer and more coherent. Specifically, we introduce the Fourier Mamba Synergy (FMSA) module, which adopts a dual branch hierarchical architecture design to achieve cross-domain feature collaboration and efficient long-range dependency modeling by constructing Fourier Spectrum Modulation (FSM) and Mamba Spatial Mixing (MSM). In addition, we introduced Dynamic Convolutional Attention (DCA), which separates feature branches based on channel splitting strategy and achieves multi-scale feature enhancement through long-range dependency capture and instance weight adaptation. Finally, we designed a multi-scale encoding and decoding architecture and dynamic fusion mechanism. HMT exceeds HGFormer by 0.07dB and 0.08dB in PSNR metrics at a scaling factor of 4 on BSD100 and Manga109 datasets respectively. Extensive experiments on public datasets show good performance in both objective metrics and visual quality. The code can be available at https://github.com/lingyuyan2014/HMT.
Colorectal cancer is the second most common cancer globally. Its high mortality necessitates early polyp detection to mitigate the risk of the disease. However, conventional segmentation methods are susceptible to noise interference and have a limited accuracy in complex environments. To address these challenges, we propose GSCCANet with an encoder-dual decoder co-design. The encoder employs hybrid Transformer (MiT) for efficient multi-scale global feature extraction. Dual decoders collaborate via SAFM and REF-RA modules to enhance segmentation precision through global semantics and boundary refinement. In particular, SAFM enhances lesion coherence via channel-space attention fusion, while REF-RA strengthens low-contrast edge response using high-frequency gradients and reverse attention, optimized through progressive fusion. Additionally, combined Focal Loss and Weighted IoU Loss mitigate the problem of undetected small polyps. Experiments on five datasets show GSCCANet surpasses baselines. It achieves 94.7% mDice and 90.1% mIoU on CVC-ClinicDB (regular) and 80.1% mDice and 72.5% mIoU on ETIS-LaribPolypDB (challenging). Cross-domain tests (CVC-ClinicDB ->$$ \to $$ Kvasir) confirm strong adaptability with 0.2% mDice fluctuation. These results prove that GSCCANet offers high-precision and generalizable solutions through global-local synergy, edge enhancement, and efficient computation.
With the continuous growth of network traffic and the evolving application requirements, the dynamic traffic management problem in Wide Area Networks (WANs) is becoming increasingly complex. Traditional traffic engineering methods often struggle to achieve rapid and efficient responses in dynamic environments involving topological changes and link failures. To address these challenges, this paper proposes GDWRO, a real-time WAN routing optimization framework that combines Graph Neural Network (GNN) with Deep Reinforcement Learning (DRL). Firstly, GDWRO designs a GNN module to enhance the representation of global topological structure and path-level features, thereby improving the accuracy of traffic feature representation, and providing critical input for routing strategy generation. Secondly, GDWRO introduces structured path embedding to optimize deep reinforcement learning policies, updating parameters through agent-environment interactions, enhancing generalization capability and robustness across diverse network topologies. Finally, GDWRO employs a Local Search (LS) mechanism to refine the initial solution of routing strategies, rapidly exploring the solution space to identify high-quality solutions, thereby enhancing the overall quality of solutions. The experimental results indicate that GDWRO is able to operate in real-world dynamic network topologies in 4.8 seconds on average for topologies up to 100 links.
In recent years, more and more scholars have applied graph neural networks in combination with other modules to the field of traffic flow forecasting and achieved outstanding results. Most of these graph-based methods describe pairwise relationships between two objects, but in real transportation networks, the relationships between objects are often of a high order. To effectively learn the higher-order relationships between objects, this paper proposes a dynamic spatio-temporal residual hypergraph convolutional network for traffic forecasting (Res-DSTHGCN). In this paper, we combine dynamic spatio-temporal graph convolution and dynamic spatio-temporal residual hypergraph convolution to capture more global spatio-temporal features than models with only spatio-temporal graph convolution or only spatio-temporal hypergraph convolution. Meanwhile, the hypergraph convolutional network enhanced by residual connectivity improves the oversmoothing problem that may occur in the traditional hypergraph convolutional network (HGCN) along with the continuous stacking of layers. It is proved that the prediction accuracy of the method proposed in this paper is improved by conducting experiments for comparison with other baselines.
The interconnection between errors due to drift and due to inhomogeneity of thermocouples is considered in this paper. A mathematical model was developed to describe the process of developing a thermocouple electromotive force. A formula is derived to describe this interconnection, and modeling is carried out to prove the formula. It is shown that the sum of absolute values of errors due to drift and inhomogeneity is always equal.
In our article, U-Net network is selected as the basic network architecture, and a super-resolution reconstruction algorithm of multi-level image feature fusion is designed. Among them, the network architecture of Residual-in-Residual Dense Block (RRDB) is introduced, and light-weight attention with channel attention mechanism (CA) and spatial attention (SA) mechanism is added to the RRDB residual block. On this basis, the multi-level attention residual block R are designed through the nesting of residual blocks, and the high-frequency residual features of different levels are extracted through the dense connection of the multi-level residual network, so that the generated super Resolution reconstructed images are richer in texture detail. The network model is trained on the DIV2K and Flickr2K training sets and fine-tuned. Finally, the proposed reconstruction algorithm and comparison algorithms are tested on five commonly used test sets. The experimental results demonstrate that, compared with the comparative algorithm, the improved reconstruction algorithm performs better in terms of objective evaluation metrics PSNR and SSIM, generating super-resolution images that are closer to real high-definition images and more in line with human visual perception.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
Insulator defect detection plays a critical role in ensuring electrical equipment’s safe and stable operation, meeting the public’s demand for electricity consumption. However, extracting features of insulator defects poses challenges due to complex backgrounds, variations in target sizes leading to potential oversights, and low detection accuracy. We propose an improved YOLOv8n-based insulator defect detection model to achieve timely and precise real-time detection. Firstly, the TripletAttention Module is introduced to enhance the network’s ability to extract insulator defect features and reduce background interference in detection. Secondly, SCConv (Spatial and Channel Reconstruction Convolution) is utilized to redesign the detection head, proposing a more lightweight SC-Detect to replace the original one, thereby restricting feature redundancy and enhancing feature representation capability. Finally, Slim-neck based on GSConv is employed to reconstruct the neck structure, enabling the network to achieve lightweight while possessing relatively stronger feature extraction and perceptual capabilities. Experimental results demonstrate that the improved insulator defect detection network achieves an accuracy of 96.1 - 0.95 of 72
This work proposes efficient multi-scale object detection model with space-to-depth convolution and BiFPN combined with FasterNet (ES-BiCF-YOLOv8), a deep learning method, to address the problems associated with detecting steel surface defects in contemporary industrial production. The method makes innovative improvements based on the YOLOv8 algorithm and enhances the performance of the novel model mainly through the following aspects. First, the space-to-depth layer followed by a non-strided convolution layer (SPD-Conv) and the efficient multi-scale attention mechanism is introduced into the feature extraction network to enhance the model's ability to capture fine-grained information and the fusion of multi-scale features. Second, the feature fusion network is optimized by utilizing a weighted bi-directional feature pyramid network and a lightweight network, FasterNet, to improve computational efficiency. Finally, it is shown that ES-BiCF-YOLOv8 reduces the complexity and computational requirements of the model while increasing the detection accuracy utilizing the NEU-DET dataset and deepPCB dataset with substantial experimental validation. The ES-BiCF-YOLOv8 model achieves a 5% improvement of the mean average precision value on the NEU-DET dataset, with the number of parameters and the computational amount only being the baseline 89% and 27%, and also demonstrates good generalization performance on the deepPCB dataset. Furthermore, the experiments demonstrate that ES-BiCF-YOLOv8 can be used for steel surface defect detection in industrial production because it uses less computational resources and can detect in real-time while maintaining high accuracy, in comparison to other popular object detection algorithms. The results of this work not only improve the efficiency and accuracy of steel surface defect detection but also provide ideas for the application of deep learning in the field of industrial detection.
The DCELANM-Net structure, which this article offers, is a model that ingeniously combines a Dual Channel Efficient Layer Aggregation Network (DCELAN) and a Micro Masked Autoencoder (Micro-MAE). On the one hand, for the DCELAN, the features are more effectively fitted by deepening the network structure; the deeper network can successfully learn and fuse the features, which can more accurately locate the local feature information; and the utilization of each layer of channels is more effectively improved by widening the network structure and residual connections. We adopted Micro-MAE as the learner of the model. In addition to being straightforward in its methodology, it also offers a self-supervised learning method, which has the benefit of being incredibly scaleable for the model.
Graph convolutional networks (GCN) are an important research method for intelligent transportation systems (ITS), but they also face the challenge of how to describe the complex spatio-temporal relationships between traffic objects (nodes) more effectively. Although most predictive models are designed based on graph convolutional structures and have achieved effective results, they have certain limitations in describing the high-order relationships between real data. The emergence of hypergraphs breaks this limitation. A dynamic spatio-temporal hypergraph convolutional network (DSTHGCN) model is proposed in this paper. It models the dynamic characteristics of traffic flow graph nodes and the hyperedge features of hypergraphs simultaneously, achieving collaborative convolution between graph convolution and hypergraph convolution (HGCN). On this basis, a hyperedge outlier removal mechanism (HOR) is introduced during the process of node information propagation to hyper-edges, effectively removing outliers and optimizing the hypergraph structure while reducing complexity. Through in-depth experimental analysis on real-world datasets, this method has better performance compared to other methods.
As a challenge in the field of smart medicine, medical picture segmentation gives important decisions and is the basis for future diagnosis by doctors. In the past decade, FCN-based network topologies have made amazing progress in the field. However, the limited perceptual capacity of convolutional kernels in FCN network topologies limits the network's ability to acquire a global field of view. We propose BSANet, a 3D medical image segmentation network based on self-focus and multi-scale information fusion with a high-performance feature extraction module. BSANet can help the network to extract deeper features by obtaining a larger range of perceptual capabilities by using its self-focus and multi-scale information aggregation pooling modules. Brain tumor segmentation dataset and multi-organ segmentation dataset are used to train and evaluate our model. BSANet produces excellent results with its high-performance feature extraction network with an attention module and multi-scale information fusion module.
The low-illumination image enhancement method based on the adaptive MSRCR algorithm is proposed to address the problems of the Retinex algorithm in processing low-illumination images, such as the need to manually adjust parameters and blurred details. In the HSV color space of the original image, the luminance V component is decomposed by mean filtering to create a detail layer, and the detail layer information is enhanced by using enhancement weights. The improved Salp Swarm Algorithm (LLSSA) is proposed for adaptive parameter adjustment of Multi-Scale Retinex with Colour Recovery (MSRCR) and detail layer weights, which uses Logistic Chaos to initialise the salps population and introduces Lévy flights into the updated positions of leaders and followers to enhance the global search capability. Finally, the adaptive MSRCR enhancement map and the detail layer enhancement map are images fused to produce a final enhanced image with clear details. The experimental results show that compared with several typical algorithms, the algorithm in this paper can effectively maintain the image details, improve the image brightness and have better visual effects.
The assessment of temperature measurement errors by platinum resistance temperature detectors (RTD) was carried out. High measurement accuracy assured with their individual calibration, the voltage divider circuit for measuring resistance, the substitution method and the transitional measure. In this case, error due heating the RTDs by their operating current needs correction. The proposed method of correction of RTD’s error due to heating by the operating current decreased this error in two times. The residual error was estimated to be no more than 0.004°C.
In recent years, how to forecast traffic flow quickly and accurately has become a key issue in building an intelligent transportation system. Due to the temporal and spatial correlation of traffic flow data, we propose a prediction model combining convolutional neural network (CNN), gated recurrent unit (GRU) and improved slime mould algorithm (ISMA). The basic idea is to construct the traffic flow data as a two-dimensional matrix containing temporal and spatial information, and use CNN to obtain location-related spatial features and use GRU's memory function to obtain the temporal distribution features. Secondly, for the shortcomings of the slime mould algorithm with low initial population quality, this paper adds Tent chaos mapping and adaptive inertia weighting strategy to obtain the ISMA algorithm and uses it to find the optimal combination of hyperparameters for the GRU network to construct the ISMA-CNN-GRU prediction model. Finally, simulation experiments are conducted on the traffic flow dataset of Heathrow Airport, UK. The experiments confirm that the ISMA-CNN-GRU exhibits higher prediction accuracy compared to the APSO-GRU model, the unoptimized CNN-GRU model and the SMA-CNN-GRU model.
An important component of the computer systems of medical diagnostics in dermatology is the device for recognition of visual images (DRVI), which includes identification and segmentation procedures to build the image of the object for recognition. In this study, the peculiarities of the application of detection, classification and vector-difference approaches for the segmentation of textures of different types in images of dermatological diseases were considered. To increase the quality of segmented images in dermatologic diagnostic systems using a DRVI, an improved vector-difference method for spectral-statistical texture segmentation has been developed. The method is based on the estimation of the number of features and subsequent calculation of a specific texture feature, and it uses wavelets obtained by transforming the graph of the power function at the stage of contour segmentation. Based on the above, the authors developed a modulus for spectral-statistical texture segmentation, which they applied to segment images of psoriatic disease; the Pratt's criterion was used to assess the quality of segmentation. The reliability of the classification of the spectral-statistical texture images was confirmed by using the True Positive Rate (TPR) and False Positive Rate (FPR) metrics calculated on the basis of the confusion matrix. The results of the experimental research confirmed the advantage of the proposed vector-difference method for the segmentation of spectral-statistical textures. The method enables further supplementation of the vector of features at the stage of identification through the use of the most informative features based on characteristic points for different degrees and types of psoriatic disease.
Feature selection is the procedure of extracting the optimal subset of features from an elementary feature set, to reduce the dimensionality of the data. It is an important part of improving the classification accuracy of classification algorithms for big data. Hybrid metaheuristics is one of the most popular methods for dealing with optimization issues. This article proposes a novel feature selection technique called MetaSCA, derived from the standard sine cosine algorithm (SCA). Founded on the SCA, the golden sine section coefficient is added, to diminish the search area for feature selection. In addition, a multi-level adjustment factor strategy is adopted to obtain an equilibrium between exploration and exploitation. The performance of MetaSCA was assessed using the following evaluation indicators: average fitness, worst fitness, optimal fitness, classification accuracy, average proportion of optimal feature subsets, feature selection time, and standard deviation. The performance was measured on the UCI data set and then compared with three algorithms: the sine cosine algorithm (SCA), particle swarm optimization (PSO), and whale optimization algorithm (WOA). It was demonstrated by the simulation data results that the MetaSCA technique had the best accuracy and optimal feature subset in feature selection on the UCI data sets, in most of the cases.
Intent-Based Networking (IBN) focuses on technologically independent, fast, and reliable interaction between network infrastructure management systems and users. This concept is being developed to automate and accelerate the deployment of the network lifecycle. Thus, IBN includes mechanisms for recognizing, understanding, and extending the intents that define the quality-of-service (QoS) requests. Therefore, this paper proposes an intent-based software-defined wireless network (IBSDWN) approach in which the controller intelligently decides when to initiate the handover of services. For this purpose, the controller selects the access point (AP) to which the client device should connect, based on the set quality of experience (QoE) requirements using an intent processing technique. A method for initiating handover in IBSDWN based on machine learning (ML) algorithms and an integral QoE criterion formed from real-time measurements of parameters: received signal strength indication (RSSI), throughput, packet loss, and delay is developed. Implementation of ML module in IBN architecture for monitoring system allowed to reduce the volume of signal traffic in communication channels between network equipment and controller. Also, the developed ML module made it possible to detect the degradation of QoE values and prevent situations when the user is not satisfied with the received QoS for adaptive prediction of the moment of network reconfiguration. On the basis of a simulation model, it is proved that the proposed solutions can improve the quality of experience of multimedia services to end-users.
Network planning of multi-layer heterogeneous mobile networks with complex topology is an important task. In this paper the spectral and energy efficiency of integrated LTE/Wi-Fi technologies for 5G are improved. A method of adaptive formation of the structure of radio access level (RAL) with the provision of the required quality of service (QoS) and the possibility of broadband data transmission is proposed. The Voronoi tessellation is used for designing the RAL of 5G mobile networks for the placement of base stations, which allowed optimum delimiting the coverage area for each base station and provide users the cell interface services. To minimize interference there was proposed a method of dynamic frequency reuse for different sizes of Voronoi cells. Modelling shows the developed method is effective both at low and at high network load, however, at high load there is a slightly smaller gain in energy efficiency than at low load.
Anatoly Sachenko合作论文数Department of Information Computing Systems and Control
Ternopil National Economic University6