In order to make full use of spatial information for classification and address the challenges posed by the differences in medical images of different modalities, we propose the SFCF Net (Spatial and Frequency Domain Fusion Network) framework. This framework adopts a three-stage network design: In the first stage, a spatial method is used to extract features to obtain features with sufficient spatial details and semantic information. In the second stage, these features are mapped in both the spatial and frequency domains. In the frequency domain mapping, we introduce the Discrete Wavelet Transform Feature Mapping (DWTF) structure. This structure uses the Haar wavelet transform to decompose the features into low-frequency and high-frequency components and integrates them with the spatial features. In the third stage, to bridge the semantic gap between the features in the frequency domain and the spatial domain and facilitate the combination of important features from different representation domains, we design the Multi-Domain Cross-Fusion Module (MDCF). This structure utilizes multi-scale vertical and horizontal convolutions, channel attention, and cross-attention to achieve the cross-fusion between the spatial domain and the frequency domain. The comprehensive experimental results show that compared with existing methods, SFCF Net achieves superior performance in terms of the main performance indicators (Accuracy, Precision, Recall, and F1) for medical images of different modalities, reaching 90%, 90.2%, 90%, and 89.97% respectively.
This paper explores the application of cooperative learning in junior high school English writing instruction, emphasizing its significance and effectiveness in en-hancing students' writing skills and fostering a more interactive and engaging classroom atmosphere. Drawing upon theoretical frameworks and empirical evi-dence, the study discusses various strategies for integrating cooperative learning into English writing classes, aiming to stimulate students' creativity, im-prove their writing quality, and promote peer learning. The paper concludes by highlighting the positive impacts of cooperative learning and suggesting avenues for further research.
Medical image classification is crucial for clinical diagnosis. However, medical datasets often face challenges such as limited characterization capabilities, difficulties in category differentiation, and the presence of individual differences. Although attention mechanisms can enhance feature representation, existing methods often struggle to utilize spatial information effectively and lack modeling of inter-channel interactions. We propose star-shaped multi-scale attention (StarMA). (1) It retains spatial information along different orientations within each channel through axial decomposition, enhancing the model's perception of complex structures. (2) A star-shaped structure is employed to project feature maps into a high-dimensional nonlinear space, strengthening inter-channel interactions and improving inter-class discriminability. (3) A cross-spatial aggregation learning strategy is introduced to integrate multi-scale contextual information, improving the model's ability to handle intra-class variability. Based on StarMA, we developed StarMA Net and conducted comparative experiments on five different datasets. Compared to advanced algorithms, StarMA Net demonstrates better effectiveness and robustness.
The issue of water pollution critically affects all living beings. The implementation of a smart water quality monitoring system, based on the Internet of Things, enables advancements in efficiency, security, and costeffectiveness while providing real-time capabilities. Current water quality prediction models often fail to fully utilize data characteristics shared by water quality indicators, resulting in poor predictive accuracy. This study introduces a novel water quality prediction model named TGMHSA, which utilizes tensor decomposition combined with a gated neural network and a multi-head self-attention mechanism. The aim is to tackle the difficulty of forecasting water quality indicators using time series data while minimizing the risk of plagiarism. The proposed model utilizes standard delay embedding transformation (SDET) to convert the time series data into tensor data, extracting data characteristics by Tucker tensor decomposition, and then combines a multi-head self-attention mechanism to discover potential relationships among data characteristics of multiple water quality indicators. Finally, the utilization of the GRU model enables accurate prediction of multi-index water quality. In order to compare its performance, we consider four indices: root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE), and coefficient of determination represented as R2. The outcomes demonstrate that this model outperforms traditional methods for predicting water quality in terms of accuracy and resilience, thereby establishing a scientific foundation for effective water quality prediction and environmental monitoring management.
This research proposes a multi-stage feature fusion network (MSFF) for medical image classification. In view of the problems existing in medical images, such as noise, diversity, and similarity among different classes, MSFF enhances the global context perception in the window partitioning framework through Context Modulation Attention (CMA). Meanwhile, it extracts fine-grained local information via the multi-stage Contextual Information Refinement (CIR) module and gradually fuses multi-stage local and global features to generate richer semantic representations. The experimental results demonstrate that MSFF significantly outperforms existing methods in multiple performance metrics (including accuracy, precision, recall, F1-score, Matthews Correlation Coefficient (MCC), Kappa coefficient, Area Under the Curve (AUC), balanced accuracy, and geometric mean) on four datasets (Endoscopic Bladder Tissue, Kvasir, SARS-COV-2 Ct-Scan, and Thyroid Nodule), showing its excellent performance in the task of medical image classification.
Automatic and accurate classification of thyroid nodules is of great significance to doctors for clinical diagnosis and subsequent treatment recommendations. Since there are no obvious features between benign and malignant nodules, enlarging or reducing the image will result in blurred edges and image distortion, thus limiting the accuracy of clinical diagnosis. Furthermore, the prevalence of sample class imbalance in medical images poses significant challenges to applying convolutional neural networks in thyroid nodule classification methods. This paper proposes a network Tnc-Net for thyroid nodule classification. The network backbone can adapt to the problem of small data volume, capture global features with the help of simple channel attention, and effectively extract image information. The branch network supplements the feature extraction from the backbone network, and the information extracted from the backbone and branch networks is effectively utilized through the fusion module. In addition, this article designs training strategies suitable for this network to deal with category imbalance, improve model classification performance, and make classification results more clinically referenceable. The method test accuracy is 0.902, which exceeds other classic deep learning models in classification. This result demonstrates the effectiveness of our method in achieving automatic classification of thyroid nodules.
Optical coherence tomography (OCT) is an important basis for retinal diagnosis. Traditional OCT image analysis methods not only require a lot of manual operation and time, but also have a certain risk of error. Machine Learning (ML) and Deep Learning (DL) have made significant achievements in the medical field. The Convolutional Neural Networks (CNN) model performs well in extracting local features, but is less effective in extracting global features. The transformer-based structure has an advantage in extracting global features and can make up for this deficiency. This paper proposes a hybrid multi-scale network model based on CNN and Swin transformer, called HRS-Net. The model splits into two branches after the convolutional layers of ResNet50. One branch combines attention modules and residual blocks to extract local features, while the other branch primarily uses Swin transformer blocks to extract global features. Finally, the two branches are fused for the multi-classification task of retinal diseases. Experimental results show that on two public datasets, the accuracy of three-classification and four-classification reached 98.76% and 97.16% respectively, which is better than the previous classic model algorithm.
Esophageal cancer (EC) is the sixth leading cause of cancer mortality and has one of the poorest survival rates among all cancers. Endoscopy is the usual method for diagnosing esophageal disease, but it may be subject to doctors' subjective judgments. Objective tools can assist doctors in enhancing their accuracy. This paper proposes a deep fusion network (DFY-Net) based on the weight transfer (WL) method, which uses the Kvasir public dataset, high-resolution images cropped from public video databases, and some private Barrett's esophagitis (BE) Endoscopy pictures. The FDY-Net architecture utilizes a modified ResNet50 network (named RY_RNet50) and VGG16 for feature extraction and selectively freezes specific layers during training. Its top-level part uses support vector machine (SVM) as a classifier to classify esophagitis and BE. Experimental results show that the overall classification accuracy of DFY-Net is as high as 97.71%, of which the precision of esophagitis is 97.22% and the precision of BE is 99.11%. Additionally, we use Grad-CAM (Gradient Weighted Class Activation Map) to increase the interpretability of the model. This study can provide an objective reference for endoscopists' diagnosis, help improve the diagnostic accuracy of esophageal diseases, and play an important role in the prevention of esophageal cancer.
Classroom learning behavior is an important factor affecting students' classroom learning effectiveness. Analyzing students' classroom behavior and exploring its influence on learning effectiveness can provide an important basis for classroom teaching evaluation and form effective feedback information and teaching guidance. Traditional observation of students' classroom behavior is often recorded and marked manually by teachers, which is labor-intensive, subjective and inaccurate. Artificial intelligence computer vision technology brings the possibility of automated annotation of students' classroom behaviors. In this paper, artificial intelligence computer vision recognition technology is applied to the traditional observation of students' classroom behaviors. The classroom learning behaviors of 25 students in a class at a university are used as research samples to test their learning effects, and then deep learning computer vision technology is used to automatically label each classroom learning behavior of different students, and correlation analysis and regression analysis are used to explore the relationship between different classroom learning behaviors and learning effects. It was found that 1) positive learning behaviors have a positive impact on the learning effect and help to improve the classroom learning effect, while negative learning behaviors have a negative impact on the learning effect, 2) note-taking behavior has a more obvious effect on the learning effect than listening to the lecture, and looking at the phone is more likely to significantly reduce the learning effect of students than drifting off. In response to the results, this paper puts forward corresponding improvement measures and suggestions. From the perspectives of both students and teachers, so as to improve students' classroom learning effect, this paper provides an important basis and reference suggestions for classroom teaching evaluation and students' classroom behavior analysis.
The restoration of hyperspectral images (HSIs) is a crucial process that eliminates various types of noise to improve subsequent applications. To effectively utilize the inherent low-rank and spatial smoothness of HSI data, this letter proposes a method that employs multimodal low-rank tensor subspace learning with total variation regularization (MLTSL-TV) model to denoise HSI data based on the observed measurements. The proposed approach utilizes a low-rankness measure of subspace tensors and learnable transform basis to represent a low-rank perspective, which enables adaptive exploitation of potential low-rank structures through multimodal tensor factorization in multiple orientations based on the observed HSI data. More importantly, we put forward a proximal alternating minimization (PAM) algorithm for efficiently solving the proposed model. Experiments were conducted on two simulated and one real HSI dataset, which were compared with representative approaches through both visual and quantitative analysis. The experimental results demonstrate that the proposed MLTSL-TV approach achieves satisfactory performance when compared to the state-of-the-art methods.
Restoration of hyperspectral images (HSI) is a crucial step in many potential applications as a preprocessing step. Recently, low-rank tensor ring factorization was applied for HSI reconstruction, which has high-order tensors’ powerful and generalized representation ability. Although low-rank TR-based approaches with nuclear norm regularization achieved successful results for restoring hyperspectral images, there is still room for improved tensor low-rank approximation. In this article, we propose a novel Auto-weighted low-rank Tensor Ring Factorization with Hybrid Smoothness regularization (ATRFHS) for mixed noise removal in HSI. Nonlocal Cuboid Tensorization (NCT) is leveraged to transform HSI data into high-order tensors. TR factorization using latent factors rank minimization removes the mixed noise in HSI data. To highlight nuclear norms of factor tensors differently effective, an auto-weighted strategy is employed to reduce the more prominent factors while shrinking the smaller ones. A hybrid regularization combining total variation (TV) and phase congruency (PC) is incorporated into a low-rank tensor ring factorization model for the HSI noise removal problem. This efficient combination yields sharper edge preservation and resolves this weakness of existing pure TV regularization. Moreover, we develop an efficient algorithm for solving the resulting optimization problem using the framework of alternating minimization. Extensive experimental results demonstrate that our proposed method can significantly outperform existing approaches for mixed noise removal in HSI. The proposed algorithm is validated on synthetic and natural HSI data.
Time-varying Quality-of-Services (QoS) data describes the non-functional characteristics of Web service, which plays a key role in service selection. Whereas, QoS data is often high-dimensional and incomplete (HDI) due to the impossibility for users to request all services. A Latent Factorization of Tensors (LFT)-based QoS predictor proves to be efficient in predicting time-varying QoS data. However, current LFT models mostly use $L_{2}$ -norm-oriented Loss. Yet $L_{2}$ norm is sensitive to outlier data, resulting model robustness not being guaranteed. Moreover, although $L$ , norm has intrinsically robustness, it is less sensitive to error. To address the above problems, this study proposes a Double-norm Aggregated Latent factorization of tensors (DAL) model. Its main idea is to aggregate $L_{2}$ -norm and smooth $L_{1}$ -norm to form its Loss, making it have both high accuracy and strong robustness in predicting the unobserved time-varying QoS data. Empirical studies on two time-varying QoS datasets shows that the proposed model has higher prediction accuracy and better convergence rate than state-of-the-art models.
Automatic segmentation of skin lesions is of great significance for assisting doctors in early diagnosis. However, due to these problems such as color, size, boundary blur, low contrast, and hair occlusion, it is more difficult to automatically segment skin diseases. In order to overcome these challenges, reduce the missed diagnosis rate and misdiagnosis rate, prevent patients from missing the best treatment period, and improve patient survival. We proposed a new lightweight network based on multi-scale Transformer, which extracts rich global context dependencies through the multi-scale Transformer of four parallel paths, and adds a simple boundary enhancement structure to the last layer of the network as local information , and finally segmented by a layer-by-layer decoding structure. A large number of experiments were carried out on four public data sets. Through the analysis of objective evaluation indicators, our method has improved compared with other excellent methods. From the comparison of visual effects, we can also see that our method is more accurate in boundary segmentation. Our method improved the segmentation accuracy and reduces the scale of the model, which is conducive to the rapid automatic segmentation of skin diseases, assists skin doctors in diagnosis, and improves efficiency and accuracy.
Hyperspectral imagery (HSI) restoration is a fundamental problem as a preprocessing step. In this letter, we present a novel auto-weighted nonlocal tensor ring rank minimization (ANTRRM) to reduce noise in HSI. First, nonlocal cuboid tensorization (NCT), built by similar grouping cuboids in HSI data, exploits the nonlocal self-similarity and the spatial–spectral correlation simultaneously. Then, the proposed model introduces nuclear norm (NN) regularization via nonlocal tensor ring with mode-{ $d$ , $l$ } unfolding. An auto-weighted optimization is employed to represent the different importance of TR unfolding. Finally, the alternating direction method of multipliers (ADMM) scheme is employed to solve the proposed model efficiently. Experiments on two simulation HSIs datasets and a real HSI dataset were carried out, compared with representative approaches in visual and quantitative comparison. The proposed ANTRRM method is superior except in a few cases.
A vast amount of images has been generated due to the diversity and digitalization of devices for image acquisition. However, the gap between low-level visual features and high-level semantic representations has been a major concern that hinders retrieval accuracy. A retrieval method based on the transfer learning model and the relevance feedback technique was formulated in this study to optimize the dynamic trade-off between the structural complexity and retrieval performance of the small- and medium-scale content-based image retrieval (CBIR) system. First, the pretrained deep learning model was fine-tuned to extract features from target datasets. Then, the target dataset was clustered into the relative and irrelative image library by exploring the Bayes classifier. Next, the support vector machine (SVM) classifier was used to retrieve similar images in the relative library. Finally, the relevance feedback technique was employed to update the parameters of both classifiers iteratively until the request for the retrieval was met. Results demonstrate that the proposed method achieves 95.87% in classification index F1 - Score, which surpasses that of the suboptimal approach DCNN-BSVM by 6.76%. The performance of the proposed method is superior to that of other approaches considering retrieval criteria as average precision, average recall, and mean average precision. The study indicates that the Bayes + SVM combined classifier accomplishes the optimal quantities more efficiently than only either Bayes or SVM classifier under the transfer learning framework. Transfer learning skillfully excels training from scratch considering the feature extraction modes. This study provides a certain reference for other insights on applications of small- and medium-scale CBIR systems with inadequate samples.
During the normalization of the anti-epidemic period, hearing impaired children in remote areas need to take distance learning to ensure their own rehabilitation. Based on the existing rehabilitation process of rehabilitation institutions, combined with the design model of distance education platform, this research builds an online rehabilitation learning platform for hearing-impaired children, and applies it to actual rehabilitation teaching. The platform changes the traditional offline education model and the child-centered intervention model, taking "parental intervention" as the starting point, and adopts the form of Internet + education. The platform offers five series of courses, including personal training, language listening, cognition, psychosocial and knowledge activities, to provide comprehensive and comprehensive rehabilitation training for hearing-impaired children. The application effect shows that the online rehabilitation training learning platform designed for hearing-impaired children can effectively solve the rehabilitation needs of hearing-impaired children and improve the rehabilitation effect of hearing-impaired children to a certain extent.
At present, there is a certain lag in the teaching ability of computational thinking in our country. This article is based on STEM86 platform, ISTE published by the author "Educator standards: Under the guidance of computational Thinking ability, a standard system for evaluating teachers' computational thinking teaching ability was constructed, and a test scale for evaluating teachers' computational thinking teaching ability was designed. The reliability and validity test and difficulty differentiation test proved that the designed test scale had good scientificity and reliability. On this basis, using this scale, through the empirical way, from three dimensions of computer discipline knowledge and skills, teaching design ability, evaluation and reflection ability to evaluate teachers before and after using STEM86 platform computational thinking teaching ability changes. The results show that STEM86 platform not only solves teachers' programming teaching resource needs, but also effectively improves teachers' computational thinking teaching ability, and provides more references and ideas for teachers' computational thinking teaching ability improvement research.
Wavelet threshold denoising and non-local mean denoising are traditional image denoising methods, but wavelet hard threshold denoising will produce pseudo Gibbs phenomenon due to some discontinuous wavelet coefficients after wavelet reconstruction; wavelet soft threshold denoising is in After wavelet reconstruction, the image accuracy will be reduced due to the constant deviation between the approximate wavelet coefficients and the original wavelet coefficients; traditional non-local mean filtering will increase the noise of similar local blocks due to the search for weights with the increase of noise, resulting in low confidence The noise denoising effect of the noise ratio is not good. In view of the above situation, this paper proposes an improved wavelet threshold combined with non-local mean filtering to denoise images. The image denoising effect is better than wavelet threshold denoising or non-local mean denoising.
The topological information of a dynamic network varies over time, making it crucial to capture its temporal patterns for predicting missing links accurately. A latent factorization of tensors (LFT)-based model has proven to be efficient to solve this problem, where a dynamic network is represented as a three-way high-dimensional and sparse (HiDS) tensor. However, currently LFT-based models do not consider multiple biases in analyzing an HiDS tensor for accomplishing dynamic link prediction. To address this issue, this paper proposes a multiple biases-incorporated latent factorization of tensors (MBLFT) model, which integrates short-term bias, preprocessing bias and long-term bias into an LFT model. Empirical studies on two large-scale dynamic networks from real applications show that compared with state-of-the-art predictors, an MBLFT model achieves higher prediction accuracy and computational efficiency for missing links in dynamic network.
An effective fraction of data with missing values from various physiochemical sensors in the Internet of Things is still emerging owing to unreliable links and accidental damage. This phenomenon will limit the predicative ability and performance for supporting data analyses by IoT-based platforms. Therefore, it is necessary to exploit a way to reconstruct these lost data with high accuracy. A new data reconstruction method based on spectral k-support norm minimization (DR-SKSNM) is proposed for NB-IoT data, and a relative density-based clustering algorithm is embedded into model processing for improving the accuracy of reconstruction. First, sensors are grouped by similar patterns of measurement. A relative density-based clustering, which can effectively identify clusters in data sets with different densities, is applied to separate sensors into different groups. Second, based on the correlations of sensor data and its joint low rank, an algorithm based on the matrix spectral k-support norm minimization with automatic weight is developed. Moreover, the alternating direction method of multipliers (ADMM) is used to obtain its optimal solution. Finally, the proposed method is evaluated by using two simulated and real sensor data sources from Panzhihua environmental monitoring station with random missing patterns and consecutive missing patterns. From the simulation results, it is proved that our algorithm performs well, and it can propagate through low-rank characteristics to estimate a large missing region’s value.