
Different from traditional measuring method based on the shipborne special equipment, satellite-based method has various potential advantages. This paper investigates the bathymetry retrieval problem from the hyperspectral remote sensing image. The idea is to make use of each spectral data, and fifind the relevant ones for water depth through the spectral attention weights. To fully exploit the spatial-spectral data, a convolutional neural network (CNN) is employed. It is fed with the small square hyperspectral patches and is required to output the bathymetry for the center point in the patch. The CNN has a side branch which outputs the attention weight for each channel, and it emphasizes the important ones by specifying a large value for it. Together with the model parameters, the attention weight helps mining the hyperspectral data for the accurate prediction.
CMOS sensors, which can directly interact with high-energy radiation and are widely distributed at an affordable price, have opened up a new avenue for nuclear security. When using CMOS sensors to detect nuclear radiation events, the main sources of interference are moving objects and changes in lighting conditions. This paper presents a radiation event extraction algorithm based on multi-scale spatial processing, combining frame differencing and the CNT algorithm. The input video sequence undergoes a pyramid transformation to obtain three layers of images at different resolutions. Three-frame differencing is applied at each resolution level, followed by thresholding and result fusion to extract radiation and motion information. Simultaneously, the CNT algorithm, morphological operations, and median filtering are applied at different resolution levels to extract motion information. Finally, radiation information is obtained by subtracting motion information from the fused radiation and motion result. Experimental results demonstrate that multi-scale spatial processing can reduce the impact of dynamic backgrounds and lighting variations. Additionally, the number of radiation event frames at various distances fits a square inverse relationship with the distance from the radiation source, allowing for accurate radiation event extraction.
Intracranial pressure (ICP) monitoring is of great significance in the treatment of neurological diseases. Currently, the probe of intracranial pressure monitors is commonly a Wheatstone bridge-type pressure sensor. However, it can be affected by temperature changes, resulting in changes in output pressure signals. This paper implements two signal processing schemes based on the structural characteristics of Wheatstone bridges, one for programmable resistors and the other for dual current sources, to reduce changes in output pressure signals due to temperature changes. After verification, both proposed signal processing schemes can work stably and achieve the desired results. Compared with the existing processing schemes of intracranial pressure monitors, the proposed schemes are simple and easy to implement, with good processing performance and can work well in actual environments.
The human brain is a highly connected network with complex patterns of correlated and anticorrelated activity. Analyzing functional connectivity matrices derived from neuroimaging data can provide insights into the organization of brain networks and their association with cognitive processes or disorders. Common approaches, such as thresholding or binarization, often disregard negative connections, which may result in the loss of critical information. This study introduces an adaptive signed random walk (ASRW) model for analyzing correlation- based brain networks that incorporates both positive and negative connections. The model calculates transition probabilities between brain regions as a function of their activities and connection strengths, dynamically updating probabilities based on the differences in node activity and connection strengths at each time step. Results show that the classical random walk approach, which only considers the absolute value of connections, underestimates the mean first passage time (MFPT) compared to the proposed ASRW model. Our model captures a wide range of interactions and dynamics within the network, providing a more comprehensive understanding of its structure and function. This study suggests that considering both positive and negative connections, has the potential to offer valuable insights into the interregional coordination underlying various cognitive processes and behaviors.
Most state-of-the-art semantic segmentation methods in the medical domain rely on fully supervised deep learning. However, creating high-quality annotated datasets demands extensive labor and domain expertise, leading to significant time and cost investments. While previous approaches have explored semi-supervised and unsupervised learning to alleviate the scarcity of annotated data by leveraging unlabeled samples and achieving promising results, they still fall short of replicating the nuanced annotations provided by medical professionals. In this paper, drawing inspiration from self-training techniques in semi-supervised learning, we propose an innovative approach to address the issue of limited annotated data, called medical images pixel rearrangement. Our method combines image editing with the pseudo-label methodology to generate labeled image pairs. Our method supports iterative training, and with each iteration, the edited images progressively resemble the originals, and the resulting annotations align more closely with the annotations made by medical experts. This novel process allows us to derive labeled pairs of data directly from the pool of unlabeled data. The methodology is implemented through a purpose-built conditional generative model and a dedicated segmentation network. Experimental results on the ISIC18 dataset demonstrate that the annotation quality achieved by our segmentation method is not only comparable to, but potentially surpasses the quality of annotations provided by medical experts.
Based on 20 wind power datasets from different regions, this article uses a series of feature engineering, data normalization, construction of training and validation sets, and five models including TCN, MLP, RNN, Transformer, and LSTM for training and testing. The wind power prediction system is established based on these five models. Finally, the web interface is implemented using Flask and Bootstrap frameworks, demonstrating data classification and test results.
Normal light images are a prerequisite for many visual tasks, making single image enhancement crucial in computer vision. While physical model-based approaches have demonstrated promising results in this domain at an early stage, they often produce undesired anomalies in their outputs, making it difficult to apply them universally across all scenarios. Conversely, learning-based methods yield more authentic outcomes. However, the enhancement capabilities of machine learning algorithms are limited by the lack of pairs of normal light images, and many algorithms focus on a single enhancement task, ignoring the role of high-level tasks in guiding enhancement. This paper introduces a low-light image enhancement algorithm rooted in a fusion strategy that consists of two subtasks, visibility restoration and realism improvement tasks, and combines the advantages of physical modeling and learning methods. In the visibility restoration stage, the fusion of features from upstream and downstream tasks enables our network to take full advantage of the guidance of the upstream task to the downstream task to recover results closer to nature; in the fidelity enhancement stage, the Retinex enhancer and the refinement enhancement network are introduced to enhance the fidelity of the image. Finally, the results of the two subtasks are fused to obtain the output. Experiments substantiate that the algorithm excels in terms of visual effects and objective metrics across both synthetic data and real-life scenes.
In the real world, blind face restoration from face images with unknown degradation is a challenging problem. Directly training deep neural networks often fails to achieve reasonable results due to unknown severe degradations. Existing methods based on generative models produce good results but produce restoration results with overly smooth textures or unnatural overall structures. To address this, we propose a novel approach that combines the strengths of convolutional neural networks and Multi-head Cross-attention. Initially, we employ ResNet50 to generate preliminary W+ latent vectors and multi-resolution scale feature maps. We then enhance the multi-scale feature map using cross attention in a Transformer-based feature match module. This process improves the correlation between spatial features and intermediate semantic features of generativa priors, enhances the distance dependence among different semantic features, and extend latent space to improve semantic expressiveness. We also leverage pre-trained GANs to enhance image reconstruction and introduce an FFT-based frequency-domain loss function to pay more attention to high-frequency features. Experiments show that our model outperforms existing Blind Face Restoration methods in both objective metrics and subjective quality comparisons, especially when restoring severely degraded images in real-world scenarios, resulting in visually realistic results.
Medical image segmentation is an important means to assist doctors in making accurate diagnoses. However, medical images are often subject to noise interference, and the high similarity between the background and target areas also increases the difficulty of segmentation, hence there is room for improvement in segmentation accuracy. To address these challenges, this paper proposes a multi-scale mixed convolutional image segmentation network. The network consists of two main parts: an encoder and a decoder. The encoder is formed by a combination of mixed convolutional module and multi-scale Transformer module, and the mixed convolutional module uses a blend of 2D and 3D convolutions to enhance feature extraction capabilities. Additionally, we designed a multi-scale Transformer module aimed at capturing long-range correlation information at multiple scales. In the decoding phase, we introduce coordinate convolution for the first time in the U-Net architecture to replace traditional convolution, thereby enhancing the model's spatial localization capabilities. In experimental validations on multiple public datasets, MTC-TransUNet outperforms other networks.
The relationship between lower limb length discrepancy (LLD) and gait symmetry in non-clinical environments has consistently garnered significant attention in the relevant research field. Accurate and objective quantification of human gait symmetry has remained a challenging task in recent studies. In this research, we propose a novel method that utilizes deep learning and wearable sensors to precisely measure the dynamic variations in LLD gait symmetry. In this study, we have developed an Att-CNN-LSTM hybrid deep learning model, which extracts spatiotemporal features closely associated with LLD gait symmetry from gait data measured by wearable sensors attached to the ankle, knee, and hip joints of the lower limbs. This allows for accurate differentiation between the left and right lower limbs of various levels of LLD. Our approach offers a new perspective for the early identification of LLD gait symmetry and presents a fresh solution for long-term outdoor monitoring of LLD gait symmetry. Experimental results demonstrate that the developed model can accurately identify LLD asymmetrical gaits at three different degrees , mild, moderate, and severe, with recognition accuracies exceeding 95% across all levels. Specifically, the recognition accuracy for severe LLD asymmetry gait can reach 97.89%. Additionally, for normal gait, our model achieves approximately 50% accuracy in gait symmetry identification.
Tuberculosis (TB) is one of the deadliest infectious diseases and the 13th leading cause of death worldwide. Because of its infectivity, screening in crowded places is particularly important. Although the PPD test, which is currently commonly used for community screening, has the advantage of simplicity and low cost compared to other TB detection methods such as IGRA and X-rays, the high number of false positives does not allow it to be used as a mass screening method. To address these issues, this paper proposes a computer-assisted screening-based method that can differentiate between normal individuals and TB patients by cough sounds. We designed a multi-model voting mechanism to test the classification of 1000 cough tone segments. The experiments were conducted by soft-voting tests with different weight assignments for three models: the resnet50 model, the improved Googlenet model, and the fusion model of Bi-LSTM fused with Googlenet. After averaging 10 cross-validations, the experimental results showed that screening for TB using cough sounds could achieve better results, obtaining 98.05% accuracy, 98.82% sensitivity, and 97.49% specificity. At the same time, this screening method has the advantages of non-invasiveness, repeatable detection, and low cost, which can be applied to the initial screening of crowded places such as communities.
With the increasing diversification of information sources, controlling information quality has become increasingly challenging, leading to the emergence of polarized news reporting. This paper primarily investigates the propagation of polarized news among users on social networks, examines the impact of polarized news propagation on user opinions, and establishes a three-party game model involving users, news self-media, and official media based on game theory. Utilizing the Hegselmann-Krause model, this paper develops the Polarized News Propagation Model (PNPM), simulating the propagation process of polarized news among users to explore its role and impact on the polarization of public opinion. This paper employs a three-party game to investigate the strategic choices of each party under varying parameter influences, providing guidance strategies for governments and official media to direct public opinion on polarized news. Through model simulation experiments, this paper analyzes the game outcomes and influencing factors of polarized news propagation under different parameter conditions, offering valuable insights for managing and addressing the challenges posed by polarized news in today’s complex information landscape.
As a rapid, simple and low-cost diagnostic technology, immunochromatography has a wide range of application in point-of-care testing (POCT). Here, a handheld reader based on image analysis was developed for the detection of HIV in serum or whole blood sample using gold colloid lateral flow strips. The control and test lines from the image of the strip collected by the reader were extracted with an image analysis algorithm including image preprocessing (image crop, grayscale of image, image equalization, median filter and image denoise) and feature extraction. Based on the preprocessed image, the control line was located based on its color difference from the background, and the test line was accurately identified with both preliminary searching and precise boundary location. The performance of the reader was evaluated by experiments using a series of different concentration of HIV serum samples. Comparing to the commercial reader, comparable performance could be achieved with the handheld reader with significantly reduced size.
In this work, extensive simulations are done to compare the performance of the 4 filter types; Linear Kalman filter (LKF), Extended Kalman Filter (EKF), Unscented Kalman Filter (UKF), and Particle Filter (PF). A simple nearly constant velocity (NCV) motion model is used with a Gaussian noise measurement model. Simulations were done with different ground truths, different measurements covariance matrices, and different speeds of the drone. Stone soup software was used in the simulations. The analyses revealed informative results that gave us more understanding of the behavior of the four filters when a common type of motion model such as the NCV model is used.
Alzheimer's disease (AD) is a progressive degenerative brain disease and one of the most common dementias in the elderly population. In recent years, deep learning and artificial intelligence techniques have been widely used in the diagnosis and research of AD. The purpose of this study is to develop an improved hybrid model based on a genetic algorithm (GA) and a particle swarm algorithm (PSO) to find the optimal feature combination for Alzheimer's disease classification and detection. First, the phase synchronization index (PSI) method is used to extract geometric features from the Alzheimer's disease ROI(Regions of Interests) signals extracted from the AAL(Automated Anatomical Labeling )-116 atlas; second, the genetic algorithm (GA) with Gaussian variation (GDM) and the particle swarm algorithm (PSO) are combined with the improved optimization algorithm for feature selection, the GA-PSO; finally, the combination of features is fed to an SVM classifier for detection. The method proposed in this paper uses a 5-fold cross-validation strategy in feature combination and achieves a classification accuracy of 84.92%, which is an effective AD detection method that can help clinicians to quickly diagnose Alzheimer's disease based on PSI brain network topology analysis and feature selection optimization algorithm.
Fatigue driving is a major contributor to traffic accidents, as it reduces alertness and can even be fatal. To investigate alertness changes during prolonged driving, we first built a simulated driving experiment platform and designed an alertness detection experiment. Electroencephalogram (EEG) signals were collected during prolonged fatigue driving from 16 subjects. For each subject, the driving process was divided into eight stages from T0-T7. Then, the alertness level was analyzed by temporal complex network method based on multiplex limited penetrable visibility graphs from the collected EEG signals. The clustering coefficient, path length and global efficiency of the constructed complex network were extracted for each driving stage under five typical brain rhythms including delta, theta, alpha, beta, and gamma. The results indicate that after a long stage of driving, the clustering coefficient and path length show an overall increasing trend, while the global efficiency reveals an overall decreasing trend. Our findings may provide indicators for monitoring alertness level.
Fatigue driving is an important cause of traffic accidents. Compared with other traffic accidents, traffic accidents caused by fatigue driving have a higher fatality rate and greater damage to the surrounding environment. Therefore, many cars manufactured by automobile companies are equipped with driver assistance functions as driver assistance, and many third-party companies are also following up to manufacture fatigue detection equipment. Although many companies and researchers have made a lot of achievements in the field of fatigue detection, there are still many areas for improvement in the research of fatigue detection. The methods of fatigue detection on the market are generally divided into invasive detection and non-invasive detection. In this paper, a non-invasive fatigue detection system based on deep learning is proposed. The system conducts face detection by YOLOv5 algorithm. At the same time, 12 key points in the eyes and 5 key points in the mouth are added to the recognized faces to detect facial expression changes, and on this basis, PERCLOS algorithm is used as the judgment basis. The fatigue state of the current face is evaluated.When epoch=80, mAP@0.5(mean accuracy) can reach 0.995, mAP@0.5:.95(the weighted average of mean accuracy under different IoU thresholds, generally speaking, the higher the accuracy, the better), also reaches 0.81.
The existence of asymptomatic carriers, individuals infected without showing symptoms, exacerbates the spread of infectious diseases, especially when people migrate between communities, leading to large-scale outbreaks. Therefore, we established the SAIR infectious disease model on a network to investigate the impact of individual migration on disease transmission in the presence of asymptomatic carriers. Based on this model, we studied how to control the spread of the epidemic when a community gets infected and migrates between communities. Consequently, we calculated the threshold value R 0 , and found that when R 0 < 1, the disease-free equilibrium point exists and is stable. Furthermore, we proved its global asymptotic stability using the decomposition theorem, and demonstrated the global stability of the local disease equilibrium point by constructing a Lyapunov function.
Keyword spotting (KWS) plays a crucial role in enabling voice-based user interactions on smart devices. However, conventional KWS methods require a large number of predefined keywords to achieve acceptable detection accuracy, which users may find challenging to provide. In recent years, self-supervised training and large models have excelled in various audio tasks. Their general audio feature extraction capabilities align well with the low-resource nature and scalability requirements of KWS tasks. In this paper, we integrate self-supervised models with keyword transformers to tailor them for KWS tasks. Experiments show that our approach significantly outperforms previous supervised methods. Moreover, our method’s advantages become even more pronounced under extremely limited resource conditions, which is of great importance for the rapid deployment of KWS systems.
To address the lack of fundamental physiological parameters assessment for hand function rehabilitation training devices, this paper designs an Android based smartphone application for acquisition of human hand physiological parameters, which can communicate and control hardware detection devices. The hand function can be evaluated by precisely measuring four physiological parameters, namely skin temperature, blood oxygen saturation, blood flow perfusion index n , and nerve conduction velocity within each finger ioint. The results may provide objective, quantitative basis for personalized rehabilitation treatment. Additionally, the software incorporates a database of patient information, facilitating long-term monitoring by healthcare professionals and allowing them to easily track the hand function status of patients, thereby assisting to improve treatments efficacy.