3-D magnetic recording (3DMR) technologies, such as 3-D heat-assisted magnetic recording (3-D HAMR) and 3-D bit-patterned magnetic recording (3-D BPMR), enhance areal density by stacking multiple recording layers and utilizing heat-assisted writing to encode multi-bit data through vertically combined magnetization states. Current 3DMR systems primarily employ dual-layer media [also referred to as dual-layer magnetic recording (DLMR)], where the readback signals from a single head consist of superimposed responses from both the top and bottom layers, leading to significant 3-D interference, including inter-symbol interference (ISI), inter-track interference (ITI), and inter-layer interference (ILI). To improve bottom-layer detection reliability in 3DMR, this article proposes a scaling-based signal extraction method that refines the bottom-layer signal using detected top-layer data. This approach enables efficient 1-D detection while significantly reducing the bottom-layer bit error rate (BER). Compared to the previously proposed dual-layer PRML detection, the method maintains the same top-layer BER while achieving an 80% reduction in the bottom-layer BER at the areal density of 3.3 Tb/in(2 ). Moreover, it demonstrates superior bottom-layer detection performance across areal densities ranging from 1 to 4.5 Tb/in(2 )per layer (2-9 Tb/in(2 )for dual-layer systems).
Three-dimensional magnetic recording (3DMR) technologies, such as three-dimensional heat-assisted magnetic recording (3D HAMR) and three-dimensional bit-patterned magnetic recording (3D BPMR), enhance areal density by stacking multiple recording layers and leveraging heat-assisted writing to encode multi-bit data through vertically combined magnetization states. Current 3DMR systems primarily employ dual-layer media (also termed dual-layer magnetic recording), where the readback signal from a single head comprises superimposed responses from a top layer and a bottom layer, introducing severe three-dimensional interference—including inter-symbol interference (ISI), inter-track interference (ITI), and inter-layer interference (ILI). To improve bottom-layer detection reliability in 3DMR, this paper proposes a scaling-based signal extraction method that refines the bottom-layer signal using detected top-layer data. This approach enables efficient one-dimensional detection while significantly reducing the bottom-layer bit error rate (BER). Compared to the previously proposed dual-layer PRML detection, the method maintains the same top-layer BER while achieving a 80% reduction in the bottom-layer BER under the areal density of 3.3 Tbpsi. Moreover, it demonstrates superior bottom-layer detection performance across areal densities ranging from 1 to 4.5 Tbpsi per layer (2–9 Tbpsi for dual-layer systems).
In real-world physiological and psychological scenarios, there often exists a robust complementary correlation between audio and visual signals. Audio-Visual Event Localization (AVEL) aims to identify segments with Audio-Visual Events (AVEs) that contain both audio and visual tracks in unconstrained videos. Prior studies have predominantly focused on audio-visual cross-modal fusion methods, overlooking the fine-grained exploration of the cross-modal information fusion mechanism. Moreover, due to the inherent heterogeneity of multi-modal data, inevitable new noise is introduced during the audio-visual fusion process. To address these challenges, we propose a novel Cross-modal Contrastive Learning Network (CCLN) for AVEL, comprising a backbone network and a branch network. In the backbone network, drawing inspiration from physiological theories of sensory integration, we elucidate the process of audio-visual information fusion, interaction, and integration from an information-flow perspective. Notably, the Self-constrained Bi-modal Interaction (SBI) module is a bi-modal attention structure integrated with audio-visual fusion information, and through gated processing of the audio-visual correlation matrix, it effectively captures inter-modal correlation. The Foreground Event Enhancement (FEE) module emphasizes the significance of event-level boundaries by elongating the distance between scene events during training through adaptive weights. Furthermore, we introduce weak video-level labels to constrain the cross-modal semantic alignment of audio-visual events and design a weakly supervised cross-modal contrastive learning loss (WCCL Loss) function, which enhances the quality of fusion representation in the dual-branch contrastive learning framework. Extensive experiments conducted on the AVE dataset for both fully supervised and weakly supervised event localization, as well as Cross-Modal Localization (CML) tasks, demonstrate the superior performance of our model compared to state-of-the-art approaches.
Hypernetworks, or hypernets in short, are neural networks that generate weights for another neural network, known as the target network. They have emerged as a powerful deep learning technique that allows for greater flexibility, adaptability, dynamism, faster training, information sharing, and model compression etc. Hypernets have shown promising results in a variety of deep learning problems, including continual learning, causal inference, transfer learning, weight pruning, uncertainty quantification, zero-shot learning, natural language processing, and reinforcement learning etc. Despite their success across different problem settings, currently, there is no review available to inform the researchers about the developments and to help in utilizing hypernets. To fill this gap, we review the progress in hypernets. We present an illustrative example to train deep neural networks using hypernets and propose categorizing hypernets based on five design criteria as inputs, outputs, variability of inputs and outputs, and architecture of hypernets. We also review applications of hypernets across different deep learning problem settings, followed by a discussion of general scenarios where hypernets can be effectively employed. Finally, we discuss the challenges and future directions that remain under-explored in the field of hypernets. We believe that hypernetworks have the potential to revolutionize the field of deep learning. They offer a new way to design and train neural networks, and they have the potential to improve the performance of deep learning models on a variety of tasks. Through this review, we aim to inspire further advancements in deep learning through hypernetworks.
Tetanus, a life-threatening bacterial infection prevalent in low- and middle-income countries like Vietnam, impacts the nervous system, causing muscle stiffness and spasms. Severe tetanus often involves dysfunction of the autonomic nervous system (ANS). Timely detection and effective ANS dysfunction management require continuous vital sign monitoring, traditionally performed using bedside monitors. However, wearable electrocardiogram (ECG) sensors offer a more cost-effective and user-friendly alternative. While machine learning-based ECG analysis can aid in tetanus severity classification, existing methods are excessively time-consuming. Our previous studies have investigated the improvement of tetanus severity classification using ECG time series imaging. In this study, our aim is to explore an alternative method using ECG data without relying on time series imaging as an input, with the aim of achieving comparable or improved performance. To address this, we propose a novel approach using a 1D-Vision Transformer, a pioneering method for classifying tetanus severity by extracting crucial global information from 1D ECG signals. Compared to 1D-CNN, 2D-CNN, and 2D-CNN + Dual Attention, our model achieves better results, boasting an F1 score of 0.77 ± 0.06, precision of 0.70 ± 0. 09, recall of 0.89 ± 0.13, specificity of 0.78 ± 0.12, accuracy of 0.82 ± 0.06 and AUC of 0.84 ± 0.05.
One approach considered promising for advancing the development of magnetic data storage systems is three-dimensional magnetic recording. But when the recording density rises, media noise becomes a bigger problem for magnetic recording technology. Machine learning algorithms have been proven to outperform traditional numerical techniques in solving the magnetic recording signal quality issue. We compared the detection bit error rate (BER) performance of a per-layer shallow neural network equalizer and a quaternary neural network equalizer for two recording layers. For a two-layer recording density of 4 Tb/in 2 , the results show that the combined quaternary neural network equalization and detection scheme outperforms the per-layer shallow neural network equalization and one-dimensional Viterbi detection and outperforms the conventional two-dimensional generalized partial response equalizer combined with a quaternary Viterbi detector, especially for the bottom layer.
Three-dimensional magnetic recording (3DMR) is a crucial technology for significantly increasing the storage capacity of hard disk drives (HDDs). However, the presence of inter-symbol interference (ISI), intertrack interference (ITI), and interlayer interference (ILI) poses significant challenges to the accurate detection of data stored on multiple layers. This study addresses the impact of interlayer misalignment on the bit error rate (BER) performance in 3DMR systems. We evaluate the performance based on a neural network estimator for reconstructing the top-layer read response signal using feedback from a Viterbi detector. This enables the separation of bottom-layer signals by subtraction from the mixed readback signal. We introduce a dual-layer partial response maximum likelihood (PRML) detector for simultaneous bit retrieval from both layers. Furthermore, we investigate methods of a per-layer binary classifier and a dual-layer four-class classifier based on neural networks. Our study demonstrates the BER performance of these detection schemes influenced by the interlayer misalignment, especially when the offset is 0, 10%, 50%, and 90% of the bit dimensions between the two layers. The results show that the neural network-based reconstruction and separation method achieves better bottom-layer BER performance under a slight interlayer misalignment. The BER performance benefits more from a mild downtrack offset than the crosstrack offset. The neural network-based separation detection and the dual-layer PRML achieve the lowest top-layer BER and the worst bottom-layer BER when the interlayer misalignment is half a bit.
The global datasphere is experiencing significant expansion, largely driven by the relentless advancement of artificial intelligence technology. Cold data, in particular, is garnering increasing attention and is now considered a vital asset, constituting a substantial portion of data capacity. Optical discs stand out as a highly promising solution for cold data archiving due to their robust data security, longevity, durability, and non-volatile nature. Additionally, their total cost of operation (TCO) further enhances their appeal. To increase the recording density, signal processing techniques must evolve along with the degradation of signal quality. In this paper, we conduct a comprehensive study of partial response maximum likelihood (PRML) detection based on simulation and real signal testing and compare an adaptive equalized detection scheme with a neural network equalized detection method for BD-ROM and BDXL systems. Results show the effectiveness of PRML method for both BD-ROM and BDXL. The neural network equalized detection method can also improve the detection performance of BD-ROM and BDXL at lower signal-to-noise ratio (SNR) levels.
In this study, we present a middleware-based approach for detecting anomalies in distributed systems. Our method facilitates the dynamic collection of logs at various levels of detail and incorporates an a priori dictionary-based compression strategy to process and transmit logs, minimizing the performance impact on distributed systems. Additionally, we utilize a dual feature fusion technique to analyze the logs. We evaluate the effectiveness of our approach by performing anomaly detection in a publish/subscribe distributed system and comparing it with existing methods. The results illustrate that our method outperforms other approaches, demonstrating superior performance.
To achieve higher magnetic recording areal densities, signal distortions caused by the crosstalk among consecutive magnetized transitions that lead to nonlinear transition shifts (NLTS) must be eliminated. We calculated and sorted the transition shifts of different recorded patterns and classified corresponding patterns based on an analytical method. We proposed a graded write precompensation scheme according to the classified patterns and compared the detecting performance of a perpendicular magnetic recording system using the graded compensation and traditional pattern-aware compensation during the writing process. Simulation results show our proposed graded precompensation can effectively mitigate the nonlinear distortion effect caused by NLTS and outperforms the traditional pattern-aware compensation method. The detecting BER decreases by 14.7% when the SNR of electronic noise is only 20 dB and the bit length is 10 nm.
Audio-visual event (AVE) localization aims to detect whether an event exists in each video segment and predict its category. Only when the event is audible and visible can it be recognized as an AVE. However, sometimes the information from auditory and visual modalities is asymmetrical in a video sequence, leading to incorrect predictions. To address this challenge, we introduce a dynamic interactive learning network designed to dynamically explore the intra- and inter-modal relationships depending on the other modality for better AVE localization. Specifically, our approach involves a dynamic fusion attention of intra- and inter-modalities module, enabling the auditory and visual modalities to focus more on regions deemed informative by the other modality while focusing less on regions that the other modality considers noise. In addition, we introduce an audio-visual difference loss to reduce the distance between auditory and visual representations. Our proposed method has been demonstrated to have superior performance by extensive experimental results on the AVE dataset. The source code will be available at https://github.com/hanliang/DILN .
Methods based on convolutional neural networks have achieved excellent performance in the image dehazing task. Unfortunately, most of the dehazing methods that exist suffer from loss of detail in the convolution and activation operations and failure to consider the effects of superimposing different intensities of haze, such as under-exposed and over-exposed images. To address these issues, we propose a dynamic dehazing convolution (DDC) based on attentional weight calculation and dynamic weight fusion and a dynamic dehazing activation (DDA) based on the input global context encoding function to address the problem of detail loss. And we propose a multi-scaled feature-fused image dehazing network (MFID-Net) based on DDC and DDA to address the effects of haze superposition. We also design a loss function based on the physical model with dynamic weights. Extensive experimental results demonstrate that the proposed MFID-Net performs favorably against the state-of-the-art algorithms on the hazy dataset while improving further on hazy images with large differences in haze concentration, and producing satisfactory dehazing results. The code is available at https://github.com/awhitewhale/MFID-Net.
Gestational diabetes mellitus (GDM) is a subtype of diabetes that develops during pregnancy. Managing blood glucose (BG) within the healthy physiological range can reduce clinical complications for women with gestational diabetes. The objectives of this study are to (1) develop benchmark glucose prediction models with long short-term memory (LSTM) recurrent neural network models using time-series data collected from the GDm-Health platform, (2) compare the prediction accuracy with published results, and (3) suggest an optimized clinical review schedule with the potential to reduce the overall number of blood tests for mothers with stable and within-range glucose measurements. A total of 190,396 BG readings from 1110 patients were used for model development, validation and testing under three different prediction schemes: 7 days of BG readings to predict the next 7 or 14 days and 14 days to predict 14 days. Our results show that the optimized BG schedule based on a 7-day observational window to predict the BG of the next 14 days achieved the accuracies of the root mean square error (RMSE) = 0.958 ± 0.007, 0.876 ± 0.003, 0.898 ± 0.003, 0.622 ± 0.003, 0.814 ± 0.009 and 0.845 ± 0.005 for the after-breakfast, after-lunch, after-dinner, before-breakfast, before-lunch and before-dinner predictions, respectively. This is the first machine learning study that suggested an optimized blood glucose monitoring frequency, which is 7 days to monitor the next 14 days based on the accuracy of blood glucose prediction. Moreover, the accuracy of our proposed model based on the fingerstick blood glucose test is on par with the prediction accuracies compared with the benchmark performance of one-hour prediction models using continuous glucose monitoring (CGM) readings. In conclusion, the stacked LSTM model is a promising approach for capturing the patterns in time-series data, resulting in accurate predictions of BG levels. Using a deep learning model with routine fingerstick glucose collection is a promising, predictable and low-cost solution for BG monitoring for women with gestational diabetes.
Tetanus is a life-threatening infectious disease, which is still common in low- and middle-income countries, including in Vietnam. This disease is characterized by muscle spasm and in severe cases is complicated by autonomic dysfunction. Ideally continuous vital sign monitoring using bedside monitors allows the prompt detection of the onset of autonomic nervous system dysfunction or avoiding rapid deterioration. Detection can be improved using heart rate variability analysis from ECG signals. Recently, characteristic ECG and heart rate variability features have been shown to be of value in classifying tetanus severity. However, conventional manual analysis of ECG is time-consuming. The traditional convolutional neural network (CNN) has limitations in extracting the global context information, due to its fixed-sized kernel filters. In this work, we propose a novel hybrid CNN-Transformer model to automatically classify tetanus severity using tetanus monitoring from low-cost wearable sensors. This model can capture the local features from the CNN and the global features from the Transformer. The time series imaging - spectrogram - is transformed from one-dimensional ECG signal and input to the proposed model. The CNN-Transformer model outperforms state-of-the-art methods in tetanus classification, achieves results with a F1 score of $\mathbf {0.82\pm 0.03}$ , precision of $\mathbf {0.94\pm 0.03}$ , recall of $\mathbf {0.73\pm 0.07}$ , specificity of $\mathbf {0.97\pm 0.02}$ , accuracy of $\mathbf {0.88\pm 0.01}$ and AUC of $\mathbf {0.85\pm 0.03}$ . In addition, we found that Random Forest with enough manually selected features can be comparable with the proposed CNN-Transformer model.
Tetanus is a life-threatening bacterial infection that is often prevalent in low- and middle-income countries (LMIC), Vietnam included. Tetanus affects the nervous system, leading to muscle stiffness and spasms. Moreover, severe tetanus is associated with autonomic nervous system (ANS) dysfunction. To ensure early detection and effective management of ANS dysfunction, patients require continuous monitoring of vital signs using bedside monitors. Wearable electrocardiogram (ECG) sensors offer a more cost-effective and user-friendly alternative to bedside monitors. Machine learning-based ECG analysis can be a valuable resource for classifying tetanus severity; however, using existing ECG signal analysis is excessively time-consuming. Due to the fixed-sized kernel filters used in traditional convolutional neural networks (CNNs), they are limited in their ability to capture global context information. In this work, we propose a 2D-WinSpatt-Net, which is a novel Vision Transformer that contains both local spatial window self-attention and global spatial self-attention mechanisms. The 2D-WinSpatt-Net boosts the classification of tetanus severity in intensive-care settings for LMIC using wearable ECG sensors. The time series imaging—continuous wavelet transforms—is transformed from a one-dimensional ECG signal and input to the proposed 2D-WinSpatt-Net. In the classification of tetanus severity levels, 2D-WinSpatt-Net surpasses state-of-the-art methods in terms of performance and accuracy. It achieves remarkable results with an F1 score of 0.88 ± 0.00, precision of 0.92 ± 0.02, recall of 0.85 ± 0.01, specificity of 0.96 ± 0.01, accuracy of 0.93 ± 0.02 and AUC of 0.90 ± 0.00.
Snow is a harsh natural phenomenon that greatly affects the performance of advanced computer vision tasks. Recently, most image desnowing methods rely on complex model structures, leading to increased carbon emissions and impossibilities to deploy on lightweight computing devices. Additionally, most deep learning-based methods do not effectively utilize the position information of snow particles, which will limit the performance of the model. To address these issues, considering that snow particles have block-shaped shape distribution and color uniformity similar to a mask, we propose a snowed autoencoder (SAE) desnowing method. Specifically, the SAE is composed of four parts: snowed masking process, SAE encoder, SAE decoder, and prediction stage. First, a snowy image is passed through the snowed masking process to generate a mask that exactly covers the snow particles and output the coordinates of the four vertices of the mask and the index of each segmented image patch. The input image and each mask coordinate and index are passed through the SAE encoder based on the snow particle attention module to generate an image representation for identifying typical features of snow particles. Then, the generated typical features of snow particles are fed into the SAE decoder to reconstruct image patches without snow. Finally, all the patches will be passed through the prediction stage to reconstruct the original size of the snow-free image. A large number of experiments show that the proposed SAE desnowing method achieves the state-of-the-art desnowing performance on three synthetic and one real-world desnowing datasets. The images after desnowing by SAE have better color details and are more consistent with human visual habits. In addition, SAE has faster desnowing speed and fewer parameters. The code is available at https://github.com/awhitewhale/SAE.
With the development of smart surveillance systems, multimodal violence video data provides enough data source for smart surveillance systems. At present, intelligent violence detection faces two main challenges. One of the important challenges to apply weakly supervised learning for violence detection is to accurately identify normal segments of abnormal videos, and another challenge is how to make full use of visual and audio features of videos. Therefore, we explored methods for fusing visual and audio information together with temporal information. We propose a novel neural network containing three parts: 1) co-attention module fusing audio and video features with LSTM to extract temporal information, 2) mutual learning branches with threshold to generate high-quality pseudo labels, and 3) a simple and effective post-processing method to ensure continuity of forecast results. Our experiment results show that the proposed model exceeds the existing state-of-art models on the XD-Violence dataset by 1.92% in AP.
Deep learning-based methods have achieved excellent performance in image-deraining tasks. Unfortunately, most existing deraining methods incorrectly assume a uniform rain streak distribution and a fixed fine-grained level. And this uncertainty of rain streaks will result in the model not being competent at repairing all fine-grained rain streaks. In addition, some existing convolution-based methods extend the receptive field mainly by stacking convolution kernels, which frequently results in inaccurate feature extraction. In this work, we propose momentum-contrast and large-kernel for multi-fine-grained deraining network (MOONLIT). To address the problem that the model is not competent at all fine-grained levels, we use the unsupervised dictionary contrastive learning method to treat different fine-grained rainy images as different degradation tasks. Then, to address the problem of inaccurate feature extraction, we carefully constructed a restoration network based on large-kernel convolution with a larger and more accurate receptive field. In addition, we designed a data enhancement method to weaken features other than rain streaks in order to be better classified for different degradation tasks. Extensive experiments on synthetic and real-world deraining datasets show that the proposed method MOONLIT achieves the state-of-the-art performance on some datasets. Code is available at https://github.com/awhitewhale/moonlit .
Cross-camera pedestrian tracking continuously tracks pedestrians in a monitoring network composed of multiple cameras. Due to the large degree of freedom of the human body and the influence of the environment variations, the pedestrian recognition is challenging for such tasks. We propose a method of dynamically matching local features of horizontal and vertical stripes (DFM+) to automatically align pedestrian features and obtain more stable pedestrian appearance features. In addition, we design a three-branched local-global dynamic feature matching network (3bDFM-Net) framework including the person orientation, and the local and global branches for cross-camera multi-person re-identification (ReId). Experiment results show that the accuracy of cross-camera multi-person tracking is improved by introducing the pedestrian orientation. The Rank-1 accuracies on the Market1501 and DukeMTMCReID datasets reach 96.2% and 90.1%, respectively. Further tests verified its capabilities of cross-camera tracking of pedestrians in the surveillance video of real scenes.
Nonlinear transition shift (NLTS) produced by the interaction between transitions is a major source of distortions in magnetic recording. NLTS must be reduced to realize higher density and maintain signal quality. However, there is some confusion about the production and behavior of NLTS in previous reports. In this paper, combined with micromagnetic simulation, we have developed a complete theoretical model to predict NLTS on perpendicular magnetic media. The results show that one-level NLTS (caused by one previous transition) will monotonically decrease as the present transition becomes further away from the previous one, and multilevel NLTS (caused by consecutive previous transitions) will firstly increase and quickly approach a constant value as there are more consecutive transitions. Particularly, it is demonstrated that NLTS and its impact on the media signal-to-noise ratio can be significantly limited by pattern constraints.