Weight initialization plays a decisive role in the success of deep neural networks, yet existing strategies remain general-purpose and overlook the structural properties of time series. To address this gap, we introduce the Wavelet-Enhanced Adaptive Network (WEANet), a modular architecture that embeds wavelet-based inductive bias directly into its initial parameters. By initializing convolutional kernels with coefficients from multiple wavelet families, WEANet transforms convolutional layers into adaptive multi-resolution analyzers, uniting the interpretability of classical signal processing with the flexibility of deep learning. To preserve this structured initialization, we propose a dual-objective loss with a reconstruction term that regularizes training and safeguards the integrity of wavelet-induced representations. Extensive experiments on classification, forecasting, imputation, and anomaly detection demonstrate that WEANet consistently outperforms state-of-the-art baselines, achieving the best results on 24 of 30 UEA benchmark datasets and delivering up to 5% lower error in long-term forecasting. We further show its effectiveness as a plug-and-play tokenizer for Transformer backbones and validate each component through in-depth ablation studies. Our findings highlight wavelet-guided initialization as a general paradigm for principled, domain-aware deep time series modeling.
As a core component of intelligent transportation systems, vehicle localization technology enables accurate positioning, supporting comprehensive insights into traffic flow, vehicle status, and environmental changes. Current vehicle position localization technologies primarily rely on visual sensors and the Global Positioning System, while their performance can be affected by extreme weather conditions and signal stability. However, vehicle-generated sound, as a stable data source unaffected by environmental conditions and free from signal limitations, is often underutilized by existing studies. In this paper, we construct a vehicle localization dataset based on sound signals and further propose a combined filtering strategy that integrates adaptive filtering with spectral subtraction filtering, dynamically adjusting the filter parameters to suppress time-correlated noise within the signal. We also remove broadband noise in the frequency domain while preserving high-frequency signal details, offering a significant advantage over existing methods in terms of signal-to-noise ratio improvement. The proposed dataset and filtering strategy are validated using the EfficientNet-1D Fusion model. Experimental results demonstrate that the proposed combined filtering method excels in recognition accuracy and computational efficiency.
Tabular data supports intelligent decision-making in key areas such as finance and healthcare, but its heterogeneous feature interaction modeling and large-scale computing efficiency issues have long restricted the application of deep learning technology. This paper proposes SparseLinTab, an efficient tabular data modeling framework based on bidirectional sparse linear self-attention. By decoupling the interactions between rows (sample level) and columns (feature level), SparseLinTab reduces the quadratic computational complexity of traditional Transformer self-attention O(N2M + NM2) to linear O(NM), while retaining the ability to model global dependencies. SparseLinTab outperforms all traditional gradient boosting models and deep learning models in terms of average performance across 7 public datasets. Specifically, the core metrics are calculated separately for each dataset (accuracy is used for classification tasks, and MAE for regression tasks), and then the simple arithmetic mean of the metric scores from all datasets is computed to balance the impact of different scenario characteristics on the overall performance evaluation. In addition, ablation experiments and attention visualization show that the sparse mechanism significantly enhances the robustness and generalization of the model by filtering noise interactions and focusing on key feature combinations. (c) 2025 The Author(s). Published by Elsevier B.V. on behalf of Shandong University. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
The rapid growth of the Internet of Things (IoT) has extended into marine environments, giving rise to the Internet of Underwater Things (IoUT). Within this context, Underwater Named Data Networking (UNDN) has emerged as a promising paradigm, yet it faces severe challenges such as limited bandwidth, long delays, energy scarcity, etc. We propose an Energy-Balanced Multi-Path Interest Forwarding (EBMIF) strategy for UNDN. The strategy has three parts. First, a depth-based name discovery builds a small set of candidate paths and uses a budget to bound diffusion. Second, each packet piggybacks the sender’s residual-energy ratio, which gives neighbors passive and timely energy visibility without extra control packets. Third, a lexicographic rule selects next hops sequentially by residual energy, depth progress, and path cost. EBMIF is implemented in Aqua-Sim-NG and compared with CBILEM-U, BestRoute, and Dual-Mode Interest Forwarding (DMIF). Experimental results illustrate that the EBMIF strategy demonstrates superior performance in energy balancing, network lifetime optimization, and overall operational efficiency, making it well suited for long-term and stable underwater network deployments.
As space communication technology advances at a rapid pace, integrated networks spanning air, space, land, and sea have demonstrated significant potential applications in various fields. However, traditional Transmission Control Protocol/Internet Protocol (TCP/IP) may not cope well with the high dynamics of satellite networks. Thus, the virtual information-defined satellite network (VIDSN) architecture is proposed, adopting the virtual node strategy. The pending interest table (PIT) and forwarding information base (FIB) are redesigned accordingly, and the FIB table is updated through controller polling. To realize efficient data transmission in a dynamic satellite network, an appropriate routing strategy is essential. A novel multilayer satellite routing (MLSR) algorithm is proposed for satellite IoT networks, which adopts a hierarchical routing control architecture to enable scalable and coordinated routing. By exploiting the hierarchical structure of multilayer satellite networks, the proposed MLSR algorithm effectively supports dynamic topology and traffic variations. Simulation results demonstrate that MLSR comprehensively outperforms baseline methods.
Distributed optical fiber sensing systems based on phase-sensitive optical time domain reflectometry are increas ingly playing a critical role in long-range mobile sensing within resource-constrained environments. However, the inherent medium dependency and spatial energy distribution of DOFS signals generate complex latent spa tiotemporal patterns that conventional models fail to represent effectively, especially for low-amplitude human activity events. To address these challenges, this paper first proposes a dual-threshold preprocessing approach specifically designed for fiber-optical data, which rapidly filters high-value spatiotemporal features while pre serving their intrinsic dependencies. Subsequently, Hierarchical Multi-scale Spatiotemporal Adaptive Network (HMSAnet) is proposed to process the rearranged data through a hierarchical design that captures distinctive spatiotemporal representations of fiber-optical data, and preserves spatial integrity. The multi-scale dynamic ar chitecture of HMSAnet further enables adaptive integration of fine-grained details from small receptive fields with long-term regularities from large receptive fields. Furthermore, a five-class human activity dataset is constructed, which effectively reflects the influence of diverse human activities on fiber-optical signals. Experimental results on the self-constructed dataset demonstrate that the proposed method achieves a 99.54 % detection accuracy, while maintaining an effective balance between computational efficiency and accuracy, outperforming existing state-of-the-art methods across multiple evaluation metrics.
Contrastive representation learning has gained prominence in multivariate time series classification (MTSC) due to its ability to leverage unlabeled data. Existing methods commonly propose consistency assumptions that focus solely on temporal or contextual features, which are inadequate for capturing the complex patterns and shifts found in real-world time-series data across diverse domains. In this paper, we present a novel contrastive learning approach for MTSC that introduces an innovative strategy for positive sample generation, leveraging the temporal-frequency consistency that naturally exists in diverse time-series data. Unlike existing methods that rely on heuristic temporal augmentations, which can inadvertently distort the underlying semantics of non-stationary signals, our approach leverages inherent structural and rhythmic properties of the original data. Specifically, we apply the Discrete Wavelet Transform (DWT) to extract wavelet coefficients, which represent time-frequency information and serve as positive samples. To improve negative sample selection, we develop a distance-based time series bank to minimize the risk of selecting similar samples from the same class as negative pairs. Furthermore, we introduce hierarchical-connected convolutional filter-based encoders to map samples and their wavelet coefficients into a joint time-frequency space, enabling hierarchical contrastive learning based on multi-scale temporal-frequency consistency. Extensive evaluations on 30 UEA datasets and 6 industrial sensing datasets show that our method achieves a state-of-the-art average accuracy of 0.777 and a top rank of 1.367, outperforming several competitive baselines. These results, supported by rigorous statistical analysis, confirm the effectiveness of our framework for MTSC.
WiFi-based human activity recognition (HAR) plays a pivotal role in applications such as elderly care, health monitoring, and smart home systems. Unlike the traditional IEEE 802.11n that rely on full channel state information (CSI), modern WiFi standards, including IEEE 802.11ac/ax, use compressed beamforming reports (CBR) instead of exchanging CSI between commercial off-the-shelf (COTS) routers and wireless devices. This partial CSI poses significant challenges for activity recognition. In this letter, we introduce CBR-HAR, a real-time activity recognition system leveraging WiFi signals compliant with the IEEE 802.11ac/ax standards. CBR-HAR consists of a CBR capture and decoding module, a time-domain and frequency-domain feature extraction module, and an activity recognition module. In the activity recognition module, we propose a dual-branch residual network to effectively utilize both the time-domain and frequency-domain information of the CBR to classify human activities. Additionally, to better distinguish between falls and similar activities such as standing up and sitting down, we integrate a BERT-based semantic embedding loss. Extensive evaluations show that CBR-HAR achieves an impressive accuracy of 90.56% in classifying six kinds of human activities.
The Internet of Underwater Things (IoUT) is essential for ocean research, but energy consumption is a major obstacle. Named Data Networking (NDN), with its in-network caching, offers a potential solution for efficient data transmission and reduced energy use in underwater environments. However, NDN's caching strategy can lead to a low cache hit ratio, high redundancy, resulting in energy waste and performance degradation. This paper proposes a novel caching strategy called Energy-Saving Caching (ESC), based on energy profit. Firstly, the new naming discovery package is used to announce and network the content to surrounding nodes, then the energy profit of the cached data is calculated through the content popularity and hop count information. When the interest arrives, the caching decision is made according to different profit information, aiming to reduce the overall energy consumption of the underwater network. Experimental results demonstrate that the proposed caching strategy effectively decreases the total energy consumption of the underwater network while improving the cache hit ratio and reducing the average retrieval delay.
Existing deep hashing methods mainly focus on preserving pairwise image similarity or reducing quantization error, often overlooking the discriminative capacity of real valued features learned by neural networks, which limits retrieval performance. To address these issues, we propose a dual -stream Collaborative Asymmetric Similarity -preserving Hashing (CASpII) deep hashing method that preserves semantic structure across categories while generating discriminative hash codes. Specifically, the Cross Attention Feature Enhancement. Block (CAFEB) is designed to mitigate information loss from feature dimensionality reduction during extraction. Furthermore, two asymmetric deep networks are constructed to capture image similarity based on semantic labels. To ensure binary codes in Hamming space retain semantic similarity from the original space, an asymmetric loss is introduced to capture the similarity between binary codes and real-valued features. This asymmetric loss not only enhances retrieval performance but also aids in faster convergence during training. Extensive experiments on three benchmark datasets demonstrate that the CASp1I method outperforms other comparative methods.
With the breakthrough of convolutional neural networks, deep hashing methods have demonstrated remarkable performance in large-scale image retrieval tasks. However, existing deep supervised hashing methods, which rely on pairwise or triplet labels, typically learn the hash function via random or hardest sample mining within training batches. This strategy primarily captures local sample similarities, causing a distribution shift and limiting retrieval performance. Furthermore, most methods emphasize global features while overlooking structural information, which is essential for understanding spatial relationships in images. To solve these limitations, we propose a Centripetal Intensive Deep Hashing (CIDH) method for remote sensing image retrieval. Initially, we design a Hybrid-Attention Guided Multiscale Refinement Network that integrates channel and spatial attention to capture multiscale visual features and highlight salient regions at different scales. Subsequently, we introduce a central similarity loss via class-centered labels to optimize the spatial distribution of global samples, which can encourage hash codes with similar semantics to cluster around centroids and reduce distribution shift. Meanwhile, we incorporate a central intensive loss into the Hamming space to shorten intraclass Hamming distances, generating more compact and discriminative hash codes. Extensive experiments demonstrate the superiority of our CIDH method compared with current state-of-the-art deep hashing methods.
Forecasting short-term fishing effort distribution is crucial for fishery management to promote sustainable development and dynamically protect the ecosystem. Utilization of impact factors, such as marine hydrological factors and chlorophyll concentration distribution, can aid in short-term forecasting. However, two significant challenges emerge: firstly, the relationship between these impact factors and fishing effort distributions require comprehensive analysis; secondly, the forecasting model must effectively integrate all spatio-temporal features derived from these factors. In addressing these challenges, this study commences with a quantitative analysis of the relationships between these impact factors and fishing effort distributions. Subsequently, we introduce TransFish, a deep learning approach that forecasts day-level fishing effort distribution by harnessing these impact factors. TransFish integrates ResNet and Transformer network, seamlessly synthesizing spatial features from historical fishing effort distributions, marine hydrological factor fields, and chlorophyll concentration distributions, while also accounting for temporal relations within forecasting sequences. The performance of TransFish is evaluated using the Vessel Monitoring System dataset, which contains spatial and temporal fishing effort information from 1589 trawlers in the East China Sea, spanning September 2015 to May 2017. Additionally, ocean biogeochemical factors, such as sea surface temperature and chlorophyll concentration, are used as environmental forecasting variables. The dataset from September 2015 to May 2016 is utilized for relationship analysis and model training, while the dataset from September 2016 to May 2017 is employed to evaluate forecasting accuracy. The results reveal that the daily forecasting error ratio for the subsequent week ranges from 4.51 to 6.01
In recent years, the significant increase in traffic accidents caused by distracted driving has drawn widespread attention. As a result, identifying distracted drivers has become essential for improving road safety and advancing intelligent driver assistance systems. Deep learning has been used for real-time driver monitoring to detect risks like distraction, fatigue, and unsafe behaviors. To deploy CNNs on mobile devices, it is important to reduce number of parameters and computation while keeping detection accuracy high. In this paper we propose a lightweight convolutional neural network for distracted driver detection named Si-CA MobileNet. The model incorporates the Si-Block module based on the Simple, Parameter-free Attention Mechanism (SimAM) and the CA-Block module based on the Coordinate Attention (CA) mechanism. Experimental results show that Si-CA MobileNet achieves Top-1 accuracy of 99.80 % and 92.35 % on the StateFarm and 100-Driver datasets, respectively, surpassing most existing methods. Additionally, the model has only 3.17 M parameters and 175.95 M FLOPs. It also supports real-time processing at 82.95 FPS, making it well-suited for lightweight deployment scenarios.
With the proliferation of smartphone application services, the pattern lock remains widely used for authentication. Notably, the risks associated with password entry in public spaces have attracted significant attention from researchers. Various attacks have been explored to steal passwords, but each comes with limitations, such as requiring good lighting conditions, close proximity, pre-deployed devices, or system intrusion. To address these challenges, we propose an attack method called BEyes, which utilizes beamforming feedback information (BFI) to eavesdrop on pattern passwords drawn on smartphone screens. Since BFI is transmitted in clear text and describes the downlink channel state information (CSI), any Wi-Fi 5-enabled device can capture it out of the victim’s view, reducing the likelihood of the attack being detected. To avoid missing critical pattern drawing information, we propose a traffic generation mechanism based on traffic competition, which ensures stable BFI. To mitigate the effects of frequency-selective fading and noise, we apply subcarrier alignment and principal component analysis (PCA) to improve efficiency. Additionally, we introduce a motion-based joint inference model, enabling BEyes to generalize its inference from a few known pattern passwords to unknown ones. Extensive experiments demonstrate that BEyes achieves an accuracy of 89.2% in inferring a 3-line pattern password within the Top-10 attempts.
Transformers have made significant progress in dealing with computer vision tasks. However, most existing Vision Transformers (ViT) typically focus on single-scale information, limiting their capability to model interactions when processing multi-scale features. Moreover, the explicit mapping of continuous real-valued features to discrete hashing codes via a quantization layer is suboptimal for retrieval tasks. To overcome the above problem, we propose a novel deep hashing based on a Cross-scale Transformer named CrossHash, aiming to extract multi-scale features. Furthermore, a relative similarity quantization method is first introduced, which maximizes the similarity between the relative positional representation of continuous codes and the normalized centroids; this method effectively reduces quantization error. It optimizes feature distribution via contrastive learning loss, maximizing the inter-class distance and minimizing the intra-class distance of the learned feature. Extensive experiments on three benchmark datasets demonstrate that the proposed model outperforms other state-of-the-art deep hashing methods. Source code is available https://github.com/wwg1010/CrossHash.
Time series data, such as sound waves, bio-signals, and user trajectories, are prevalent in social application scenarios. While single-modal time series data often proves inadequate for addressing challenges in complicated environments, it necessitates integrating multiple modalities to understand real-world phenomena. Utilizing multimodal data improves deep learning systems’ effectiveness, generalizability, and robustness. This survey comprehensively reviews recent advancements, especially methodologies for general multimodal time series analysis and efforts in various social computing contexts, including autonomous driving, healthcare, audiovisual speech recognition, and gesture recognition. We highlight the key techniques of existing studies, namely modality selection, feature extraction, and information fusion and detail the solutions under various circumstances. Finally, we discuss the unresolved challenges and suggest potential future research directions. Our survey aims to provide researchers and industries insights into trends, behaviors, and preferences for performing multimodal time series analysis in social computing applications.
Recent advancements in deep learning (DL) have introduced new security challenges in the form of side-channel attacks. A prime example is the website fingerprinting attack (WFA), which targets anonymity networks like Tor, enabling attackers to unveil users' protected browsing activities from traffic data. While state-of-the-art WFAs have achieved remarkable results, they often rely on unrealistic single-website assumptions. In this paper, we undertake an exhaustive exploration of multi-tab website fingerprinting attacks (MTWFAs) in more realistic scenarios. We delve into MTWFAs and introduce MTWFA-SEG, a task involving the fine-grained packet-level classification within multi-tab Tor traffic. By employing deep learning models, we reveal their potential to threaten user privacy by discerning visited websites and browsing session timing. We design an improved fully convolutional model for MTWFA-SEG, which are enhanced by both network architecture advances and traffic data instincts. In the evaluations on interlocking browsing datasets, the proposed models achieve remarkable accuracy rates of over 68.6%, 71.8%, and 76.1% in closed, imbalanced open, and balanced open-world settings, respectively. Furthermore, the proposed models exhibit substantial robustness across diverse train-test settings. We further validate our designs in a coarse-grained task, MTWFA-MultiLabel, where they not only achieve state-of-the-art performance but also demonstrate high robustness in challenging situations.
3D Gaussian Splatting (3DGS) is gaining popularity in fields such as robotics, autonomous driving, and virtual reality, due to its effectiveness and efficiency. Given that some tasks involve high risks, it is crucial to investigate the adversarial robustness of 3DGS and its downstream tasks—a topic that remains largely unexplored. In this study, we introduce a framework, Multi-Parametric Adversarial Manipulation for 3D Gaussian Splatting (MPAM-3DGS), that allows to attack 3DGS and its downstream tasks, such as object detection and classification, by perturbing a specified subset of parameters. Leveraging this framework, we examine the adversarial sensitivity of each 3DGS parameter and propose two strategies to attack multiple parameters based on our observations. To our knowledge, this is the first study to explore the adversarial robustness of 3DGS. Our experimental results demonstrate the effectiveness of our attacks on downstream tasks and the invisibility of perturbations in 3DGS. The code can be found at https://github.com/jiang-wenxiang/MPAM-3DGS.
The symbiotic Internet of Things (IoT) computing paradigm leverages high-speed transmission technologies, such as 6G, to execute large-scale model computations while preserving data privacy, thereby mitigating significant economic losses from data breaches. This paradigm has emerged as a pivotal focus in industrial intelligent manufacturing research. However, within this framework, edge devices must concurrently support both large-scale model computations and high-precision industrial core tasks, creating critical challenges in system efficiency and stability that impede technological advancement. To address these issues, this article introduces a novel Quality of Experience (QoE)-driven heuristic task resource scheduling algorithm that employs Composite Differential Evolution (CoDE) integrated with a tabular Kolmogorov-Arnold Network (KANTab), specifically designed for the comprehensive processes of industrial intelligent manufacturing systems. This methodology enables precise simulation of extensive industrial task requirements and efficient allocation of computational resources under constrained conditions, effectively resolving the resource allocation conflict between large-scale model computation tasks at the edge and primary tasks on terminal devices. We evaluate our approach on a custom-built large-scale task demand dataset from refrigerator manufacturing and demonstrate that it achieves superior overall performance compared to state-of-the-art algorithms.
Detecting anomalies in multivariate time series data is crucial for various industries. However, the increasing volume and dimensionality of data have driven up the cost of data labeling, making supervised or semi-supervised methods less effective. Unsupervised learning methods, which do not require labeled training data, have steadily gained importance due to their ability to reduce data processing costs. This paper proposes an unsupervised multivariate time series anomaly detection method - Res2coder. It decomposes time series data into trend and residual components for separate reconstruction, and incorporates frequency-domain analysis. Res2coder uses an MLP-based autoencoder module to reconstruct each component. Additionally, two error feedback mechanisms are designed in the reconstruction module to enable the model to more accurately capture the features and changing patterns of the data. This approach improves detection accuracy and reduces model training costs without relying on complex networks like convolutions or attention mechanisms. We compare Res2coder with several baseline models (e.g., TranAD and ATF-UAD) on six datasets (e.g., SWaT, WADI, and SMD). It achieves higher scores in six evaluation metrics—precision, recall, F1, ROC/AUC, Composite F-score (Fc1), and Real Under Point Adjustment
Hanjiang Luo (罗汉江)合作论文数16