Accurate authentication of paper-based materials is crucial for secure document verification, counterfeit currency detection, and intelligent packaging across industries such as finance, logistics, and security. Differentiating between subtle variations, such as glossy, recycled, coated, or counterfeit paper, remains a significant challenge due to minimal material-level differences. We introduce PSense, the first millimeter-wave MIMO radar-based sensing framework for paper material analysis, offering a fully contactless solution that relies solely on reflected signals, eliminating the need for transmission-based or dual-sided setups. To characterize intrinsic paper properties, we propose a novel material-sensitive feature called the Energy Reflection Ratio (ERR), which encodes key physical interactions including surface reflectivity, internal attenuation, and dispersion, uniquely capturing the electromagnetic signature of each paper type. We further present Trans-PapNet, a dual-model transfer learning framework designed to enhance classification robustness and generalization. One model extracts spatial and spectral patterns from radar signal scalograms using ResNet-50, while the other processes ERR-based physical features. These complementary representations are fused via a shared attention mechanism, enabling effective domain adaptation across paper types. Our method achieves an overall classification accuracy of 97.85%, with strong performance in three knowledge-based transfer scenarios (94.8%, 93.0%, and 90.0%) and under both seen and unseen material conditions (95.3%). These results underscore PSense's robustness and scalability, paving the way for real-world deployment in mobile, secure, and intelligent paper-based authentication systems.
Maintaining focus is essential for carrying out tasks accurately and efficiently, but it can be difficult to do so in settings that are full of distractions and continuous change in daily life, the workplace, and education. Traditional assessment methods such as self-reporting, eye tracking, or camera-based observation are often intrusive, subjective, or limited by privacy concerns. To address these limitations, this study proposes a novel radar-based framework for continuous estimation of human attention during fine-grained hand-object interaction tasks using frequency-modulated continuous-wave (FMCW) millimeter-wave (mmWave) radar. We capture fine-grained Doppler-time motion patterns. The considered activities included pouring water, stacking cups, and writing, representing different motion types (translational, repetitive, and fine-motor), collected under focused and distracted conditions. A multi-input deep regression network is introduced, which combines handcrafted behavioral descriptors with temporal Doppler-time feature embeddings using a 1-D-convolutional neural network (CNN)-bidirectional long short-term memory (BiLSTM)-attention fusion pipeline. This network simultaneously encodes motion smoothness, energy dynamics, and temporal regularity to derive attention scores on a 0-100 scale. Extensive experiments show the model's capability to generalize across different subjects and activities, achieving an overall RMSE of approximate to 2.21 and a coefficient of determination (R-2 ) of 0.9959. Ablation analysis validates the needs of both handcrafted features and multihead-attention fusion for performance. signalto-noise ratios below 15 dB, noise tests revealed a performance decline of less than 5%. Changes in attention impact behavioral patterns were gained through correlation analysis. Our approach offers a privacy-preserving and contactless alternative to existing methods, making it ideal for use in classroom monitoring, workplace productivity assessment, and cognitive health evaluation.
We propose a non-contact, privacy-preserving emotion recognition framework using millimeter-wave (mm-Wave) radar and deep learning, addressing the limitations of traditional wearable and camera-based approaches. By broadcasting frequency-modulated radar pulses, the system isolates heart rate signals even in dynamic scenarios such as gameplay Fig. 1. The design integrates a hybrid 1D-CNN for efficient feature extraction and Bi-LSTM for temporal analysis, with a computational complexity of $O(N \cdot F + N \cdot H)$ , ensuring real-time capability. Validation through ROC curves, alongside F1-scores and precision-recall metrics ranging from 0.98 to 0.99, confirms the system's reliability. Unlike existing methods, this framework investigates the robustness of mm-wave radar to function independently of environmental factors like lighting or clothing, making it scalable for applications in healthcare, human-computer interaction, and educational settings. These findings establish mm-wave radar as a transformative tool for emotion recognition, offering enhanced comfort, privacy, and adaptability.
Camera-based facial emotion recognition (FER) suffers from poor lighting, occlusion, and privacy exposure, whereas mmWave-only solutions lack the spatial detail required for fine-grained affect analysis. To close this gap, we present a shifted-window transformer + long short-term memory emotion recognition (SwinLSTM-EmoRec). This noncontact dual-modal framework fuses micro-Doppler signatures captured by a TI IWR1443 mmWave radar with red-green-blue image modality (RGB) imagery while treating radar as the primary, identity-obscured source and adaptively limiting reliance on RGB. Privacy is preserved because the cross-attention gate downweights or bypasses RGB when illumination is poor or when potential identity exposure is detected, leaving decisions dominated by illumination-invariant radar dynamics. A shifted-window Swin Transformer extracts spatial facial cues, a long short-term memory (LSTM) models temporal radar dynamics, and a lightweight cross-attention layer aligns the two streams, boosting F-1 by up to 4% over early, late, and self-attention baselines. On a 50-participant interactive-gaming dataset recorded under varied lighting and distances of 0.5-2 m, the system achieves 98.5% accuracy ( F-1 approximate to 0.98 ). It maintains 33.9-ms end-to-end latency on a 15-W Jetson Xavier NX edge device. Performance remains > 92% at 2 m, demonstrating robust, privacy-preserving FER robust, privacy-aware emotion sensing suitable for smart-home, tele-health, and e-sports internet of things (IoT) applications.
In recent years, more and more people choose to work out at home or in the office to improve their physique and build muscle. However, the lack of professional guidance makes it difficult for many fitness practitioners to achieve optimal results. Consequently, research on noncontact fitness monitoring using wireless signals has gained attention. Existing studies primarily focus on isolated exercises, while compound exercises, which involve multiple isolated exercises, remain underexplored. This combination introduces new challenges for fitness recognition and monitoring. To this end, we propose M-Fitness, a millimeter-wave (mmWave) radar-based fitness assistant system capable of recognizing and monitoring both isolated and compound exercises. First, we capture fine-grained motion features and design image enhancement algorithms to generate intuitive motion images. Next, we design a novel motion segmentation method for fitness actions. We further formulate compound exercise recognition as a sequential task and develop customized deep learning models that allow users to incorporate new compound exercises. Finally, we perform a comprehensive fitness assessment based on the frequency, intensity, time, and type (FITT) principle. Experiments on a dataset of over 5000 movement samples from 18 volunteers demonstrate that M-Fitness achieves 96.5% accuracy for isolated exercise recognition and 92.7% accuracy for compound exercise recognition, exhibiting strong adaptability to diverse environments.
Reliable, privacy-preserving emotion sensing is essential for next-generation IoT applications; however, vision or audio-only pipelines often break down when faces are masked, environments are noisy, or multiple speakers coexist. We present AuraVox, a contactless system that couples millimeter-wave lip micro-Doppler with speech acoustics and fuses them through AuraNet, a bespoke cross-modal transformer. The radar first localizes each talker via MUSIC-based direction-of-arrival (DOA) estimation and Bartlett beamforming, then captures high-resolution Doppler signatures of lip motion; in parallel, prosodic and spectral speech cues are extracted from a lapel microphone. AuraNet holistically attends to these heterogeneous streams and produces a unified representation for emotion classification. Evaluated on a 30-subject bilingual (English/Mandarin) corpus that includes mask-wearing, multispeaker overlap, and 30-90-cm ranges, AuraVox attains 96% macro-F1, outperforming radar-only and audio-only baselines by up to eight percentage points. End-to-end latency is 12.8 ms per frame on a 10-W Jetson Xavier NX, meeting real-time constraints for edge deployment. By unifying beamformed lip kinematics with speech cues through AuraNet, AuraVox delivers the first multispeaker, cross-language, mask-resilient emotion recognizer that runs on commodity hardware. Representative use cases include stress-aware in-cabin driver assistance, hospital check-in triage, and mood-adaptive smart-home interfaces.
As the size of large language models (LLMs) increases, the limitations of a single data center, such as constrained computational resources and storage capacity, have made distributed training across multiple data centers the preferred solution. However, a primary challenge in this context is reducing the impact of gradient synchronization on the training efficiency across multiple data centers. In this work, we propose a distributed training scheme for LLMs, named parallel gradient computation and synchronization (PGCS). Specifically, while one expert model is being trained to compute gradients, another expert model performs gradient synchronization in parallel. In addition, a gradient synchronization algorithm named BLP is developed to find the optimal gradient synchronization strategy under arbitrary network connectivity and limited bandwidth across multiple data centers. Ultimately, the effectiveness of PGCS and BLP in enhancing the efficiency of distributed training is demonstrated through comprehensive simulations and physical experiments.
Radar systems are increasingly being used for human activity recognition (HAR), healthcare monitoring, and smart environments due to their privacy-preserving nature, contactless operation, and robustness in varying lighting conditions. Traditional HAR approaches often rely on wearable sensors or vision-based methods, which may pose privacy concerns and have practical limitations in real-world settings. mmWave radar has emerged as a viable alternative, enabling detailed motion analysis without requiring intrusive sensors. In this study, we employ a low-cost mmWave radar to recognize student study-related activities in tabletop scenarios by generating micro-Doppler (mD) spectrograms and extracting Region of Interest (ROI) features. Our dataset, collected from 10 participants of varying age, height, and weight performing six different activities, captures fine-grained motion dynamics in a different meeting rooms setting at distances of 1 m and 2 m from the radar. This dataset captures fine-grained motion dynamics, particularly subtle hand movements, such as page-flipping in studying or picking up a cup while drinking. To enhance feature extraction, we implement a dynamic thresholding mask that emphasizes ROI regions in the spectrograms, ensuring that the most relevant motion signatures are captured. We evaluate classification performance using machine learning models. Initially, classifiers trained on extracted ROI features achieved up to 80% accuracy. To further improve recognition, we introduce a feature-level fusion strategy that integrates structured ROI-based features with spectrogram-based representations. Additionally, we employ hierarchical classification to first distinguish between major activity groups before refining classifications at a finer granularity. Furthermore, decision-level fusion is incorporated to combine classifier predictions, enhancing robustness across varying distances and environments. These optimizations resulted in a peak classification accuracy of up to 99% at 2 m and 98% at 1 m, demonstrating the effectiveness of our fusion-based approach. The generalization analysis across different participants and controlled indoor environments confirms the model’s adaptability within structured tabletop settings, highlighting its potential as a foundation for future privacy-sensitive indoor monitoring and smart environments.
Sitting posture is closely related to our health. Poor sitting posture can cause various diseases and jeopardize our health. Among the current methods for detecting sitting posture, computer vision solutions suffer from privacy leakage and wearable sensor solutions suffer from inconvenience and cost of wearing. In this study, we introduce 3D-Sitpose, which leverages millimeter-wave radar to detect human sitting posture. 3D-Sitpose utilizes wireless signal transmission for non-contact detection, ensuring privacy protection and cost reduction. Firstly, we analyze the impact of variations in human sitting posture on millimeter-wave radar signals, and design sophisticated signal processing methods to refine the collected radar data, yielding clearer point cloud information for volunteers in various sitting postures. Secondly, we develop a two-channel neural network to extract fine-grained features related to volunteers from the point cloud data. Finally, we obtain coordinates for 25 human skeletal points. 3D-Sitpose can instruct users to maintain correct sitting posture based on a set of six key angles. We recruit 20 volunteers from our institute to conduct comprehensive evaluations of 3D-Sitpose. Experiments are conducted in two indoor environments to estimate sitting posture. The results reveal the mean Euclidean distance error for all skeletal point locations is 6.65 cm. This demonstrates that our method is able to estimate various sitting changes in volunteers.
DriveEmo-FL presents a privacy-preserving, radar-based emotion-recognition framework tailored for autonomous-vehicle (AV) cabins. Leveraging a compact embedded mmWave radar, the system captures upper-body emotional gestures using a CFAR-enhanced preprocessing pipeline, extracting both micro-Doppler signatures and velocity-time profile (VTP) features. These features are processed via EmoNet, a lightweight dual-stream deep learning model that performs early fusion of spatial-temporal and motion statistics. EmoNet achieves a top classification accuracy of 94.5% (Precision: 0.945, Recall: 0.943, F1-Score: 0.944) with an average latency of 9.7 ms on edge hardware. DriveEmo-FL’s effectiveness is validated through extensive testing across four real-world vehicular scenarios: motion, low light, direct sunlight, and gesture overlap, demonstrating robust performance under diverse conditions. Additionally, we incorporate federated learning to preserve passenger privacy and enable model generalization across multiple AV fleets without sharing raw data. Comparative evaluation against six state-of-the-art models confirms EmoNet’s superiority in both accuracy and computational efficiency. By linking emotional state detection to adaptive AV behaviors, DriveEmo-FL offers a proactive, intelligent interface for future emotion-aware intelligent transportation systems.
This paper introduces a novel mm-wave radar-based system for detecting upper-body emotional gestures using micro-Doppler(m-D) signatures and motion profiling to enable emotion-driven IoT device control in smart homes. The system recognizes five key gestures,single thumbs up, double thumbs up, single thumbs down, double thumbs down, and the OK gesture, each mapped to emotional states such as approval, disapproval, and strong emotional responses. By leveraging m-D features and motion profiling, the system captures both gestures’ dynamic and temporal characteristics to infer the corresponding emotional context with high precision. Motion profiling is used to analyze gesture velocity, acceleration, and duration, providing insights into each motion’s intensity and emotional significance. Combined with m-D data, the system achieves robust emotion detection, even under challenging conditions such as low-light conditions or occlusions. The proposed framework is computationally efficient, has a complexity of O(N log N + M ) and delivers an average accuracy in emotion recognition. The system demonstrates low latency (ms) and seamless integration with IoT devices, enabling real-time smart home automation based on emotional cues.
Reliable, privacy-preserving emotion sensing is essential for next-generation IoT applications; however, vision or audio-only pipelines often break down when faces are masked, environments are noisy, or multiple speakers coexist. We present AuraVox, a contactless system that couples millimetre-wave lip micro-Doppler with speech acoustics and fuses them through AuraNet, a bespoke cross-modal transformer. The radar first localizes each talker via MUSIC-based direction-of-arrival estimation and Bartlett beamforming, then captures high-resolution Doppler signatures of lip motion; in parallel, prosodic and spectral speech cues are extracted from a lapel microphone. AuraNet holistically attends to these heterogeneous streams and produces a unified representation for emotion classification. Evaluated on a 30-subject bilingual (English/Mandarin) corpus that includes mask-wearing, multi-speaker overlap, and 30–90 cm ranges, AuraVox attains 96% macro-F1, outperforming radar-only and audio-only baselines by up to eight percentage points. End-to-end latency is 12.8 ms per frame on a 10 W Jetson Xavier NX, meeting real-time constraints for edge deployment. By unifying beamformed lip kinematics with speech cues through AuraNet, AuraVox delivers the first multi-speaker, cross-language, mask-resilient emotion recognizer that runs on commodity hardware. Representative use cases include stress-aware in-cabin driver assistance, hospital check-in triage, and mood-adaptive smart-home interfaces.
Traditional RF-based liquid identification methods generally rely on a single characteristic such as refractive index or permittivity and often assume prior container knowledge, limiting their versatility. These approaches also face challenges in scenarios involving gradual state changes in the liquid. We propose LiqState, a contactless framework for fine-grained liquid identification and continuous state monitoring, capable of operating without prior container information. To mitigate container effects, we developed a LiqState reflection model that analyzes frequency-dependent changes, leveraging the diverse permittivity profiles of liquids across the mmWave frequency range. Our approach introduces a novel feature extraction method, VRCP, which captures four distinct physical and chemical properties for robust identification and state monitoring. Using LiqNet, a service-oriented and customized deep learning model, LiqState achieves an average classification accuracy of 97.3% across diverse conditions, accurately distinguishing 12 liquid types. Additionally, case studies highlight LiqState’s capability to monitor complex processes, such as milk fermentation (RMSE: 0.251) and fruit juice ripening (RMSE: 0.162), and differentiate between similar liquids with minimal alcohol concentration variations.
Drunk-driving is an important factor causing road traffic accidents and deaths, which deserves a lot of research. However, most current methods for detecting drunk-driving depend on customized hardware or require users' active participation, making it impractical to monitor blood alcohol content (BAC) during driving. This article introduces BACFuse, a device-free, contactless, and noninvasive system utilizing smartphone in driving environments, which achieves relatively high accuracy in drunk-driving monitoring by integrating various voice sensing modalities. BACFuse first captures vocal cord vibration from ultrasonic signals, then records voice commands from audio signals. BACFuse combines the ultrasonic signals with audio signals and effectively detects drunk-driving and BAC. A key enabler lies in our modeling of latent interaction between acoustic and ultrasonic signals to mitigate ambient noise, realizing noise-resistant drunk-driving detection. Additionally, we propose an effective modules within the co-attention method to fuse the multimodal signals, further enhancing the accuracy of drunk-driving detection. We conduct extensive experiments to evaluate BACFuse's performance on 20 participants in safe laboratory experiments. The results demonstrate that our system achieves BAC measurement with an MAE of 2.13 mg/dl, showing promise for future in-car driving management paradigms.
This study introduces a novel dual-modality emotion recognition system that combines mm-wave radar with camera-based labeling to provide accurate and privacy-preserving emotion detection. The mm-wave radar captures subtle physiological signals through micro-Doppler and time-frequency characteristics, while the camera assists in labeling facial expressions. The radar data is transformed into spectrograms, which are then fused with camera datasets to train deep learning models, employing convolutional layers for feature extraction and recurrent layers for temporal pattern recognition. Performance evaluation, conducted across a wide range of real-world occlusion and interference scenarios, shows that the system achieves 98.5% accuracy, 0.98 F1-score, and 0.98 recall, significantly outperforming traditional systems. Other experiments, including those for multi-person interference, hand-held paper occlusion, and industrial goggles, achieved accuracy rates of 92%, 91%, and 92%, respectively. The system’s latency for real-time processing is 60.5 ms on edge devices like the NVIDIA Jetson, making it suitable for applications requiring low-latency emotion recognition. Additionally, radar parameter optimization, such as adjusting the ADC sample rate and chirp size, has been shown to improve classification accuracy. These findings highlight the system’s robustness and adaptability to varying environmental conditions and its potential use in privacy-sensitive applications, including healthcare, security, and interactive media. Future work will explore radar-only systems, further reducing dependence on visual data, and investigate more advanced deep learning techniques to improve performance, scalability, and real-time deployment.
mm-FERP (millimeter wave Facial Expression Recognition for Personality) explores the use of mm-Wave radar technology, specifically the TI IWR1443, to assess personality traits based on the OCEAN model through facial expression analysis. This research uniquely combines psychological profiling with state-of-the-art technology to predict the OCEAN (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) personality traits by carefully analyzing facial muscle movements collected through mm-wave radar alongside detailed questionnaire analysis. Our advanced mm-FERP system employs mm-wave radar technology for the detection and analysis of facial expressions in a manner that is both non-intrusive and privacy-centric, handling the ethical and privacy concerns associated with traditional camera- based methods. Using a convolutional neural network (CNN), mm-FERP effectively analyzes the complex patterns in mm-wave signals. This approach enables the smooth transfer of model knowledge from extensive image-based (Scalograms) datasets to the detailed understanding of mm-wave radar signals, significantly enhancing the model's predictive accuracy and efficiency in identifying personality traits via emotional behavior. Our in-depth evaluation reveals mmFERP's remarkable potential to predict personality traits through emotion recognition (Neutral, Smile, Angry, Sad, Amazed) with an impressive accuracy of 97% across distances up to 0.47 m. We experiment in a controlled environment with more than 50 participants from different age groups (18-35) including males and females of different continents to train our model on different facial symmetry. Each participant gives 50 samples 10 for each expression making a total of 2500 samples. We also collected a self-assessment report from the same participants of 64 questions related to psychological behavior to validate personality by correlating it with radar signal features on question value weight (0.5-1.5). mm-FERP achieve an average score of 97.8% in precision, 97.2% in Recall, and 97.2% of F1. These results show mm-FERP's ability as an innovative approach for psychological behavioral analysis through mm-wave emotion recognition, improving user experience design, and paving the path for interactive technologies that are both personalized and psychologically insightful.
The massive-antenna wideband millimeter wave (mmWave)/terahertz (THz) systems inevitably suffer from a severe beam split effect due to the non-negligible signal propagation delays, which dramatically reduces communication efficiency. Nevertheless, if the wideband split effect is properly utilized, it can also bring benefits via sensing split directions for channel training. Hence, this paper proposes a novel wideband beam alignment framework with true-time-delayer (TTD) modules, which can fully exploit the controllable split beams for efficient angle-of-arrivals (AoAs) estimation. Moreover, we develop a hierarchical posterior matching (PM) enabled wideband beam alignment approach, which proactively configures the split beams to accelerate the estimation of AoAs posterior probability distributions. To deal with the computational complexity of the predesigned codebook and the insensitivity of the Gaussian distribution assumption in PM, we further introduce a low-complex and high-flexible wideband beam alignment approach based on a deep unfolding mechanism. Numerical results verify that: 1) The proposed framework can significantly improve the AoAs estimation accuracy at the cost of the same pilot overheads. 2) The proposed low-complexity deep unfolding approach outperforms the conventional PM mechanism even in low signal-to-noise-ratio (SNR) scenarios.
Early detection of myocardial infarction (MI) is essential for alleviating symptoms and improving daily activity performance. Researchers typically employ continuous segments of heartbeat signals (20-30 seconds), such as ECG signals, for MI detection, as MI often induces changes in heartbeat patterns. Current MI detection methods, like wearable sensors, may induce discomfort from prolonged wear, and Radio Frequency (RF) based approaches might fail to extract fine-grained heartbeat signals during vigorous movement. This article presents a reliable and motion-robust MI detection method based on RF signals. By developing a series of advanced signal processing algorithms, MI-Ra can capture fine-grained heartbeat signals during various daily activities. Our design is inspired by the fact that RF reflections caused by heartbeat signals are mixed with other motion-induced reflections in a nonlinear manner. We utilize the Taylor series expansion method to extract the linear component of these mixed non-linear signals and propose a novel Generative Adversarial Networks (GAN) method, named IQ-TransGAN, to separate the heartbeat signal. To enhance MI detection reliability, MI-Ra employs a multi-periodicity modeling method to extract refined signal representations from recovered heartbeat signals. We have recruited 50 volunteers with MI from Zhongnan Hospital of Wuhan, China, and 50 volunteers without MI, for comprehensive evaluations. The results demonstrate that MI-Ra achieves an average MI detection accuracy of 95.2% when user is quasi-stationary. Even during user non-stationary conditions, MI-Ra maintains an average detection accuracy of 90.5%. MI-Ra shows promise in paving way for smart home healthcare.
Wireless sensing offers a promising approach for non-destructive and contactless identification of the moisture content in fruits. Traditional methods assess fruit quality based on external features such as color, shape, size, and texture. However, fruits often appear perfect externally while being rotten inside. Thus, accurately measuring internal conditions is crucial. This paper introduces mmFruit, a non-destructive and ubiquitous system that employs mmWave signals for precise and robust moisture level sensing in thin and thick pericarp fruits. We propose a novel dual incidence moisture estimation model for regular moisture monitoring to achieve high granularity and eliminate fruit type and size dependency. Additionally, we leverage unique reflection responses across different mmWave frequencies to provide discriminative information about fruit moisture levels. Our comprehensive theoretical model demonstrates how fruits' refractive index, attenuation factor, and elasticity can be estimated by eliminating fruit type dependency. We developed an electric field distribution model utilizing two receiving antennas to address the challenge of varying fruit sizes through a differential approach, aiming to improve overall robustness. mmFruit integrates a customized Spatial-invariant network (SpI-Net) to eliminate interference from different frequencies and locations, ensuring stable moisture monitoring regardless of target displacement. Extensive experiments were conducted over a month in varied environments on seven types of fruits with thin and thick pericarps (apple, pear, peach, mango, orange, dragon fruit, and watermelon). The results demonstrate that mmFruit achieves a commendable RMSE of 0.276 in moisture estimation. It accurately distinguishes fruits with minor moisture level differences (0% to 7%) with 93.6% accuracy and higher moisture differences (45% to 65%) with over 95.1% accuracy, even in scenarios involving diverse displacements and rotations.
Target material sensing in non-invasive and ubiquitous contexts plays an important role in various applications. Recently, a few wireless sensing systems have been proposed for material identification. In this article, we introduce mm-CUR, A Novel Ubiquitous, Contact-free, and Location-aware Counterfeit Currency Detection in Bundles using a Millimeter-Wave Sensor. This system eliminates the need for individual note inspection and pinpoints the location of counterfeit notes within the bundle. We use Frequency Modulated Continuous Wave (FMCW) radar sensors to classify different counterfeit currency bundles on a tabletop setup. To extract informative features for currency detection from FMCW signals, we construct a Radio Frequency Snapshot (RFS) and build signal scalogram representations that capture the distinct patterns of currency received from different currency bundles. We refine the RFS by eliminating multi-path interference, and noise cancellation and apply high pass filters for mitigating the smearing effect with the continuous wavelet transform (CWT). To broaden the usage of mm-CUR, we built a transferable learning model that yields robust detection results in different scenarios. The classification results demonstrated that the proposed counterfeit currency detection system can detect counterfeit notes in 100-note bundles with an accuracy greater than 93%. Compared to the standard CNN and DNN methods, the proposed mm-CUR model showed superior performance in distinguishing each bundle data, even for a limited-size dataset.