
This research focuses on the development and performance evaluation of a recommendation system model that leverages time-based data and user preference scores. The model recommends movies to users by analyzing movie data and the preference scores previously assigned by users in the system, with these scores decreasing in significance over time. The research utilizes the MovieLens dataset, which is divided into 80% for model training and 20% for testing. The model employs a Collaborative Filtering technique to determine user similarity, utilizing Pearson Correlation and Cosine Similarity methods. It then selects the nearest neighbors to the user through a two-step process: in the first step, the top K most similar neighbors are identified, and in the second step, the top 10 “friends of friends” for each neighbor are selected. The study tests K values ranging from 10 to 40, incrementing by 5 at each step. The results indicate that as the K value increases, the model's performance also improves. At a K value of 10, the model using Pearson Correlation slightly outperforms the one using Cosine Similarity, with scores of 0.53 and 0.51, respectively. When the K value increases to 15, both models perform similarly, with scores of 0.77 and 0.75, respectively. At K values of 30, 35, and 40, both models achieve equivalent performance, with scores of 0.95, 0.97, and 0.98, respectively. This research demonstrates that the developed model can be effectively applied to recommend movies or other products with similar characteristics, providing a robust foundation for time-sensitive and user-preference-based recommendation systems.
Measuring remote photoplethysmography (rPPG), a contactless facial video-based PPG estimation, requires a large amount of labeled data via supervised methods, leading to significantly increased labor and costs. Existing rPPG studies based on self-supervised methods have utilized temporal and spatial similarities between rPPG data, which help reduce manual labeling costs. Here, we propose a novel self-supervised rPPG estimation method with contrastive learning. The proposed method not only leverages temporal and spatial similarities as representations but also maintains a well-contained representation. A well-contained representation reduces the computational cost associated with calculating negative loss. We use a 3D-CNN model to capture well-contained representation and multiple rPPG signals, which perform on different spaces but are temporally similar. We show that the proposed method outperforms current rPPG methods in terms of heart rate estimation accuracy on two public datasets, i.e., PURE and UBFC.
This study proposed a body movement-based system for efficiently controlling robots in large indoor environments. Traditional hand gesture recognition systems experience significant accuracy loss in large indoor spaces, particularly as the distance between the user and the camera grows. To address this issue, clear and large body movements were used as control signals instead of hand movements. For example, participants could give commands to the robot by raising both hands or lifting one hand in a specific direction. The system utilized MediaPipe Pose to track the user's posture in real time and communicated commands from the server to the robot via TCP/IP socket communication. This structure ensured stable data transmission, enabling smooth robot control even in expansive spaces. Consequently, this research demonstrated the feasibility of effective robot control in large indoor areas, such as screen golf courses, and suggested the potential for system expansion in various indoor environments.
This paper proposes a low-power single-stage single-balanced radio frequency (RF) front-end receiver in which a high performance Colpitts based self-oscillating mixer (SOM) is folded at the output of a cascode low-noise amplifier (LNA). Since the cascode LNA is used in the front-end, the proposed SOM has good reverse isolation from the oscillator output to the RF input. Using a 65 nm CMOS technology, the proposed SOM is optimized for the best phase noise performance. Targeted for WLAN frequency band application, the proposed SOM oscillates at around 2.4 GHz and achieves the phase noise of -77 dBc/Hz, -100 dBc/Hz, and -120 dBc/Hz at 10 kHz, 100 kHz, and 1 MHz offset frequency, respectively. The simulated voltage conversion gain is about 44 dB. The double side-band (DSB) noise figure (NF) is about 6.8 dB at 1 MHz intermediate frequency (IF). The SOM cell consumes 755 μW de power from a 1-V supply.
In recent years, there has been a spike in demand for wearable devices. Low dropout regulators are an integral part of these devices, owing to their high Power Supply Rejection (PSR) and stable output. This paper proposes a high gain, low quiescent current $(10.298\mu \mathrm{A})$, Output Capacitorless Low Dropout Regulator (OCL-LDO) with a fast local loop and unity feedback. Nested Miller frequency Compensation (NMC) technique is used to stabilize the system. The proposed circuit can operate with a load current ranging from $20\text{pA}$ to 20mA. The design encompasses three gain stages with a total open loop gain of 132 dB resulting in enhanced load regulation (0.0831 µV/mA) and line regulation (0.6917 mV/V). The proposed LDO has been simulated in gpdk 90nm CMOS technology using Cadence Virtuoso. The simulation results demonstrate a PSR of −61.03 dB and −41.17 dB at 1kHz and 10kHz respectively.
Shadows appear in image areas where an object obstructs the light path. These areas, having lower values than non-shadow areas, degrade image quality and lead to issues in object recognition and segmentation. Supervised shadow removal depends on datasets with shadow and shadow-free image pairs, sparking increased interest in unpaired techniques. However, unpaired shadow removal faces challenges in accurately focusing on shadow areas without direct supervisory signals from ground truth. Consequently, shadows are not effectively removed, and non-shadow areas may suffer from unintended distortions. In this paper, we introduce a transformer-based network designed to identify shadow areas accurately, leveraging global context through spatial and channel attention mechanisms. Additionally, in the training phase, the transformer network is trained to precisely concentrate on shadow areas using a domain classifier and a dropkey mechanism, which randomly drops features of the keys to enhance focus. Our method is tested across several benchmark datasets for shadow removal, such as ISTD, ISTD+, SRD and WSRD, demonstrating better performance compared to current unpaired approaches.
Inrush current is commonly generated when transitioning from an idle to active state, thus causing power noise and voltage drops. This disrupts core stability, leading to potential errors, increased power delivery challenges, and reduced reliability in high-performance processors. This article presents a novel approach to Power Gating (PG) through a soft-start header that enables a gradual activation, allowing a controlled, stepwise increase in current and thereby alleviating the inrush current effect. In this approach, the width (W) of the header sleep transistor is strategically divided into smaller segments, such as W/2, W/5, W/10, or W/20, where each segment of the sleep transistor is activated sequentially, with fixed delay cells introduced between these segments to introduce a deliberate delay in activation. This staggered approach allows only a fraction of the total transistor width to turn on at any given moment, spreading out the current demand over time. The simulation results for UMC65nm CMOS on Cadence Virtuoso with 20% sleep signal demonstrate a notable reduction (54.55%, 82.01 %, 89.94%, and 93.65% in Plan 1 to Plan 4 respectively) in peak inrush current compared to conventional approaches, underscoring the efficacy of soft-start header PG in power-sensitive applications.
In general, determining the authenticity of damaged banknotes can often be challenging. To address this, the Bank of Korea exchanges banknotes that are not suitable for circulation due to damage or wear. In this study, we propose a banknote authenticity judgment system using MemSeg based on K-means memory update. This system identifies damage to banknotes more objectively and quickly. When memory update is performed using K-means, normal images with various feature patterns are updated uniformly in the memory. The experimental results show that the image-AUROC is improved by an average of 4.27% compared to the existing method. This proves that the proposed method is more objective.
Analyzing respiratory patterns is essential for diagnosing and monitoring various health conditions, particularly during sleep when irregularities such as apneas are prevalent. This study presents RespireSegNet, a deep audio segmentation method tailored for sleep breathing analysis, which addresses limitations of traditional signal processing techniques. Utilizing PSG-Audio dataset with tracheal sound recordings and respiratory belt data, RespireSegNet applies WhisperSeg, a pretrained Transformer-based model, to segment and analyze breathing cycles. The model captures subtle respiratory sounds amidst noise, demonstrating high precision in detecting respiratory rates and cycle durations across sleep stages. Compared with FFT and PeakFinding methods, RespireSegNet achieved superior accuracy in both breathing rate detection and cycle length estimation. These results highlight RespireSegNet's potential as a robust tool for non-invasive sleep disorder diagnostics, paving the way for improved respiratory sound analysis in healthcare applications.
This paper proposes a method that combines imitation learning and reinforcement learning to improve the learning performance of autonomous vehicles. The actor model was pre-trained using expert demonstration data from the CARLA simulator, and then integrated into PPO-based reinforcement learning. The experimental results showed that the model combining imitation learning outperformed the model using only reinforcement learning, particularly in the early stages of training. However, this study is limited by using only image data, and future research will incorporate vehicle position and speed information for enhanced results.
In this study, we propose a deep learning-based noise reduction and signal classification model designed to detect and classify victim signals in noisy disaster environments. The proposed model combines a CNN-based noise reduction model and an STFT-based signal classification model to enable real-time detection and classification of victim signals. This approach allows effective extraction of victim signals even in various noisy conditions, supporting rapid rescue operations in disaster situations. The experimental results show that the proposed model achieved 87.87% accuracy and an Fl-score of 0.8573, demonstrating high performance across a range of noisy con-ditions. Furthermore, the model was trained and evaluated on a dataset that reflects real disaster environments, proving its excellent performance in detecting victim signals. This study is expected to make a significant technical contribution to the prompt and accurate detection of victims in disaster situations.
In this paper, a non-return-to-zero (NRZ) transmitter (TX) with a 2-tap de-emphasis feed-forward equalization (FFE) is proposed for low-power double data rate memory. The TX adopts a quarter-rate clocking structure, which reduces the on-chip bandwidth and relieves timing margin. The low-voltage swing terminated logic voltage-mode driver performs the 2-tap FFE using 1 unit interval delayed data, and ZQ calibration is implemented for impedance matching with the channel. The driver is operated at 0.5-V VDDQ, while the remaining blocks are supplied with VDD of 1.05 V. The TX is designed using a 28-nm CMOS process and occupies an area of 0.0145 mm2. It achieves a data rate and energy efficiency of 15.6 Gb/s and 0.76 pJ/bit, respectively.
This paper presents a capacitive-to-digital converter (CDC) based on a time-to-digital converter (TDC). The proposed three-stage Vernier TDC achieves both high time resolution and wide time conversion range. In comparison to alternative CDC implementations, the proposed TDC is wholly digital, which results in reduced power consumption and more straightforward integration with other digital circuits within the system. The proposed CDC was implemented in a standard 0.18 μm CMOS process and demonstrated a capacitance conversion range of 150 pF to 700 pF and TDC with a time resolution of 14 ps. These results indicate that this capacitive digital converter is suitable for sensor applications and other systems requiring high conversion resolution.
In this paper, we examine the cybersecurity vulnerability assessment method of medical software. Medical software processes patient sensitive data and is linked to various medical devices and systems in real time. Due to these characteristics, medical software is highly likely to be exposed to various cybersecurity threats such as ransomware, data leakage, and medical device hacking. Based on the international standard IEC TS 60601-4-5, we propose threat modeling, vulnerability scanning, and penetration testing as a methodology for assessing the security vulnerabilities of medical software. Through this, we can identify security vulnerabilities in advance and prepare measures to respond quickly. We can prevent security threats and improve the safety of medical software through response strategies such as security patches and updates, network separation, data encryption, and security education. In conclusion, strengthening the security of medical software is essential to maintain patient safety and the reliability of the medical system, and systematic security assessment and continuous response are required.
In this paper, we propose a method to improve the accuracy of speech emotion recognition in extreme situations where noise is injected into the speech signal, which is equivalent to the signal magnitude of the speech signal. The proposed method is carried out in the following steps. i: Random Gaussian noise is injected into the speech signal and converted into a log-Mel spectrogram image by performing short-term Fourier transform (STFT), etc. ii: ResNet50 and a deconvolution layer are used to generate an implicitly filtered image with the same size as the input log-Mel spectrogram image. iii: Perform speech emotion recognition from the implicitly filtered image using the vertically long patch vision transformer (ViT), which is excellent for spectrogram image analysis. We experimentally demonstrate that by generating implicitly filtered images of the same size as the input image instead of conventional filtering using a CNN, we minimise the intervention of noise and facilitate the detection of high importance features. The proposed method is objectively evaluated using the noise-injected CREMA-D dataset. As a result, the proposed method achieves an accuracy of 39.02%, which is a significant improvement over the 11.95% and 17.66% accuracy of speech emotion recognition using existing methods.
Video-based person re-identification (ReID) has emerged as a pivotal task in multi-camera surveillance and security systems, enabling the accurate identification of individuals across diverse viewpoints. While traditional ReID approaches predominantly rely on image-based methodologies, video-based ReID introduces distinct challenges, including spatial distractions such as background clutter and temporal variations across consecutive frames, which frequently impede robust identity recognition. Building upon the Spatial and Temporal Memory Networks (STMN) [2] architecture, this study proposes an advanced framework for video-based ReID by integrating Orthogonal Projection [1], aiming to enhance model robustness in highly cluttered and dynamic environments with numerous distractors. The proposed method leverages spatial memory modules to identify and suppress distracting artifacts, thereby mitigating the influence of background noise on the learned person representations. Simultaneously, temporal memory modules are employed to model repetitive patterns in background dynamics, enabling the model to focus on temporally consistent identity-related features across video frames. To further enhance discriminative capabilities, Orthogonal Projection is introduced, which enforces orthogonal constraints within the embedding space. This mechanism ensures better separation between identity clusters by creating well-defined, non-overlapping decision boundaries. This integration of orthogonal regularization not only improves the discriminative power of the learned feature space but also establishes clear classification criteria, effectively reducing feature overlap and enhancing identity representation fidelity. Extensive experiments conducted on the MARS [3] dataset demonstrate that the proposed STMN with Orthogonal Projection significantly outperforms existing state-of-the-art methods, particularly under challenging scenarios involving partial oc-clusions, dynamic lighting conditions, and complex background noise. The proposed framework offers a robust and scalable solution for video-based ReID, extending its applicability to real-world surveillance settings with diverse operational constraints. These findings underscore the potential of the proposed method to advance the state of video-based ReID, offering a pathway toward more reliable and efficient multi-camera identification systems in complex and unconstrained environments.
This paper presents a novel deep learning-based approach to address time-delay challenges arising from nonlinear transformations in Gait Phase Estimation, a critical factor in analyzing human gait for lower limb exoskeleton robots. The proposed model leverages hip joint angles and angular velocities as feature data, organized using a sliding window technique, to predict the Linearized Gait Phase (LGP)‖a continuous repre-sentation of gait progression between Heel Strikes. Motivated by the need for real-time and precise gait phase estimation to enhance exoskeleton control, a Bidirectional LSTM network was employed to reduce delays and improve accuracy. The model's effectiveness was validated through comprehensive performance evaluations using metrics such as Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Mean Square Error (MSE), and R2 score, alongside a temporal analysis of delay reduction. The results demonstrated superior performance, with an MAE of 2.8±0.8%, RMSE of 3.9±1.2%, MSE of 0.1±0.1%, and an R2 score of 98.1±1.8%. Particularly, temporal analysis revealed a marked improvement, achieving an MAE of 0%, compared to 2.74±0.92% reported in prior studies. These findings underscore the proposed model's potential for real-time gait phase estimation, offering significant implications for enhancing the responsiveness and adaptability of lower limb exoskeleton robots.
Detecting pedestrians in blind spots is a critical component of ensuring vehicle safety, especially in urban environments. This paper presents a pedestrian localization system utilizing a single ultra-wideband (UWB) anchor installed in a moving vehicle. The proposed system leverages UWB's robust ranging capabilities to estimate the position of a pedestrian carrying a UWB-enabled smartphone, even in Non-Line-of-Sight conditions. The focus is on analyzing the UWB ranging protocol performance when the anchor is in motion and how the position of pedestrians can be estimated in both two-ranging and three-ranging point scenarios. Experiments are conducted at various vehicle speeds and demonstrated the effectiveness of the approach.
This article is studies of the noninvasive blood glucose measurement based on light diffuse reflection ratio of two PPG signal of red and infrared is proposed. Which the red light led and infrared light led is driven by 80 Hz and 190 Hz cosine wave frequency, respectively to generating two amplitude modulation that have photoplethysmography (PPG) as an information. It causes the two PPGs signal have immunity from ambient light and moving artifact noise and after that the amplitude modulation of two PPGs signal is taken to laptop computer by using soundcard interface and they are processed by two band pass filter center frequency 80 Hz and 190 Hz, respectively and next they are demodulated by squaring and passing through the low pass filter and take the output of LPF to be square rooting. the PPGs signal is appeared. Next, the two PPGs signal is to be calculated the reflection ratio by minimum and maximum of them, which the reflection ratio is to be evaluation Blood Glucose, in finally. The results from 10 volunteers show that this technique is possibility to measure BGP with noninvasive fashion. In the future work, the circuit should be improved, and more subject is experiment in order to more reliable.
In this paper, we present energy-efficient single-ended non-return-to-zero (NRZ) receiver with a 1-tap decision feedback equalizer (DFE) for low-power memory interfaces. While multi-level pulse amplitude modulation (PAM) signaling methods, such as P AM-3 and P AM-4, have been recently proposed to increase data rate, these techniques face challenges in maintaining signal integrity due to reduced voltage sensing margins and increased susceptibility to noise. To address these issues, our proposed receiver adopts NRZ signaling, which is more robust in low supply voltage conditions. The receiver employs a quarter-rate clocking scheme that reduces the operating clock frequency, thereby increasing timing margins and reducing clocking power. The receiver's three-stage sampler design minimizes the number of MOSFET stacks, improving performance at a low supply voltage. Additionally, the implemented 1-tap direct DFE effectively compensates for inter-symbol interference. Designed using a 28-nm CMOS process, the proposed receiver achieves a data rate of 15.6 Gb/s with a best energy efficiency of 0.11 pJ/bit.