Robust modulation classification remains challenging under low-SNR and non-ideal channel conditions. Due to the diversity of modulation formats and signal characteristics, a single signal representation often fails to capture the full range of discriminative features needed for accurate classification. In addition, most existing AMC models are trained from scratch, which is often inefficient in terms of training time and data usage. Fine-tuning large-scale models leverages transferable features learned from diverse visual data, enabling more efficient and generalized learning. In this work, we propose CHEST, a dualstrategy framework that combines multi-representation feature fusion with fine-tuning of pre-trained models to improve classification robustness. Experimental results show that the proposed method outperforms existing baselines.
Electroencephalography (EEG) based pattern recognition systems faces challenges in cross-subject generalization due to signal non-stationarity and inter-subject variability. Unsupervised domain adaptation (UDA) leverages labeled data from multiple source subjects to address this, yet existing methods often use a single distance metric for alignment and treat all sources equally, risking suboptimal or negative transfer. This paper presents an Enhanced Multi-Source Domain Adaptation (EMSDA) approach for EEG classification problems, enhancing the Multi-Source Marginal Distribution Adaptation (MSMDA) framework with multiple distance metrics (e.g., Maximum Mean Discrepancy (MMD), Correlation Alignment (CORAL)) for robust distribution alignment, a top-n close source selection strategy and a domain distance based strategy to weight sources. Evaluated on the Leipzig Mind-Brain-Body (LEMON) dataset, our method significantly outperforms baseline techniques in cross-subject classification, demonstrating adaptability to heterogeneous populations.
Foundation models represented by ChatGPT, have initiated an outstanding revolution across various domains. With the pre-trained foundation model in specific fields, numerous downstream tasks exhibit state-of-the-art performances. This paper extends this paradigm to automatic modulation classification (AMC), employing a masked autoencoder vision transformer (ViT-MAE) as a foundation model to advance AMC. The experimental results show that our signal constellation diagram-based foundation model outperforms traditional deep learning methodologies, underscoring the vast potential of foundation models in AMC and wireless communication systems.
Reliably detecting and accurately counting neutrons in the presence of a strong gamma ray background is critical to nuclear security and safeguards applications. It is very challenging especially at high count rates where pileup events are more likely to happen. If not detected, pileup is likely to be misclassified. Traditional approaches to pulse shape discrimination (PSD) and pileup rejection heavily rely on the manually selected features in either time or frequency domains. Their performance is susceptible to noise and complex high count rate scenarios. In this paper, vision transformers (ViT) are trained for PSD and pileup detection based on the continuous wavelet transform of the pulses. The superiority of ViT over the traditional approaches and the other deep learning models in detecting the pileup pulses is demonstrated, even for the close pileup cases.
In many applications of wireless sensor networks (WSNs), sensor positions are often not known exactly. The existence of sensor position error (SPE) may significantly impact system performance if not appropriately modeled or considered. In this article, we address the important sensor selection problem in the presence of SPE for general nonlinear measurement models considering independent and correlated measurement noise cases. The sensor selection problem is formulated in the framework of sparse sensing, in which the number of activated sensors is minimized subject to certain predetermined performance constraints. To facilitate the challenging convex relaxation of nonconvex constraints for the case with independent measurement noise, we prove that the Fisher information matrix (FIM) remains additive even in the presence of SPE. For correlated measurement noise, quadratic inequality constraints are introduced and two suboptimal solvers are proposed. The first solver uses matrix decomposition to transform the quadratic constraint into linear matrix inequalities (LMIs), while the second solver iteratively performs a linearization procedure on the quadratic constraint to obtain a reduced dimension of LMI, thus decreasing the computational complexity significantly. The proposed algorithms are compared in terms of their computational complexity quantitatively and experimentally with suboptimal greedy approaches and existing algorithms ignoring SPE. The results show the importance and necessity of considering SPE when implementing sensor selection and demonstrate the effectiveness of the proposed three sensor selection solvers.
Prognoses of Traumatic Brain Injury (TBI) outcomes are neither easily nor accurately determined from clinical indicators. This is due in part to the heterogeneity of damage inflicted to the brain, ultimately resulting in diverse and complex outcomes. Using a data-driven approach on many distinct data elements may be necessary to describe this large set of outcomes and thereby robustly depict the nuanced differences among TBI patients' recovery. In this work, we develop a method for modeling large heterogeneous data types relevant to TBI. Our approach is geared toward the probabilistic representation of mixed continuous and discrete variables with missing values. The model is trained on a dataset encompassing a variety of data types, including demographics, blood-based biomarkers, and imaging findings. In addition, it includes a set of clinical outcome assessments at 3, 6, and 12 months post-injury. The model is used to stratify patients into distinct groups in an unsupervised learning setting. We use the model to infer outcomes using input data, and show that the collection of input data reduces uncertainty of outcomes over a baseline approach. In addition, we quantify the performance of a likelihood scoring technique that can be used to self-evaluate the extrapolation risk of prognosis on unseen patients.
Pulse shape discriminating scintillator materials in many cases allow the user to identify two basic kinds of pulses arising from two kinds of particles: neutrons and gammas, respectively. An uncomplicated solution for building a classifier consists of a two-component mixture model learned from mixtures of pulses from neutrons and gammas at a range of energies. Depending on the conditions of data gathered to be classified, multiple classes of events besides neutrons and gammas may occur, most notably pileup events. All these kinds of events that are neither neutron nor gamma are anomalous and, in cases where the class of the particle is in doubt, it is preferable to remove them from the analysis. This study compares the performance of two analytical methods for using the scores from the two-component model to identify anomalous events and in particular to remove pileup events. This study further benchmarks the analytical methods against supervised machine learning methods. This study also presents a means of assessing performance of pileup removal using ROC curves and precision–recall curves. A specific outcome of this study is to propose a novel anomaly score, denoted by G, from an unsupervised two-component model that is conveniently distributed on the interval [−1,1].
Real-time emotional state recognition from neural activities has great potentials, e.g., enabling closed-loop systems to treat neuropsychiatric disorders. Besides its utility in identifying epileptogenic brain regions, electrocorticography (ECoG) provides the opportunity to study brain activities during various emotional events. In this study, we aim to test if the second order statistics such as the electrode network connectivity are relevant in discriminating different emotions. Specifically, we adopt the statistical dependence as the connectivity measure and use sparse logistic regression for classification based on the connectivity features. The practical issues of limited samples and imbalanced data are addressed. We show that the data with the highest synchronization in the γ band provide the most competitive performance. The identified brain regions involved in the most discriminative connectivity patterns are consistent with the previous findings in the literature.
In this paper, we address the multiple sound source localization problem by associating and fusing the direction of arrival (DOA) estimates from multiple microphone arrays. For multi-source scenarios especially in indoor environments, a critical issue is to tell the correspondence among DOA estimates across different arrays, which is known as the data association problem. We propose a multi-dimensional assignment-based data association approach to find the optimal associations of DOA estimates from the same source. First, in the sense of maximum likelihood, the data association problem is formulated by finding the most likely partition of the measurement set into the source-originated and false alarm-originated subsets. Next, by defining the association costs appropriately, the problem of finding the most likely measurement partition is transformed into a generalized multi-dimensional assignment problem which can be solved efficiently by a Lagrangian relaxation algorithm. After the optimal associations of DOA estimates across different arrays are obtained, the locations of sources can be estimated by fusing the same source-originated DOA estimates. In the presence of missed detections, false alarms and the unknown number of sources, the proposed method achieves high accuracy in data association and localization, and outperforms the competing method in reverberant and noisy environments. In addition, since our method does not require additional features and uses DOA estimates only, it is more computationally efficient than the competing method.
Detailed recordings of brain activity acquired using Electrocorticography (ECoG) sensors offer an opportunity to explore the activity of the human brain. We study the information content of ECoG array data with respect to the emotional affect of the subject. Lognormal spectral models are estimated on data from patients undergoing monitoring for intractable epilepsy. A model based approach for mapping discriminative locations is developed for this problem. Differing spatial sensitivity in the ECoG array allows for retrospective analysis of localized brain activity. Our discriminability measure evaluates the information content at each sensor location relative to Positive and Negative displays of emotional affect. This measure can be approximated by a symmetrized Kullback-Leibler divergence between the estimated models.
Obstructive sleep apnea (OSA) is a serious sleep disorder affecting millions of people worldwide. There is a great need to develop an efficient, low-cost OSA detection method. Traditional OSA detection methods are purely data-driven and hence their detection performance greatly depends on the quality and quantity of the sensor data. Several mathematical models of the human cardiorespiratory system exist which can generate different physiological signals that are hard to measure using current sensor technology. In this paper, we propose a new framework for OSA detection in which we fuse the sensor data with the physiological signal data from the cardiorespiratory system models. Multivariate Gaussian processes (GPs) are used to capture and model the physiological signal variations among different individuals. We define the multivariate GP covariance function using the sum of separable kernel functions form and estimate the corresponding hyperparameters by maximizing the GP marginal likelihood function. We detect OSA using the heart rate signal on a window-by-window basis using a likelihood ratio test. We conduct several experiments on both simulated and real data to show the effectiveness of the proposed OSA detection framework. We also compare with other purely data-driven OSA detection methods to demonstrate the advantage of the proposed OSA detection fusion framework.
Bring the power system to its stable operation after the system being disturbed is one of the most important capabilities in power generation and control. A loss of synchronization is generally the result of instability. The angle of the rotor in a generator depends on the mechanical power and electrical power. The output power of a synchronous machine is linked to the mechanical torque, which is central to the angle stability of rotor. Therefore, maintaining synchronization in the presence of disturbance is the goal in power generation. In this paper, the transient stability of a power system is investigated by solving the swing equation using the particle swarm optimization (PSO). The investigation is carried out for the swing equation in terms of angle stability during the normal operation and after the fault occurs. Single machine and two-machine power systems are considered in this study. The obtained results show satisfactorily approximation solution for the swing equation. The proposed method introduces a similar and better solution compared to the other conventional numerical methods.
In this paper, we address the multiple sound source localization problem using time differences of arrival (TDOAs) of sound sources to a microphone array. Typically, TDOAs are estimated based on the peak extraction of the generalized crosscorrelation function. In multi-source cases, for any given microphone pair, it is hard to tell the correspondence between the sound sources and the extracted peaks. In this work, we develop a novel localization approach based on data association which combines multiple TDOAs from the same source across different microphone pairs. Firstly, the generalized cross correlation-phase transform (GCC-PHAT) function is evaluated and multiple peaks of the GCC function indicating candidate TDOAs are extracted for each pair of microphones. Next, we employ the multi-dimensional assignment algorithm to associate multiple TDOAs from the same source. Finally, multiple sound source localization is carried out based on the obtained TDOA associations across different microphone pairs. Experimental results show the proposed method achieves superior performance for multiple sound source localization compared to the competing algorithm, especially in noisy environments.
A good quality of electric power system is that both the frequency and voltage remain at the desired values during operation and transmission of power. If the load power changes, the frequency will oscillate and deviate from its rated value, leading to instability issues. Thus, a design of efficient load frequency control (LFC) is needed to maintain the frequency constant against continuous variation of loads, which is also referred as unknown external load disturbance. A proportional-integral-derivative (PID) controller has been used for decades as the load frequency controller to keep frequency approximately at the nominal value by tuning the proportional, integral and derivative gains of the PID controller. In this paper, we propose two methods to tune the PID controller. The first method is online tuning based on neural networks and the second method is offline tuning based on particle swarm optimization. The two tuning methods are applied on a single and two interconnected power areas. Both tuning methods are compared with each other and both show good performance in terms of the overshoot, undershoot and settling time, but online tuning method gives better results in damping out the frequency deviation compared to the offline tuning method.
In this paper, we consider the source localization problem in which several microphones collaborate to locate an active sound source in a reverberant environment. Sound source localization (SSL) based on the Generalized Cross Correlation (GCC) function is widely studied for the past few decades. However, in a reverberant environment, the maximal peak of the GCC function does not necessarily correspond to the true source location due to the multipath effect. In this case, the traditional GCC-based method performs poorly. In this paper, by combining the information from all the available microphone pairs, we aim to seek a set of source-originated peaks rather than the maximal peaks. To achieve this, for each pair of microphones, multiple peaks of the GCC function indicating candidate TDOAs are extracted firstly. A graphic model is then constructed based on the extracted TDOAs from multiple microphone pairs, and the optimal association of peaks corresponding to true time delays can be obtained by optimizing the association cost function for the given set of peaks. Finally, the source location is estimated in the least square sense. Simulation results show the superior performance of the proposed approach compared with the traditional GCC-based localization algorithm.
This paper deals with sound source localization and number estimation in indoor environments using a circular microphone array. Multiple sound source localization is achieved by performing single source localization at each selected time-frequency (TF) point of received signals after short-time Fourier transform. A TF point selection method is proposed to reduce the computational time, which depends on a trained SVM with power and power ratio of TF points as its features. Nonparametric Bayesian clustering is applied on the obtained DOA estimates to identify the number of active sources. The algorithm is shown to outperform others through simulations.
In the presence of strong noise or reverberation, the acoustic source localization and tracking problem is confronted with severe challenges. One of them is that the maximal peak of the localization function may not be generated by the real source. In this paper, the stochastic region contraction (SRC) is adopted to search for multiple peaks of the localization function rather than the maximal one only, so as to provide sufficient pseudo-measurement information. A mixed distribution model is used to describe the pseudo-measurement likelihood function within the particle filter framework, where the mixing weight is evaluated based on the idea of probability data association (PDA) by incorporating the dynamic information of the sound source. Simulation results demonstrate that the proposed approach can effectively improve the ability of anti-noise and anti-reverberation of the acoustic target tracking system.
We consider the problem of throughput maximization for MIMO cognitive radios, jointly with respect to the sensing duration and the secondary user transmit power, under imperfect null space sensing. The average interference to the primary receiver, due to imperfect null space sensing is derived. Next, the formulated non-convex throughput maximization problem is solved based on geometrical insights. Under the assumption of a large number of secondary transmit antennas and low SNR conditions, analytical closed-form approximations for the secondary user transmit power and the sensing duration are also obtained. Extensive simulation results are provided to demonstrate the efficacy of the proposed approach.
Quantitative risk assessment is a critical first step in risk management and assured design of networked computer systems. It is challenging to evaluate the marginal probabilities of target states/conditions when using a probabilistic attack graph to represent all possible attack paths and the probabilistic cause-consequence relations among nodes. The brute force approach has the exponential complexity and the belief propagation method gives approximation when the corresponding factor graph has cycles. To improve the approximation accuracy, a region-based method is adopted, which clusters some highly dependent nodes into regions and messages are passed among regions. Experiments are conducted to compare the performance of the different methods.
Traditional biometric recognition systems often utilize physiological traits such as fingerprint, face, iris, etc. Recent years have seen a growing interest in electrocardiogram (ECG)-based biometric recognition techniques, especially in the field of clinical medicine. In existing ECG-based biometric recognition methods, feature extraction and classifier design are usually performed separately. In this paper, a multitask learning approach is proposed, in which feature extraction and classifier design are carried out simultaneously. Weights are assigned to the features within the kernel of each task. We decompose the matrix consisting of all the feature weights into sparse and low-rank components. The sparse component determines the features that are relevant to identify each individual, and the low-rank component determines the common feature subspace that is relevant to identify all the subjects. A fast optimization algorithm is developed, which requires only the first-order information. The performance of the proposed approach is demonstrated through experiments using the MIT-BIH Normal Sinus Rhythm database.