This study represents the first integration of large language models (LLMs) with non-negative matrix factorization (NMF), marking a novel advancement in the source separation field. The LLM is employed in two unique ways: enhancing the separation results by providing detailed insights for disease prediction and operating in a feedback loop to optimize a fundamental frequency penalty added to the NMF cost function. We tested the algorithm on two datasets: 100 synthesized mixtures of real measurements, and 210 recordings of heart and lung sounds from a clinical manikin including both individual and mixed sounds, captured using a digital stethoscope. The approach consistently outperformed existing methods, demonstrating its potential to significantly enhance medical sound analysis for disease diagnostics.
The effectiveness of state-of-the-art cross-linking strategies and mass spectrometry (MS) detection was explored in an important biological context, namely, the ubiquitin-proteasome system, which is responsible for most of the regulated protein degradation in eukaryotic cells. The locations of possible binding sites on the S. cerevisiae 19S proteasome regulatory particle for Lys48 linked polyubiquitin chains were examined using cross-linking strategies and MS based detection by comparing two types of cross-linkers: a (bis)-sulfosuccinimidyl suberate (BS3) and diethyl suberothioimidate (DEST). The well-established BS3-based strategy produced 328 cross-linked peptides; however, no ubiquitin-19S cross-links were observed. The recently developed DEST-based approach produced fewer (146) linkages overall, but these included six ubiquitin-19S cross-links. Some of these cross-links are predicted by the canonical view of ubiquitin recognition, but others suggest novel insights into how the proteasome recognizes its substrates. A discussion of these strategies and structural implications for polyubiquitin-proteasome binding is provided.
Non-Negative Matrix Factorization (NMF) is an unsupervised learning method offering low-rank representations across various domains such as audio processing, biomedical signal analysis, and image recognition. The incorporation of α-divergence in NMF formulations enhances flexibility in optimization, yet extending these methods to multi-layer architectures presents challenges in ensuring convergence. To address this, we introduce a novel approach inspired by the Boltzmann probability of the energy barriers in chemical reactions to theoretically perform convergence analysis. We introduce a novel method, called Chem-NMF, with a bounding factor which stabilizes convergence. To our knowledge, this is the first study to apply a physical chemistry perspective to rigorously analyze the convergence behaviour of the NMF algorithm. We start from mathematically proven asymptotic convergence results and then show how they apply to real data. Experimental results demonstrate that the proposed algorithm improves clustering accuracy by 5.6
Large language models have shown a remarkable ability to extract meaning from unstructured data, offering new ways to interpret biomedical signals beyond traditional numerical methods. In this study, we present a matrix factorization framework for bioacoustic signal analysis which is enhanced by large language models. The focus is on separating bioacoustic signals that commonly overlap in clinical recordings, using matrix factorization to decompose the mixture into interpretable components. A large language model is then applied to the separated signals to associate distinct acoustic patterns with potential medical conditions such as cardiac rhythm disturbances or respiratory abnormalities. Recordings were obtained from a digital stethoscope applied to a clinical manikin to ensure a controlled and high-fidelity acquisition environment. This hybrid approach does not require labeled data or prior knowledge of source types, and it provides a more interpretable and accessible framework for clinical decision support. The method demonstrates promise for integration into future intelligent diagnostic tools.
This study introduces a novel unsupervised approach for separating overlapping heart and lung sounds using variational autoencoders (VAEs). In clinical settings, these sounds often interfere with each other, making manual separation difficult and error-prone. The proposed model learns to encode mixed signals into a structured latent space and reconstructs the individual components using a probabilistic decoder, all without requiring labeled data or prior knowledge of source characteristics. We apply this method to real recordings obtained from a clinical manikin using a digital stethoscope. Results demonstrate distinct latent clusters corresponding to heart and lung sources, as well as accurate reconstructions that preserve key spectral features of the original signals. The approach offers a robust and interpretable solution for blind source separation and has potential applications in portable diagnostic tools and intelligent stethoscope systems.
Staphylococcus aureus is one of the major community-acquired human pathogens, with growing multidrug-resistance, leading to a major threat of more prevalent infections to humans. A variety of virulence factors and toxic proteins are secreted during infection via the general secretory (Sec) pathway, which requires an N-terminal signal peptide to be cleaved from the N-terminus of the protein. This N-terminal signal peptide is recognized and processed by a type I signal peptidase (SPase). SPase-mediated signal peptide processing is the crucial step in the pathogenicity of S. aureus. In the present study, the SPase-mediated N-terminal protein processing and their cleavage specificity were evaluated using a combination of N-terminal amidination bottom-up and top-down proteomics-based mass spectrometry approaches. Secretory proteins were found to be cleaved by SPase, specifically and non-specifically, on both sides of the normal SPase cleavage site. The non-specific cleavages occur at the relatively smaller residues that are present next to the -1, +1, and +2 locations from the original SPase cleavage site to a lesser extent. Additional random cleavages at the middle and near the C-terminus of some protein sequences were also observed. This additional processing could be a part of some stress conditions and unknown signal peptidase mechanisms.
Auscultation provides a rich diversity of information to diagnose cardiovascular and respiratory diseases. However, sound auscultation is challenging due to noise. In this study, a modified version of the affine non-negative matrix factorization (NMF) approach is proposed to blindly separate lung and heart sounds recorded by a digital stethoscope. This method applies a novel NMF algorithm, which embodies a parallel structure of multilayer units on the input signal, to find a proper estimation of source signals. Another key innovation is the use of the periodic property of the signals which improves accuracy compared to previous works. The method is tested on 100 cases. Each case consists of two synthesized mixtures of real measurements. The effect of different parameters is discussed, and the results are compared to other current methods. Results demonstrate improvements in the source-to-distortion ratio (SDR), source-to-interference ratio (SIR), and source-to-artifacts ratio (SAR) of heart and lung sounds, respectively.
The human auditory system excels at detecting patterns needed for processing speech and music. According to predictive coding, the brain predicts incoming sounds, compares predictions to sensory input and generates a prediction error whenever a mismatch between the prediction and sensory input occurs. Predictive coding can be indexed in electroencephalography (EEG) with the mismatch negativity (MMN) and P3a, two components of event-related potentials (ERP) that are elicited by infrequent deviant sounds (e.g., differing in pitch, duration and loudness) in a stream of frequent sounds. If these components reflect prediction error, they should also be elicited by omitting an expected sound, but few studies have examined this. We compared ERPs elicited by infrequent randomly occurring omissions (unexpected silences) in tone sequences presented at two tones per second to ERPs elicited by frequent, regularly occurring omissions (expected silences) within a sequence of tones presented at one tone per second. We found that unexpected silences elicited significant MMN and P3a, although the magnitude of these components was quite small and variable. These results provide evidence for hierarchical predictive coding, indicating that the brain predicts silences and sounds.
LAY SUMMARY Combat Veterans are vulnerable to suicidal thoughts and behaviour. Many who die by suicide deny having suicidal ideation (SI). Typically, researchers try to find variables indicating the presence of SI using traditional statistical approaches. These approaches do not possess the capacity to detect highly complex multivariable interactions. In contrast, machine learning (ML) is designed to detect such patterns and can consequently yield much higher predictive accuracy. In this study, the authors trained ML algorithms using 192 variables extracted from questionnaires administered to 738 Veterans and serving personnel to detect the presence of self-harm and SI (SHSI). Using the 10 most predictive non-suicide-related items, the ML algorithms could detect SHSI with 75.3% accuracy. Most of these items reflect psychological phenomena that can change quickly over time, allowing repeated risk reassessment from day to day. The study’s findings suggest that ML methods may play an important role in the discovery, within a large data set, of predictive patterns that might be useful in suicide risk assessment.
Objective: In one of the largest and most comprehensive studies investigating the link between objective parameters of sleep and biological rhythms with peripartum mood and anxiety to date, we prospectively investigated the trajectory of subjective and objective sleep and biological rhythms, levels of melatonin, and light exposure from late pregnancy to postpartum and their relationship with depressive and anxiety symptoms across the peripartum period. Methods: One hundred women were assessed during the third trimester of pregnancy, of whom 73 returned for follow-ups at 1-3 weeks and 6-12 weeks postpartum. Participants were recruited from an outpatient clinic and from the community from November 2015 to May 2018. Subjective and objective measures of sleep and biological rhythms were obtained, including 2 weeks of actigraphy at each visit. Questionnaires validated in the peripartum period were used to assess mood and anxiety. Results: Discrete patterns of longitudinal changes in sleep and biological rhythm variables were observed, such as fewer awakenings (F = 23.46, P <.001) and increased mean nighttime activity (F = 55.41, P <.001) during postpartum compared to late pregnancy. Specific longitudinal changes in biological rhythm parameters, most notably circadian quotient, activity during rest at night, and probability of transitioning from rest to activity at night, were most strongly linked to higher depressive and anxiety symptoms across the peripartum period. Conclusions: Biological rhythm variables beyond sleep were most closely associated with severity of depressive and anxiety symptoms across the peripartum period. Findings from this study emphasize the importance of biological rhythms and activity beyond sleep to peripartum mood and anxiety.
Crosslinking mass spectrometry (XL-MS) of bacterial ribosomes revealed the dynamic intra- and intermolecular interactions within the ribosome structure. It has been also extended to capture the interactions of ribosome binding proteins during translation. Generally, XL-MS often identified the crosslinks within a cross-linkable distance (<40 Å) using amine-reactive crosslinkers. The crosslinks larger than cross-linkable distance (>40 Å) are always difficult to interpret and remain unnoticed. Here, we focused on stationary phase bacterial ribosome crosslinking that yields ultra-long crosslinks in an E. coli cell lysate. We explain these ultra-long crosslinks with the combination of sucrose density gradient centrifugation, chemical crosslinking, high-resolution mass spectrometry, and electron microscopy analysis. Multiple ultra-long crosslinks were observed in E. coli ribosomes for example ribosomal protein L19 (K63, K94) crosslinks with L21 (K71, K81) at two locations that are about 100 Å apart. Structural mapping of such ultra-long crosslinks in 70S ribosomes suggested that these crosslinks are not potentially formed within one 70S particle and could be a result of dimer and trimer formation as evidenced by negative staining electron microscopy. Ribosome dimerization captured by chemical crosslinking reaction could be an indication of ribosome-ribosome interactions in the stationary phase.
Background The purpose of this study was to investigate the utility of BNP, hsTroponin-I, interleukin-6, sST2, and galectin-3 in predicting the future development of new onset heart failure with preserved ejection fraction (HFpEF) in asymptomatic patients at-risk for HF. Methods This is a retrospective analysis of the longitudinal STOP-HF study of thirty patients who developed HFpEF matched to a cohort that did not develop HFpEF (n = 60) over a similar time period. Biomarker candidates were quantified at two time points prior to initial HFpEF diagnosis. Results HsTroponin-I and BNP at baseline and follow-up were statistically significant predictors of future new onset HFpEF, as was galectin-3 at follow-up and concentration change over time. Interleukin-6 and sST2 were not predictive of future development of new onset HFpEF in this study. Unadjusted biomarker combinations of hsTroponin-I, BNP, and galectin-3 could significantly predict future HFpEF using both baseline (AUC 0.82 [0.73,0.92]) and follow-up data (AUC 0.86 [0.79,0.94]). A relative-risk matrix was developed to categorize the relative-risk of new onset of HFpEF based on biomarker threshold levels. Conclusion We provided evidence for the utility of BNP, hsTroponin-I, and Galectin-3 in the prediction of future HFpEF in asymptomatic event-free populations with cardiovascular disease risk factors.
In certain conditions, dye-conjugated icosahedral virus shells exhibit suppression of concentration quenching. The recently observed radiation brightening at high fluorophore densities has been attributed to coherent emission, i.e., to a cooperative process occurring within a subset of the virus-supported fluorophores. Until now, the distribution of fluorophores among potential conjugation sites and the nature of the active subset remained unknown. With the help of mass spectrometry and molecular dynamics simulations, we found which conjugation sites in the brome mosaic virus capsid are accessible to fluorophores. Reactive external surface lysines but also those at the lumenal interface where the coat protein N-termini are located showed virtually unrestricted access to dyes. The third type of labeled lysines was situated at the intercapsomeric interfaces. Through limited proteolysis of flexible N-termini, it was determined that dyes bound to them are unlikely to be involved in the radiation brightening effect. At the same time, specific labeling of genetically inserted cysteines on the exterior capsid surface alone did not lead to radiation brightening. The results suggest that lysines situated within the more rigid structural part of the coat protein provide the chemical environments conducive to radiation brightening, and we discuss some of the characteristics of these environments.
Objective: Schizophrenia is a severe mental disorder associated with nerobiological deficits. Auditory oddball P300 have been found to be one of the most consistent markers of schizophrenia. The goal of this study is to find quantitative features that can objectively distinguish patients with schizophrenia (SCZs) from healthy controls (HCs) based on their recorded auditory odd-ball P300 electroencephalogram (EEG) data. Methods: Using EEG dataset, we develop a machine learning (ML) algorithm to distinguish 57 SCZs from 66 HCs. The proposed ML algorithm has three steps. In the first step, a brain source localization (BSL) procedure using the linearly constrained minimum variance (LCMV) beamforming approach is employed on EEG signals to extract source waveforms from 30 specified brain regions. In the second step, a method for estimating effective connectivity, referred to as symbolic transfer entropy (STE), is applied to the source waveforms. In the third step the ML algorithm is applied to the STE connectivity matrix to determine whether a set of features can be found that successfully discriminate SCZ from HC. Results: The findings revealed that the SCZs have significantly higher effective connectivity compared to HCs and the selected STE features could achieve an accuracy of 92.68%, with a sensitivity of 92.98% and specificity of 92.42%. Conclusion: The findings imply that the extracted features are from the regions that are mainly affected by SCZ and can be used to distinguish SCZs from HCs. Significance : The proposed ML algorithm may prove to be a promising tool for the clinical diagnosis of schizophrenia.
When trialkylamines are added to buffered solutions of peptides, unexpected adducts can be formed. These adducts correspond to Schiff base products. The source of the reaction is the unexpected presence of aldehydes in amines. The aldehydes can be detected in a few ways. Most importantly, they can lead to unanticipated results in proteomics experiments. Their undesirable effects can be minimized through the addition of other amines.
The current literature presents a discordant view of mild traumatic brain injury and its effects on the human brain. This dissonance has often been attributed to heterogeneities in study populations, aetiology, acuteness, experimental paradigms and/or testing modalities. To investigate the progression of mild traumatic brain injury in the human brain, the present study employed data from 93 subjects (48 healthy controls) representing both acute and chronic stages of mild traumatic brain injury. The effects of concussion across different stages of injury were measured using two metrics of functional connectivity in segments of electroencephalography time-locked to an active oddball task. Coherence and weighted phase-lag index were calculated separately for individual frequency bands (delta, theta, alpha and beta) to measure the functional connectivity between six electrode clusters distributed from frontal to parietal regions across both hemispheres. Results show an increase in functional connectivity in the acute stage after mild traumatic brain injury, contrasted with significantly reduced functional connectivity in chronic stages of injury. This finding indicates a non-linear time-dependent effect of injury. To understand this pattern of changing functional connectivity in relation to prior evidence, we propose a new model of the time-course of the effects of mild traumatic brain injury on the brain that brings together research from multiple neuroimaging modalities and unifies the various lines of evidence that at first appear to be in conflict.
Shahram Shirani合作论文数Professional Engineers Ontario;The Institute of Electrical and Electronics Engineers (IEEE);UBC Alumni Association8