Deploying Vision-Language Models (VLMs) on edge devices is challenged by resource constraints and performance degradation under distribution shifts. While test-time adaptation (TTA) can counteract such shifts, existing methods are too resource-intensive for on-device deployment. To address this challenge, we propose LQA, a lightweight, quantized-adaptive framework for VLMs that combines a modality-aware quantization strategy with gradient-free test-time adaptation. We introduce Selective Hybrid Quantization (SHQ) and a quantized, gradient-free adaptation mechanism to enable robust and efficient VLM deployment on resource-constrained hardware. Experiments across both synthetic and real-world distribution shifts show that LQA improves overall adaptation performance by 4.5%, uses less memory than full-precision models, and significantly outperforms gradient-based TTA methods, achieving up to 19.9× lower memory usage across seven open-source datasets. These results demonstrate that LQA offers a practical pathway for robust, privacy-preserving, and efficient VLM deployment on edge devices.
Surface electromyographic hand gesture recognition has gained significant attention in recent years, especially within the field of human-computer interfaces. However, cross-subject tasks remain challenging due to inherent individual differences. To address this, a novel approach for hand gesture recognition is proposed that leverages a subject-generalized variational autoencoder. This approach involves an extended variational autoencoder designed to disentangle input data into three distinct feature-specific representations. The primary classifier within the variational autoencoder focuses on gesture recognition, while two auxiliary classifiers work together to extract subject-specific and gesture-specific features. The gesture-specific features capture generalized characteristics applicable across all subjects, enabling direct application to new subjects. To enhance accuracy and stability, a competitive voting strategy is implemented. The effectiveness of the proposed method was evaluated using a dataset comprising six representative gestures performed by eight subjects. Comparative analysis with baseline models shows that our approach outperforms others, demonstrating superior generalization with an average accuracy of 90.52% in cross-subject validation.
OBJECTIVE:This study aimed to assess whether population-level patterns in seizure occurrence previously observed in self-reported diaries, medical records, and electroencephalographic recordings were also present in tonic-clonic seizure (TCS) diaries produced via the combined input of a US Food and Drug Administration-cleared wristband with an artificial intelligence detection algorithm and patient self-reports. We also investigated the characteristics of patient interactions with wearable seizure alerts. METHODS:We analyzed wristband data from patients with TCSs who had at least three reported TCSs over a minimum of 90 days. We quantified TCS frequency and cycles, and the relationship between the mean and variability of monthly TCS counts. We also assessed interaction metrics such as false alarm dismissal and seizure confirmation rates. RESULTS:Applying strict criteria for usable data, we reviewed 137 490 TCSs from 3012 patients, with a median length of TCS alert records of 445 days (range = 90-1806). Analyses showed consistency between prior diary studies and the present data concerning (1) the distribution of monthly TCS frequency (median = 3.1, range = .08-26); (2) the linear relationship (slope = .79, R2 = .83) between the logarithm of the mean and the logarithm of the SD of monthly TCS frequency (L-relationship); and (iii) the prevalence of multiple coexisting seizure cycles, including circadian (84.0%), weekly (24.6%), and long-term cycles (31.1%). SIGNIFICANCE:Key population-level patterns in seizure occurrence are recapitulated in wrist-worn device recordings, supporting their validity for tracking TCS burden. Compared to other approaches, wearables can provide noninvasive, objective, long-term data, revealing cycles in seizure risk. However, improved patient engagement with wristband alerts and further validation of detection accuracy in ambulatory settings are needed. Together, these findings suggest that data from smart wristbands may be used to derive features of TCS records and, ultimately, facilitate remote monitoring and the development of personalized forecasting tools for TCS management. Our findings may not generalize to other types of seizures.
Generative training has been demonstrated to be powerful for building visual-language models. However, on zero-shot discriminative benchmarks, there is still a performance gap between models trained with generative and discriminative objectives. In this paper, we aim to narrow this gap by improving the efficacy of generative training on classification tasks, without any finetuning processes or additional modules. Specifically, we focus on narrowing the gap between the generative captioner and the CLIP classifier. We begin by analysing the predictions made by the captioner and classifier and observe that the caption generation inherits the distribution bias from the language model trained with pure text modality, making it less grounded on the visual signal. To tackle this problem, we redesign the scoring objective for the captioner to alleviate the distributional bias and focus on measuring the gain of information brought by the visual inputs. We further design a generative training objective to match the evaluation objective. We name our model trained and evaluated from the novel procedures as Information Gain (IG) captioner. We pretrain the models on the public Laion-5B dataset and perform a series of discriminative evaluations. For the zero-shot classification on ImageNet, IG captioner achieves $> 18\%$ improvements over the standard captioner, achieving comparable performances with the CLIP classifier. IG captioner also demonstrated strong performance on zero-shot image-text retrieval tasks on MSCOCO and Flickr30K. We hope this paper inspires further research towards unifying generative and discriminative training procedures for visual-language models.
In the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system that effectively leverages the parallelization capabilities of accelerator hardware. Our approach combines the Universal Speech Model (USM) and the PaLM 2 language model in per-segment scoring mode, achieving an average relative WER improvement across all languages of 10.8% on FLEURS and 3.6% on YouTube captioning. Furthermore, our comprehensive ablation study analyzes key parameters such as LLM size, context length, vocabulary size, fusion methodology. For instance, we explore the impact of LLM size ranging from 128M to 340B parameters on ASR performance. This study provides valuable insights into the factors influencing the effectiveness of practical large-scale LM-fused speech recognition systems.
The success of large language models (LLMs) has prompted efforts to integrate speech and audio data, aiming to create general foundation models capable of processing both textual and non-textual inputs. Recent advances, such as GPT-4o, highlight the potential for end-to-end speech LLMs, which preserves non-semantic information and world knowledge for deeper speech understanding. To guide the development of speech LLMs, we propose a five-level roadmap, ranging from basic automatic speech recognition (ASR) to advanced superhuman models capable of integrating non-semantic information with abstract acoustic knowledge for complex tasks. Moreover, we design a benchmark, SAGI Bechmark, that standardizes critical aspects across various tasks in these five levels, uncovering challenges in using abstract acoustic knowledge and completeness of capability. Our findings reveal gaps in handling paralinguistic cues and abstract acoustic knowledge, and we offer future directions. This paper outlines a roadmap for advancing speech LLMs, introduces a benchmark for evaluation, and provides key insights into their current limitations and potential.
We introduce a multilingual speaker change detection model (USM-SCD) that can simultaneously detect speaker turns and perform ASR for 96 languages. This model is adapted from a speech foundation model trained on a large quantity of supervised and unsupervised data, demonstrating the utility of fine-tuning from a large generic foundation model for a downstream task. We analyze the performance of this multilingual speaker change detection model through a series of ablation studies. We show that the USM-SCD model can achieve more than 75% average speaker change detection F1 score across a test set that consists of data from 96 languages. On American English, the USM-SCD model can achieve an 85.8% speaker change detection F1 score across various public and internal test sets, beating the previous monolingual baseline model by 21% relative. We also show that we only need to fine-tune one-quarter of the trainable model parameters to achieve the best model performance. The USM-SCD model exhibits state-of-the-art ASR quality compared with a strong public ASR baseline, making it suitable to handle both tasks with negligible additional computational cost.
Style transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from reference singing voice samples. However, the endeavor to model the intricate nuances of singing voice styles is an arduous task, as singing voices possess a remarkable degree of expressiveness. Moreover, existing SVS methods encounter a decline in the quality of synthesized singing voices in OOD scenarios, as they rest upon the assumption that the target vocal attributes are discernible during the training phase. To overcome these challenges, we propose StyleSinger, the first singing voice synthesis model for zero-shot style transfer of out-of-domain reference singing voice samples. StyleSinger incorporates two critical approaches for enhanced effectiveness: 1) the Residual Style Adaptor (RSA) which employs a residual quantization module to capture diverse style characteristics in singing voices, and 2) the Uncertainty Modeling Layer Normalization (UMLN) to perturb the style attributes within the content representation during the training phase and thus improve the model generalization. Our extensive evaluations in zero-shot style transfer undeniably establish that StyleSinger outperforms baseline models in both audio quality and similarity to the reference singing voice samples. Access to singing voice samples can be found at https://stylesinger.github.io/.
Objective: Impulse control behaviors (ICBs) and apathy are believed to represent opposite motivational expressions of the same behavioral spectrum involving hypo- and hyperdopaminergic status, but this has been recently debated. Our study aims to estimate the co-occurrence of ICBs and apathy in early Parkinson's disease (PD) and to determine whether this complex neuropsychiatric condition is an important marker of PD prognoses. Methods: Neuropsychiatric symptoms, clinical data, neuroimaging results, and demographic data from de novo PD patients were obtained from the Parkinson's Progression Markers Initiative, a prospective, multicenter, observational cohort. The clinical characteristics of ICBs co-occurring with apathy and their prevalence were analyzed. We compared the prognoses of the different groups during the 8-year follow-up. Multivariate Cox regression analysis was conducted to predict the development of levodopa-induced dyskinesia (LID) using baseline neuropsychiatric symptoms. Results: A total of 422 PD patients and 195 healthy controls (HCs) were included. In brief, 87 (20.6 %) de novo PD patients and 37 (19.0 %) HCs had ICBs at baseline. Among them, 23 (26.4 %) de novo PD patients and 3 (8.1 %) HCs had clinical symptoms of both ICBs and apathy. The ICBs and apathy group had more severe non-motor symptoms than the isolated ICBs group. Cox regression analysis demonstrated that the co-occurrence of ICBs and apathy was a risk factor for LID development (HR 2.229, 95 % CI 1.209 to 4.110, p = 0.010). Conclusions: Co-occurrence of ICBs and apathy is common in patients with early PD and may help to identify the risk of LID development.
Objective: To characterize sleep duration and investigate its association with quality of life among Parkinson's Disease (PD) patients. Methods: In this multicenter cross-sectional study, 970 PD patients were divided into five groups based on selfreported sleep duration: <5,>= 5 to <6, >= 6 to <7, >= 7 to <= 8, and >8 h. The quality of life was evaluated using the 39-Item Parkinson's Disease Questionnaire (PDQ-39). Multivariable linear regression analysis, subgroup analysis, and mediation analysis were conducted to examine the association between sleep duration and quality of life. Results: In multivariable linear regression model, patients with sleep duration (<5 h) had significantly higher PDQ-39 scores (beta = 8.132, 95 % CI: 3.99 to 12.266), especially in mobility, activities of daily living, emotional well-being, stigma, social support, cognition, communication, and bodily discomfort (p < 0.05). The association between sleep duration (<5 h) and worse quality of life was more pronounced in patients with higher HY stage, longer disease duration, and sleep disorders. Moreover, a significant indirect effect of sleep duration (<5 h) on quality of life was observed, with UPDRS I, UPDRS II, and UPDRS IV scores acting as mediators. Conclusions: Short sleep duration (<5 h) is associated with worse quality of life among PD patients. This association was stronger among patients with advanced PD and sleep disorders, while non-motor symptoms and motor complications were identified as significant mediators in this association. These findings highlight the significance of adequate sleep duration and suitable interventions for sleep may help improve quality of life.
Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have been constrained to a single modality; however in the case of video data there is a wealth of other metadata that may be beneficial for this task. In this work, we propose MuSeLI, a Multimodal Spoken Language Identification method, which delves into the use of various metadata sources to enhance language identification. Our study reveals that metadata such as video title, description and geographic location provide substantial information to identify the spoken language of the multimedia recording. We conduct experiments using two diverse public datasets of YouTube videos, and obtain state-of-the-art results on the language identification task. We additionally conduct an ablation study that describes the distinct contribution of each modality for language recognition.
患者 男性,42岁.主因左手抖动18个月,左下肢抖动伴运动迟缓6个月,于2023年7月8日入院.患者18个月前(2022年1月)无明显诱因出现左上肢不自主抖动,静止时出现、持物及动作时消失,无动作迟缓、反应变慢等其他伴随症状.于
Thermoplastic polypropylene (PP) insulated cables, an alternative to cross-linked polyethylene, offer superior insulation, high operating temperature, recyclability, cost-effectiveness, and a limitless cable length. However, challenges such as brittleness at low temperatures and limited flexibility at room temperature impede the application of PP in the field of cable insulation. To address these issues, in-reactor alloy technology seems to be a promising strategy, creating a multiphase system with intrinsic elastomer dispersion in a homopolypropylene matrix. Most of the research on PP-based multiphase systems focuses on enhancing mechanical properties by controlling microscopic structures. A comprehensive understanding of structural evolution during processing and its correlation with the electrical performance of PP thermoplastic insulation materials remains in its infancy. In this study, PP in-reactor alloys with intrinsic elastomers were utilized as model polymeric materials. A novel technology of "melting extrusion-hot stretching-thermal annealing" was employed to manipulate the elastomer phase morphology and crystalline structure. Severe interfacial mismatch during hot stretching initially compromised the mechanical and electrical properties. After thermal annealing, the mechanical and electrical properties were recovered, arising from the reduced rubber deformation and increased crystalline reorganization. The work presented here is expected to help our understanding of the dependence of electrical and mechanical properties on the microstructure of PP in-reactor alloys, providing a valuable reference for the structural design of cable insulation.
GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision, and audio, meaning all inputs and outputs are processed by the same neural network. GPT-4o can respond to audio inputs in as little as 232 milliseconds, with an average of 320 milliseconds, which is similar to human response time in conversation. It matches GPT-4 Turbo performance on text in English and code, with significant improvement on text in non-English languages, while also being much faster and 50% cheaper in the API. GPT-4o is especially better at vision and audio understanding compared to existing models. In line with our commitment to building AI safely and consistent with our voluntary commitments to the White House, we are sharing the GPT-4o System Card, which includes our Preparedness Framework evaluations. In this System Card, we provide a detailed look at GPT-4o's capabilities, limitations, and safety evaluations across multiple categories, focusing on speech-to-speech while also evaluating text and image capabilities, and measures we've implemented to ensure the model is safe and aligned. We also include third-party assessments on dangerous capabilities, as well as discussion of potential societal impacts of GPT-4o's text and vision capabilities.
Recently, a number of approaches to train speech models by incorporating text into end-to-end models have been developed, with Maestro advancing state-of-the-art automatic speech recognition (ASR) and Speech Translation (ST) performance. In this paper, we expand our understanding of the resulting shared speech-text representations with two types of analyses. First we examine the limits of speech-free domain adaptation, finding that a corpus-specific duration model for speech-text alignment is the most important component for learning a shared speech-text representation. Second, we inspect the similarities between activations of unimodal (speech or text) encoders as compared to the activations of a shared encoder. We find that the shared encoder learns a more compact and overlapping speech-text representation than the uni-modal encoders. We hypothesize that this partially explains the effectiveness of the Maestro shared speech-text representations.
End-to-end models with large capacity have significantly improved multilingual automatic speech recognition, but their computation cost poses challenges for on-device applications. We propose a streaming truly multilingual Conformer incorporating mixture-of-expert (MoE) layers that learn to only activate a subset of parameters in training and inference. The MoE layer consists of a softmax gate which chooses the best two experts among many in forward propagation. The proposed MoE layer offers efficient inference by activating a fixed number of parameters as the number of experts increases. We evaluate the proposed model on a set of 12 languages, and achieve an average 11.9% relative improvement in WER over the baseline. Compared to an adapter model using ground truth information, our MoE model achieves similar WER and activates similar number of parameters but without any language information. We further show around 3% relative WER improvement by multilingual shallow fusion.
This paper introduces a new speech dataset called ``LibriTTS-R'' designed for text-to-speech (TTS) use. It is derived by applying speech restoration to the LibriTTS corpus, which consists of 585 hours of speech data at 24 kHz sampling rate from 2,456 speakers and the corresponding texts. The constituent samples of LibriTTS-R are identical to those of LibriTTS, with only the sound quality improved. Experimental results show that the LibriTTS-R ground-truth samples showed significantly improved sound quality compared to those in LibriTTS. In addition, neural end-to-end TTS trained with LibriTTS-R achieved speech naturalness on par with that of the ground-truth samples. The corpus is freely available for download from \url{http://www.openslr.org/141/}.
We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM freezes the pretrained foundation models to maximally preserves their capabilities, and only trains a simple adapter with just 1\% (156M) of the foundation models' parameters. This adaptation not only leads SLM to achieve strong performance on conventional tasks such as speech recognition (ASR) and speech translation (AST), but also introduces the novel capability of zero-shot instruction-following for more diverse tasks: given a speech input and a text instruction, SLM is able to perform unseen generation tasks including contextual biasing ASR using real-time context, dialog generation, speech continuation, and question answering, etc. Our approach demonstrates that the representational gap between pretrained speech and language models might be narrower than one would expect, and can be bridged by a simple adaptation mechanism. As a result, SLM is not only efficient to train, but also inherits strong capabilities already acquired in foundation models of different modalities.