Integrating multiple generative foundation models, especially those trained on different modalities, into something greater than the sum of its parts poses significant challenges. Two key hurdles are the availability of aligned data (concepts that contain similar meaning but is expressed differently in different modalities), and effectively leveraging unimodal representations in cross-domain generative tasks, without compromising their original unimodal capabilities. We propose Zipper, a multi-tower decoder architecture that addresses these concerns by using cross-attention to flexibly compose multimodal generative models from independently pre-trained unimodal decoders. In our experiments fusing speech and text modalities, we show the proposed architecture performs very competitively in scenarios with limited aligned text-speech data. We also showcase the flexibility of our model to selectively maintain unimodal (e.g., text-to-text generation) generation performance by freezing the corresponding modal tower (e.g. text). In cross-modal tasks such as automatic speech recognition (ASR) where the output modality is text, we show that freezing the text backbone results in negligible performance degradation. In cross-modal tasks such as text-to-speech generation (TTS) where the output modality is speech, we show that using a pre-trained speech backbone results in superior performance to the baseline.
Simultaneous interpretation is an especially challenging form of translation because it requires converting speech from one language to another in real-time. Though prior work has relied on out-of-the-box machine translation metrics to evaluate interpretation data, we hypothesize that strategies common in high-quality human interpretations, such as summarization, may not be handled well by standard machine translation metrics. In this work, we examine both qualitatively and quantitatively four potential barriers to evaluation of interpretation: disfluency, summarization, paraphrasing, and segmentation. Our experiments reveal that, while some machine translation metrics correlate fairly well with human judgments of interpretation quality, much work is still needed to account for interpretation strategies during evaluation. As a first step to addressing this problem, we develop a fine-tuned model for interpretation evaluation, which achieves better correlation with human judgments than state-of-the-art machine translation metrics.
We introduce AudioPaLM, a large language model for speech understanding and generation. AudioPaLM fuses text-based and speech-based language models, PaLM-2 [Anil et al., 2023] and AudioLM [Borsos et al., 2022], into a unified multimodal architecture that can process and generate text and speech with applications including speech recognition and speech-to-speech translation. AudioPaLM inherits the capability to preserve paralinguistic information such as speaker identity and intonation from AudioLM and the linguistic knowledge present only in text large language models such as PaLM-2. We demonstrate that initializing AudioPaLM with the weights of a text-only large language model improves speech processing, successfully leveraging the larger quantity of text training data used in pretraining to assist with the speech tasks. The resulting model significantly outperforms existing systems for speech translation tasks and has the ability to perform zero-shot speech-to-text translation for many languages for which input/target language combinations were not seen in training. AudioPaLM also demonstrates features of audio language models, such as transferring a voice across languages based on a short spoken prompt. We release examples of our method at https://google-research.github.io/seanet/audiopalm/examples
Current disfluency detection models focus on individual utterances each from a single speaker. However, numerous discontinuity phenomena in spoken conversational transcripts occur across multiple turns, hampering human readability and the performance of downstream NLP tasks. This study addresses these phenomena by proposing an innovative Multi-Turn Cleanup task for spoken conversational transcripts and collecting a new dataset, MultiTurnCleanup1. We design a data labeling schema to collect the high-quality dataset and provide extensive data analysis. Furthermore, we leverage two modeling approaches for experimental evaluation as benchmarks for future research.
In modern interactive speech-based systems, speech is consumed and transcribed incrementally prior to having disfluencies removed. This post-processing step is crucial for producing clean transcripts and high performance on downstream tasks (e.g. machine translation). However, most current state-of-the-art NLP models such as the Transformer operate non-incrementally, potentially causing unacceptable delays. We propose a streaming BERT-based sequence tagging model that, combined with a novel training objective, is capable of detecting disfluencies in real-time while balancing accuracy and latency. This is accomplished by training the model to decide whether to immediately output a prediction for the current input or to wait for further context. Essentially, the model learns to dynamically size its lookahead window. Our results demonstrate that our model produces comparably accurate predictions and does so sooner than our baselines, with lower flicker. Furthermore, the model attains state-of-the-art latency and stability scores when compared with recent work on incremental disfluency detection.
Automatic subtitle translation is an important technology to make video content available across language barriers. Subti-tle translation complicates the normal translation problem by adding the challenge of how to format the system output into subtitles. We propose a simple technique that treats subtitle translation as standard sentence translation plus alignment driven markup transfer, which enables us to reliably maintain timing and formatting information from the source subtitles. We also introduce two metrics to measure the quality of subtitle boundaries: a Timed BLEU that penalizes mistimed tokens with respect to a reference subtitle sequence, and a measure of how much Timed BLEU is lost due to suboptimal subtitle boundary placement. In experiments on TED and YouTube subtitles, we show that we are able to achieve much better translation quality than a baseline that translates each subtitle independently, while coming very close to optimal subtitle boundary placement.
Traditional translation systems trained on written documents perform well for text-based translation but not as well for speech-based applications. We aim to adapt translation models to speech by introducing actual lexical errors from ASR and segmentation errors from automatic punctuation into our translation training data. We introduce an inverted projection approach that projects automatically detected system segments onto human transcripts and then re-segments the gold translations to align with the projected human transcripts. We demonstrate that this overcomes the train-test mismatch present in other training approaches. The new projection approach achieves gains of over 1 BLEU point over a baseline that is exposed to the human transcripts and segmentations, and these gains hold for both IWSLT data and YouTube data.
Diarization partitions an audio stream into segments based on the voices of the speakers. Real-time diarization systems that include an enrollment step should limit enrollment training samples to reduce user interaction time. Although training on a small number of samples yields poor performance, we show that the accuracy can be improved dramatically using a chronological self-training approach. We studied the tradeoff between training time and classification performance and found that 1 second is sufficient to reach over 95% accuracy. We evaluated on 700 audio conversation files of about 10 minutes each from 6 different languages and demonstrated average diarization error rates as low as 10%.
Automatic Speech Recognition (ASR) systems are often optimized to work best for speakers with canonical speech patterns. Unfortunately, these systems perform poorly when tested on atypical speech and heavily accented speech. It has previously been shown that personalization through model fine-tuning substantially improves performance. However, maintaining such large models per speaker is costly and difficult to scale. We show that by adding a relatively small number of extra parameters to the encoder layers via so-called residual adapter, we can achieve similar adaptation gains compared to model fine-tuning, while only updating a tiny fraction (less than 0.5%) of the model parameters. We demonstrate this on two speech adaptation tasks (atypical and accented speech) and for two state-of-the-art ASR architectures.
Neural Machine Translation (NMT) models have demonstrated strong state of the art performance on translation tasks where well-formed training and evaluation data are provided, but they remain sensitive to inputs that include errors of various types. Specifically, in the context of long-form speech translation systems, where the input transcripts come from Automatic Speech Recognition (ASR), the NMT models have to handle errors including phoneme substitutions, grammatical structure, and sentence boundaries, all of which pose challenges to NMT robustness. Through in-depth error analysis, we show that sentence boundary segmentation has the largest impact on quality, and we develop a simple data augmentation strategy to improve segmentation robustness.
The detection and quantification of carotid artery stenosis guides the decision of the need for surgical intervention such as endarterectomy or stenting of a patient. Contrast enhanced ultrasound highlights detailed information on hemodynamics and even on pathophysiology, and 3D scans yield full in-situ information of the complete anatomy. We introduce 3D image analysis algorithms that enable quantification and visualization of the carotid lumen in these scans. Our method enhances the lumen intensity contrast, extracts lumen centerlines using a gradient concentration calculation, and then uses a graph search method to obtain a coarse lumen segmentation, which is refined through a level set method applied on an intensity-corrected image. We processed 35 images acquired from 7 patients, and demonstrated quantitative comparisons with MR/CT lumen segmentations. We obtained a segmentation correlation score of R-2 = 0.94 and a coefficient of 0.99 for US versus MR/CT through a linear regression on effective diameters extracted from slices in the XY-plane.
While pathologists can readily elucidate disease-relevant information from tissue images, automated algorithms may fail to capture the intricate details of complex biological specimens. As histology patterns vary depending on different tissue types, it is typically necessary to adapt and optimize segmentation algorithms to specific applications. To address this, we present a supervised machine learning method we call Support Vector Shape Segmentation (SVSS) to enhance and improve more general segmentation methods by utilizing a cell shape ranking function. First, we pose shape segmentation as an optimization problem that maximizes shape similarity with respect to the specific shape classes. Secondly, we propose a computationally efficient algorithm to solve the multi-scale segmentation problem in a minimum number of steps. The main advantage of the approach is that it naturally induces a ranking measure given the set of shape exemplars. We demonstrate large-scale quantitative and qualitative results on epithelial cells in a range of tissue types.
Radiologists are required to read thousands of patient images every day, and any tools that can improve their workflow to help them make efficient and accurate measurements is of great value. Such an interactive tool must be intuitive to use, and we have found that users are accustomed to clicking on the contour of the object for segmentation and would like the final segmentation to pass through these points. The tool must also be fast to enable real-time interactive feedback. To meet these needs, we present a segmentation workflow that enables an intuitive method for fast interactive segmentation of 2D and 3D objects. Given simple user clicks on the contour of an object in one 2D view, the algorithm generates foreground and background seeds and computes foreground and background distributions that are used to segment the object in 2D. It then propagates the information to the two orthogonal planes in a 3D volume and segments all three 2D views. The automated segmentation is automatically updated as the user continues to add points around the contour, and the algorithm is re-run using the total set of points. Based on the segmented objects in these three views, the algorithm then computes a 3D segmentation of the object. This process requires only limited user interaction to segment complex shapes and significantly improves the workflow of the user.
Wavelet approaches have proven effective in many segmentation applications and in particular in the segmentation of cells, which are blob-like in shape. We build upon an established wavelet segmentation algorithm and demonstrate how to overcome some of its limitations based on the theoretical derivation of the compounding process of iterative convolutions. We demonstrate that the wavelet decomposition can be computed for any desired level directly without iterative decompositions that require additional computation and memory. This is especially important when dealing with large 3D volumes that consume significant amounts of memory and require intense computation. Our approach is generalized to automatically handle both 2D and 3D and also implicitly handles the anisotropic pixel size inherent in such datasets. Our results demonstrate a 28X improvement in speed and 8X improvement in memory efficiency for standard size 3D confocal image volumes without adversely affecting the accuracy.
Introduction: Vulnerable plaque is marked by increased intra-plaque vascularity, hemorrhage and is associated with cerebrovascular (CV) events. Current use of 2D CEUS carotid imaging does not provide a true volumetric representation of the intra-plaque angiogenesis. Therefore, this is the first clinical trial using 3D CEUS to assess intra-plaque angiogenesis, in patients with clinically significant plaques prior to carotid endarterectomy (CEA). Methods: Data were acquired on 11 patients that were asked to participate in this pilot study prior to CEA. All patients had MR and/or CT angiogram as part of their standard care and all carotid artery plaque specimens were collected and sent for histologic evaluation. CEUS volumes were acquired using a GE LOGIQ E9, RSP6-16 4D probe and Optison. In order to assess plaque vulnerability, semi-automatic image analysis algorithms were developed to segment the lumen, plaque, and intra-plaque vascularity. Measurements derived from 3D CEUS images were validated against CT, MR, and histology. Results: Preliminary results from 3 patients were analyzed and lumen volumes were segmented from all available modalities (figure shows CEUS example). Lumen measurements of the CCA, ICA, and ECA from 3D CEUS were compared to MR or CT. High correlation is shown in the figure for these patients (R 2 =0.95). Further analysis of the 3D CEUS images demonstrates that we can identify intra-plaque vascularity (see segmentation from one patient). These CEUS results estimated an average cross-sectional vascular density of 7-11% in the plaque, matching histology estimates of 10% from a single slice. Conclusion: These results demonstrate that it is feasible to use 3D CEUS for quantitative, volumetric imaging of the carotid artery and intra-plaque angiogenesis in patients scheduled for CEA. Future work will extend this analysis to all patients and make a quantitative comparison of intra-plaque vascular density from 3D CEUS vs. multiple histology slices.