We develop Polaris, the first safety-focused LLM constellation for real-time patient-AI healthcare conversations. Unlike prior LLM works in healthcare focusing on tasks like question answering, our work specifically focuses on long multi-turn voice conversations. Our one-trillion parameter constellation system is composed of several multibillion parameter LLMs as co-operative agents: a stateful primary agent that focuses on driving an engaging conversation and several specialist support agents focused on healthcare tasks performed by nurses to increase safety and reduce hallucinations. We develop a sophisticated training protocol for iterative co-training of the agents that optimize for diverse objectives. We train our models on proprietary data, clinical care plans, healthcare regulatory documents, medical manuals, and other medical reasoning documents. We align our models to speak like medical professionals, using organic healthcare conversations and simulated ones between patient actors and experienced nurses. This allows our system to express unique capabilities such as rapport building, trust building, empathy and bedside manner. Finally, we present the first comprehensive clinician evaluation of an LLM system for healthcare. We recruited over 1100 U.S. licensed nurses and over 130 U.S. licensed physicians to perform end-to-end conversational evaluations of our system by posing as patients and rating the system on several measures. We demonstrate Polaris performs on par with human nurses on aggregate across dimensions such as medical safety, clinical readiness, conversational quality, and bedside manner. Additionally, we conduct a challenging task-based evaluation of the individual specialist support agents, where we demonstrate our LLM agents significantly outperform a much larger general-purpose LLM (GPT-4) as well as from its own medium-size class (LLaMA-2 70B).
Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, HealthSearchQA. We propose a human evaluation framework for model answers along multiple axes including factuality, comprehension, reasoning, possible harm and bias. In addition, we evaluate Pathways Language Model 1 (PaLM, a 540-billion parameter LLM) and its instruction-tuned variant, Flan-PaLM 2 on MultiMedQA. Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA 3 , MedMCQA 4 , PubMedQA 5 and Measuring Massive Multitask Language Understanding (MMLU) clinical topics 6 ), including 67.6% accuracy on MedQA (US Medical Licensing Exam-style questions), surpassing the prior state of the art by more than 17%. However, human evaluation reveals key gaps. To resolve this, we introduce instruction prompt tuning, a parameter-efficient approach for aligning LLMs to new domains using a few exemplars. The resulting model, Med-PaLM, performs encouragingly, but remains inferior to clinicians. We show that comprehension, knowledge recall and reasoning improve with model scale and instruction prompt tuning, suggesting the potential utility of LLMs in medicine. Our human evaluations reveal limitations of today’s models, reinforcing the importance of both evaluation frameworks and method development in creating safe, helpful LLMs for clinical applications.
Harmonization of local source concepts to standard clinical terminologies is a prerequisite for multi-center data aggregation and sharing. Challenges in automating the mapping process stem from the idiosyncratic source encoding schemes adopted by different health systems and the lack of large publicly available training data. In this study, we aim to develop a scalable and generalizable machine learning tool to facilitate standardizing laboratory observations to the Logical Observation Identifiers Names and Codes (LOINC). Specifically, we leverage the contextual embedding from pre-trained T5 models and propose a two-stage fine-tuning strategy based on contrastive learning to enable learning in a few-shot setting without manual feature engineering. Our method utilizes unlabeled general LOINC ontology and data augmentation to achieve high accuracy on retrieving the most relevant LOINC targets when limited amount of labeled data are available. We further show that our model generalizes well to unseen targets. Taken together, our approach shows great potential to reduce manual effort in LOINC standardization and can be easily extended to mapping other terminologies.
Objectives Few machine learning (ML) models are successfully deployed in clinical practice. One of the common pitfalls across the field is inappropriate problem formulation: designing ML to fit the data rather than to address a real-world clinical pain point. Methods We introduce a practical toolkit for user-centred design consisting of four questions covering: (1) solvable pain points, (2) the unique value of ML (eg, automation and augmentation), (3) the actionability pathway and (4) the model’s reward function. This toolkit was implemented in a series of six participatory design workshops with care managers in an academic medical centre. Results Pain points amenable to ML solutions included outpatient risk stratification and risk factor identification. The endpoint definitions, triggering frequency and evaluation metrics of the proposed risk scoring model were directly influenced by care manager workflows and real-world constraints. Conclusions Integrating user-centred design early in the ML life cycle is key for configuring models in a clinically actionable way. This toolkit can guide problem selection and influence choices about the technical setup of the ML problem.
During the last decade, a dramatic rise in the development and application of artificial intelligence (AI) tools for use in pathology services has occurred. This trend is often expected to continue and reshape the field of pathology in the coming years. The deployment of computational pathology and applications of AI tools can be considered as a paradigm shift that will change pathology services, making them more efficient and capable of meeting the needs of this era of precision medicine. Despite the success of AI models, the translational process from discovery to clinical applications has been slow. The gap between self-contained research and clinical environment may be too wide and has been largely neglected. In this review, we cover the current and prospective applications of AI in pathology. We examine its applications in diagnosis and prognosis, and we offer insights for considerations that could improve clinical applicability of these tools. Then, we discuss its potential to improve workflow efficiency, and its benefits in pathologist education. Finally, we review the factors that could influence adoption in clinical practices and the associated regulatory processes.
Background Breast cancer management depends on biomarkers including estrogen receptor, progesterone receptor, and human epidermal growth factor receptor 2 (ER/PR/HER2). Though existing scoring systems are widely used and well-validated, they can involve costly preparation and variable interpretation. Additionally, discordances between histology and expected biomarker findings can prompt repeat testing to address biological, interpretative, or technical reasons for unexpected results. Methods We developed three independent deep learning systems (DLS) to directly predict ER/PR/HER2 status for both focal tissue regions (patches) and slides using hematoxylin-and-eosin-stained (H&E) images as input. Models were trained and evaluated using pathologist annotated slides from three data sources. Areas under the receiver operator characteristic curve (AUCs) were calculated for test sets at both a patch-level (>135 million patches, 181 slides) and slide-level ( n = 3274 slides, 1249 cases, 37 sites). Interpretability analyses were performed using Testing with Concept Activation Vectors (TCAV), saliency analysis, and pathologist review of clustered patches. Results The patch-level AUCs are 0.939 (95%CI 0.936–0.941), 0.938 (0.936–0.940), and 0.808 (0.802–0.813) for ER/PR/HER2, respectively. At the slide level, AUCs are 0.86 (95%CI 0.84–0.87), 0.75 (0.73–0.77), and 0.60 (0.56–0.64) for ER/PR/HER2, respectively. Interpretability analyses show known biomarker-histomorphology associations including associations of low-grade and lobular histology with ER/PR positivity, and increased inflammatory infiltrates with triple-negative staining. Conclusions This study presents rapid breast cancer biomarker estimation from routine H&E slides and builds on prior advances by prioritizing interpretability of computationally learned features in the context of existing pathological knowledge.
Breast cancer is the most common cancer and second leading cause of cancer-related death worldwide. The mainstay of breast cancer workup is histopathological diagnosis - which guides therapy and prognosis. However, emerging knowledge about the complex nature of cancer and the availability of tailored therapies have exposed opportunities for improvements in diagnostic precision. In parallel, advances in artificial intelligence (AI) along with the growing digitization of pathology slides for the primary diagnosis are a promising approach to meet the demand for more accurate detection, classification and prediction of behaviour of breast tumours. In this article, we cover the current and prospective uses of AI in digital pathology for breast cancer, review the basics of digital pathology and AI, and outline outstanding challenges in the field.
BACKGROUND:Limited-channel EEG research in neonates is hindered by lack of open, accessible analytic tools. To overcome this limitation, we have created the Washington University-Neonatal EEG Analysis Toolbox (WU-NEAT), containing two of the most commonly used tools, provided in an open-source, clinically-validated package running within MATLAB.METHODS:The first algorithm is the amplitude-integrated EEG (aEEG), which is generated by filtering, rectifying and time-compressing the original EEG recording, with subsequent semi-logarithmic display. The second algorithm is the spectral edge frequency (SEF), calculated as the critical frequency below which a user-defined proportion of the EEG spectral power is located. The aEEG algorithm was validated by three experienced reviewers. Reviewers evaluated aEEG recordings of fourteen preterm/term infants, displayed twice in random order, once using a reference algorithm and again using the WU-NEAT aEEG algorithm. Using standard methodology, reviewers assigned a background pattern classification. Inter/intra-rater reliability was assessed. For the SEF, calculations were made using the same fourteen recordings, first with the reference and then with the WU-NEAT algorithm. Results were compared using Pearson's correlation coefficient.RESULTS:For the aEEG algorithm, intra- and inter-rater reliability was 100% and 98%, respectively. For the SEF, the mean±SD Pearson correlation coefficient between algorithms was 0.96±0.04.CONCLUSION:We have demonstrated a clinically-validated toolbox for generating the aEEG as well as calculating the SEF from EEG data. Open-source access will enable widespread use of common analytic algorithms which are device-independent and unlikely to become outdated as technology changes, thereby facilitating future collaborative research in neonatal EEG.
OBJECTIVE Electrical stimulation of peripheral nerve tissue has been shown to accelerate axonal regeneration. Yet existing methods of applying electrical stimulation to injured peripheral nerves have presented significant barriers to clinical translation. In this study, the authors examined the use of a novel implantable wireless nerve stimulator capable of simultaneously delivering therapeutic electrical stimulation of injured peripheral nerve tissue and providing postoperative serial assessment of functional recovery. METHODS Flexible wireless stimulators were fabricated and implanted into Lewis rats. Thin-film implants were used to deliver brief electrical stimulation (1 hour, 20 Hz) to sciatic nerves after nerve crush or nerve transection-and-repair injuries. RESULTS Electrical stimulation of injured nerves via implanted wireless stimulators significantly improved functional recovery. Brief electrical stimulation was observed to increase the rate of functional recovery after both nerve crush and nerve transection-and-repair injuries. Wireless stimulators successfully facilitated therapeutic stimulation of peripheral nerve tissue and serial assessment of nerve recovery. CONCLUSIONS Implantable wireless stimulators can deliver therapeutic electrical stimulation to injured peripheral nerve tissue. Implantable wireless nerve stimulators might represent a novel means of facilitating therapeutic electrical stimulation in both intraoperative and postoperative settings.
It is critical to understand the privacy and robustness vulnerabilities of machine learning models, as their implementation expands in scope. In membership inference attacks, adversaries can determine whether a particular set of data was used in training, putting the privacy of the data at risk. Existing work has mostly focused on image related tasks; we generalize this type of attack to speaker identification on audio samples. We demonstrate attack precision of 85.9\% and recall of 90.8\% for LibriSpeech, and 78.3\% precision and 90.7\% recall for VOiCES (Voices Obscured in Complex Environmental Settings). We find that implementing defenses such as prediction obfuscation, defensive distillation or adversarial training, can reduce attack accuracy to chance.
Adversarial training was introduced as a way to improve the robustness of deep learning models to adversarial attacks. This training method improves robustness against adversarial attacks, but increases the models vulnerability to privacy attacks. In this work we demonstrate how model inversion attacks, extracting training data directly from the model, previously thought to be intractable become feasible when attacking a robustly trained model. The input space for a traditionally trained model is dominated by adversarial examples - data points that strongly activate a certain class but lack semantic meaning - this makes it difficult to successfully conduct model inversion attacks. We demonstrate this effect using the CIFAR-10 dataset under three different model inversion attacks, a vanilla gradient descent method, gradient based method at different scales, and a generative adversarial network base attacks.
Peripheral nerve injuries represent a significant problem in public health, constituting 2–5% of all trauma cases 1 . For severe nerve injuries, even advanced forms of clinical intervention often lead to incomplete and unsatisfactory motor and/or sensory function 2 . Numerous studies report the potential of pharmacological approaches (for example, growth factors, immunosuppressants) to accelerate and enhance nerve regeneration in rodent models 3 – 10 . Unfortunately, few have had a positive impact in clinical practice. Direct intraoperative electrical stimulation of injured nerve tissue proximal to the site of repair has been demonstrated to enhance and accelerate functional recovery 11 , 12 , suggesting a novel nonpharmacological, bioelectric form of therapy that could complement existing surgical approaches. A significant limitation of this technique is that existing protocols are constrained to intraoperative use and limited therapeutic benefits 13 . Herein we introduce (i) a platform for wireless, programmable electrical peripheral nerve stimulation, built with a collection of circuit elements and substrates that are entirely bioresorbable and biocompatible, and (ii) the first reported demonstration of enhanced neuroregeneration and functional recovery in rodent models as a result of multiple episodes of electrical stimulation of injured nervous tissue.
We propose an algorithm to denoise speakers from a single microphone in the presence of non-stationary and dynamic noise. Our approach is inspired by the recent success of neural network models separating speakers from other speakers and singers from instrumental accompaniment. Unlike prior art, we leverage embedding spaces produced with source-contrastive estimation, a technique derived from negative sampling techniques in natural language processing, while simultaneously obtaining a continuous inference mask. Our embedding space directly optimizes for the discrimination of speaker and noise by jointly modeling their characteristics. This space is generalizable in that it is not speaker or noise specific and is capable of denoising speech even if the model has not seen the speaker in the training set. Parameters are trained with dual objectives: one that promotes a selective bandpass filter that eliminates noise at time-frequency positions that exceed signal power, and another that proportionally splits time-frequency content between signal and noise. We compare to state of the art algorithms as well as traditional sparse non-negative matrix factorization solutions. The resulting algorithm avoids severe computational burden by providing a more intuitive and easily optimized approach, while achieving competitive accuracy.
Genomic data are becoming increasingly valuable as we develop methods to utilize the information at scale and gain a greater understanding of how genetic information relates to biological function. Advances in synthetic biology and the decreased cost of sequencing are increasing the amount of privately held genomic data. As the quantity and value of private genomic data grows, so does the incentive to acquire and protect such data, which creates a need to store and process these data securely. We present an algorithm for the Secure Interrogation of Genomic DataBases (SIG-DB). The SIG-DB algorithm enables databases of genomic sequences to be searched with an encrypted query sequence without revealing the query sequence to the Database Owner or any of the database sequences to the Querier. SIG-DB is the first application of its kind to take advantage of locality-sensitive hashing and homomorphic encryption to allow generalized sequence-to-sequence comparisons of genomic data.
Myelin is a multilamellar sheath generated by specialized glia called Schwann cells (SCs) in the peripheral nervous system (PNS), which serves to protect and insulate axons for rapid neuronal signaling. In zebrafish and rodent models, we identify GPR56/ADGRG1 as a conserved regulator of PNS development and health. We demonstrate that, during SC development, GPR56-dependent RhoA signaling promotes timely radial sorting of axons. In the mature PNS, GPR56 is localized to distinct SC cytoplasmic domains, is required to establish proper myelin thickness, and facilitates organization of the myelin sheath. Furthermore, we define plectin—a scaffolding protein previously linked to SC domain organization, myelin maintenance, and a series of disorders termed “plectinopathies”—as a novel interacting partner of GPR56. Finally, we show that Gpr56 mutants develop progressive neuropathy-like symptoms, suggesting an underlying mechanism for peripheral defects in some human patients with GPR56 mutations. In sum, we define Gpr56 as a new regulator in the development and maintenance of peripheral myelin.
The speakers in the room (SITR) corpus is a collaboration between Lab41 and SRI International, designed to be a freely available data set for speech and acoustics research in noisy room conditions. The main focus of the corpus is on distant microphone collection in a series of four rooms of different sizes and configurations. There are both foreground speech and background adversarial sounds, played through high-quality speakers in each room to create multiple, realistic acoustic environments. The foreground speech is played from a randomly rotating speaker to emulate head motion. Foreground speech consists of files from LibriVox audio collections and the background distractor sounds will consist of babble, music, HVAC, TV/radio, dogs, vehicles, and weather sounds drawn from the MUSAN collection. Each room has multiple sessions to exhaustively cover the background foreground combinations, and the audio is collected with twelve different microphones (omnidirectional lavalier, studio cardioid, and piezoelectric) placed strategically around the room. The resulting data set was designed to enable acoustic research on event detection, background detection, source separation, speech enhancement, source distance, sound localization, as well as speech research on speaker recognition, speech activity detection, speech recognition, and language recognition.
Pseudarthrosis is an exceedingly common, costly, and morbid complication in the treatment of long bone fractures and after spinal fusion surgery. Electrical bone growth stimulation (EBGS) presents a unique approach to accelerate healing and promote fusion success rates. Over the past three decades, increased experience and widespread use of EBGS devices has led to significant improvements in stimulation paradigms and clinical outcomes. In this paper, we comprehensively review the literature and examine the history, scientific evidence, available technology, and clinical applications for EBGS. We summarize indications, limitations, and provide an overview of cost-effectiveness and future directions of EBGS technology. Various models of electrical stimulation have been proposed and marketed as adjuncts for spinal fusions and long bone fractures. Clinical studies show variable safety and efficacy of EBGS under different conditions and clinical scenarios. While the results of clinical trials do not support indiscriminate EBGS utilization for any bone injury, the evidence does suggest that EBGS is desirable and cost efficient for certain orthopedic indications, especially when used in combination with standard, first-line treatments. This review should serve as a reference to inform practicing clinicians of available treatment options, facilitate evidence-based decision making, and provide a platform for further research.
This paper introduces the Voices Obscured in Complex Environmental Settings (VOiCES) corpus, a freely available dataset under Creative Commons BY 4.0. This dataset will promote speech and signal processing research of speech recorded by far-field microphones in noisy room conditions. Publicly available speech corpora are mostly composed of isolated speech at close-range microphony. A typical approach to better represent realistic scenarios, is to convolve clean speech with noise and simulated room response for model training. Despite these efforts, model performance degrades when tested against uncurated speech in natural conditions. For this corpus, audio was recorded in furnished rooms with background noise played in conjunction with foreground speech selected from the Libri-Speech corpus. Multiple sessions were recorded in each room to accommodate for all foreground speech-background noise combinations. Audio was recorded using twelve microphones placed throughout the room, resulting in 120 hours of audio per microphone. This work is a multi-organizational effort led by SRI International and Lab41 with the intent to push forward state-of-the-art distant microphone approaches in signal processing and speech recognition.