In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio. The family includes two new models: (1) an updated Gemini 1.5 Pro, which exceeds the February version on the great majority of capabilities and benchmarks; (2) Gemini 1.5 Flash, a more lightweight variant designed for efficiency with minimal regression in quality. Gemini 1.5 models achieve near-perfect recall on long-context retrieval tasks across modalities, improve the state-of-the-art in long-document QA, long-video QA and long-context ASR, and match or surpass Gemini 1.0 Ultra's state-of-the-art performance across a broad set of benchmarks. Studying the limits of Gemini 1.5's long-context ability, we find continued improvement in next-token prediction and near-perfect retrieval (>99 as Claude 3.0 (200k) and GPT-4 Turbo (128k). Finally, we highlight real-world use cases, such as Gemini 1.5 collaborating with professionals on completing their tasks achieving 26 to 75 categories, as well as surprising new capabilities of large language models at the frontier; when given a grammar manual for Kalamang, a language with fewer than 200 speakers worldwide, the model learns to translate English to Kalamang at a similar level to a person who learned from the same content.
Cross-device federated learning (FL) has been well-studied from algorithmic, system scalability, and training speed perspectives. Nonetheless, moving from centralized training to cross-device FL for millions or billions of devices presents many risks, including performance loss, developer inertia, poor user experience, and unexpected application failures. In addition, the corresponding infrastructure, development costs, and return on investment are difficult to estimate. In this paper, we present a device-cloud collaborative FL platform that integrates with an existing machine learning platform, providing tools to measure real-world constraints, assess infrastructure capabilities, evaluate model training performance, and estimate system resource requirements to responsibly bring FL into production. We also present a decision workflow that leverages the FL-integrated platform to comprehensively evaluate the trade-offs of cross-device FL and share our empirical evaluations of business-critical machine learning applications that impact hundreds of millions of users.
This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging from complex reasoning tasks to on-device memory-constrained use-cases. Evaluation on a broad range of benchmarks shows that our most-capable Gemini Ultra model advances the state of the art in 30 of 32 of these benchmarks - notably being the first model to achieve human-expert performance on the well-studied exam benchmark MMLU, and improving the state of the art in every one of the 20 multimodal benchmarks we examined. We believe that the new capabilities of the Gemini family in cross-modal reasoning and language understanding will enable a wide variety of use cases. We discuss our approach toward post-training and deploying Gemini models responsibly to users through services including Gemini, Gemini Advanced, Google AI Studio, and Cloud Vertex AI.
The ability to use a blood sample to determine the transcriptional state of cells that are releasing DNA into the bloodstream of a patient may be helpful in a variety of clinical applications. Here we present a case study of a gene expression prediction model that uses cell-free DNA (cfDNA) fragment coverage data generated by high-throughput sequencing to predict which genes are highly or lowly expressed in the cells contributing to that cfDNA. We evaluated a number of models, including a convolutional neural network that takes cfDNA fragment information (the density of both fragment midpoint and length by genomic position) over a transcription start site (TSS) as input, and outputs a predicted probability of whether that gene is highly expressed in cfDNA-producing cells. When we trained the convolutional model on a set of 554 genes with TSSs that were either constitutively expressed or unexpressed across leukocyte samples from the NIH Roadmap Epigenome Mapping Consortium, we achieved ~0.97 AUC in cross validation. With other models and splits of the data, we observed AUCs ranging from 0.95 to 0.99 on this gene-expression task. Next, we were interested in whether this trained model could answer specific clinical questions. For example, we hypothesized that we should see an increased influence of colon gene expression profiles in colorectal cancer patients with a higher fraction of circulating tumor DNA. To test this hypothesis, we applied our models to a set of genes with colon-specific expression, which generated a list of probabilities of each gene being expressed in each sample. We then applied simple models on the these lists of probabilities to predict whether a patient had CRC or was healthy. This yielded cross validation AUCs between 0.85 and 0.95 across many of the models we tested in differentiating healthy patients from colorectal cancer patients with tumor fraction over 5%. These results suggest a path forward for modeling transcriptional states using cfDNA sequencing data, which will enable greater insights from cfDNA that could augment those provided by other analytes. Citation Format: John A. St John, Erik Gafni, Brandon White, Ajay Kannan, Loren Hansen, Artur Jaroszewicz, Anshul Kundaje, Nathan Boley. Predicting gene expression from plasma cell-free DNA using both the fragment length and fragment position [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 4349.
Blood-based methods using cell-free DNA (cfDNA) are under development as an alternative to existing screening tests. However, early-stage detection of cancer using tumor-derived cfDNA has proven challenging because of the small proportion of cfDNA derived from tumor tissue in early-stage disease. A machine learning approach to discover signatures in cfDNA, potentially reflective of both tumor and non-tumor contributions, may represent a promising direction for the early detection of cancer. Whole-genome sequencing was performed on cfDNA extracted from plasma samples (N = 546 colorectal cancer and 271 non-cancer controls). Reads aligning to protein-coding gene bodies were extracted, and read counts were normalized. cfDNA tumor fraction was estimated using IchorCNA. Machine learning models were trained using k-fold cross-validation and confounder-based cross-validations to assess generalization performance. In a colorectal cancer cohort heavily weighted towards early-stage cancer (80% stage I/II), we achieved a mean AUC of 0.92 (95% CI 0.91–0.93) with a mean sensitivity of 85% (95% CI 83–86%) at 85% specificity. Sensitivity generally increased with tumor stage and increasing tumor fraction. Stratification by age, sequencing batch, and institution demonstrated the impact of these confounders and provided a more accurate assessment of generalization performance. A machine learning approach using cfDNA achieved high sensitivity and specificity in a large, predominantly early-stage, colorectal cancer cohort. The possibility of systematic technical and institution-specific biases warrants similar confounder analyses in other studies. Prospective validation of this machine learning method and evaluation of a multi-analyte approach are underway.
Human computer interaction facilitates intelligent communication between humans and computers, in which gesture recognition plays a prominent role. This paper proposes a machine learning system to identify dynamic gestures using tri-axial acceleration data acquired from two public datasets. These datasets, uWave and Sony, were acquired using accelerometers embedded in Wii remotes and smartwatches, respectively. A dynamic gesture signed by the user is characterized by a generic set of features extracted across time and frequency domains. The system was analyzed from an end-user perspective and was modelled to operate in three modes. The modes of operation determine the subsets of data to be used for training and testing the system. From an initial set of seven classifiers, three were chosen to evaluate each dataset across all modes rendering the system towards mode-neutrality and dataset-independence. The proposed system is able to classify gestures performed at varying speeds with minimum preprocessing, making it computationally efficient. Moreover, this system was found to run on a low-cost embedded platform - Raspberry Pi Zero (USD 5), making it economically viable.
Statistical learning on biological data can be challenging due to confounding variables in sample collection and processing. Confounders can cause models to generalize poorly and result in inaccurate prediction performance metrics if models are not validated thoroughly. In this paper, we propose methods to control for confounding factors and further improve prediction performance. We introduce OrthoNormal basis construction In cOnfounding factor Normalization (ONION) to remove confounding covariates and use the Domain-Adversarial Neural Network (DANN) to penalize models for encoding confounder information. We apply the proposed methods to simulated and empirical patient data and show significant improvements in generalization.
We present Deep Voice 3, a fully-convolutional attention-based neural text-to-speech (TTS) system. Deep Voice 3 matches state-of-the-art neural speech synthesis systems in naturalness while training ten times faster. We scale Deep Voice 3 to data set sizes unprecedented for TTS, training on more than eight hundred hours of audio from over two thousand speakers. In addition, we identify common error modes of attention-based speech synthesis networks, demonstrate how to mitigate them, and compare several different waveform synthesis methods. We also describe how to scale inference to ten million queries per day on one single-GPU server.
Introduction: Despite population screening and availability of several stool-based, non-invasive screening methods, over 20% of colorectal cancers (CRC) in the US are metastatic at the time of diagnosis. Blood-based methods using cell-free DNA (cfDNA) are under development as an alternative to stool-based tests. However, early stage detection of cancer using tumor-derived mutations in cfDNA (circulating tumor DNA, or ctDNA) has proven challenging because of the small proportion of cfDNA derived from tumor tissue (tumor fraction, ctDNA/cfDNA ratio) in early stage disease. Using an artificial intelligence-driven approach based on machine learning (ML) to discover signatures in cfDNA potentially reflective of both tumor and immune contributions may represent a promising direction for the early detection of cancer. Methods: De-identified plasma samples (N=1,040) were received from academic clinical studies and commercial biobanks (n=579 CRC patients; n=461 controls). Whole-genome sequencing was performed to >50M reads on cfDNA extracted from plasma. Reads aligning to expressed sequences in the genome were extracted and read counts were normalized to account for variability in read depth, sequencecontent bias, and technical batch effects. cfDNA tumor fraction was estimated using IchorCNA. ML models were trained using 10-fold cross-validation stratified by sequencing batch to mitigate bias from sequencing batch effects. Results: In a cohort heavily weighted towards early stage cancer (82% stage I/II), our method achieves a sensitivity in cross-validation of 81% (Clopper-Pearson 95% confidence interval, 77-84%) at 85% specificity. Sensitivity generally increased by tumor stage. Stratification by sequencing batch was required for reliable generalization. Further analyses revealed susceptibility to additional confounders, including variation in preanalytical and analytical processes such as institution-specific blood collection protocols. Downsampling the dataset to balance with respect to such confounders can reduce sensitivity at 85% specificity by 20-30%. Conclusion: An ML approach using a single analyte was able to achieve high sensitivity and specificity in an predominantly early-stage CRC cohort. The observation of systematic technical and site-specific biases warrants similar confounder analyses in other retrospective studies. Prospective validation of the presented ML method and evaluation of a previously presented multi-analyte approach are underway.
Authors often convey meaning by referring to or imitating prior works of literature, a process that creates complex networks of literary relationships ("intertextuality") and contributes to cultural evolution. In this paper, we use techniques from stylometry and machine learning to address subjective literary critical questions about Latin literature, a corpus marked by an extraordinary concentration of intertextuality. Our work, which we term "quantitative criticism," focuses on case studies involving two influential Roman authors, the playwright Seneca and the historian Livy. We find that four plays related to but distinct from Seneca's main writings are differentiated from the rest of the corpus by subtle but important stylistic features. We offer literary interpretations of the significance of these anomalies, providing quantitative data in support of hypotheses about the use of unusual formal features and the interplay between sound and meaning. The second part of the paper describes a machine-learning approach to the identification and analysis of citational material that Livy loosely appropriated from earlier sources. We extend our approach to map the stylistic topography of Latin prose, identifying the writings of Caesar and his near-contemporary Livy as an inflection point in the development of Latin prose style. In total, our results reflect the integration of computational and humanistic methods to investigate a diverse range of literary questions.
We present Deep Speaker, a neural speaker embedding system that maps utterances to a hypersphere where speaker similarity is measured by cosine similarity. The embeddings generated by Deep Speaker can be used for many tasks, including speaker identification, verification, and clustering. We experiment with ResCNN and GRU architectures to extract the acoustic features, then mean pool to produce utterance-level speaker embeddings, and train using triplet loss based on cosine similarity. Experiments on three distinct datasets suggest that Deep Speaker outperforms a DNN-based i-vector baseline. For example, Deep Speaker reduces the verification equal error rate by 50 (relatively) and improves the identification accuracy by 60 text-independent dataset. We also present results that suggest adapting from a model trained with Mandarin can improve accuracy for English speaker recognition.
We present Deep Voice 3, a fully-convolutional attention-based neural text-to-speech (TTS) system. Deep Voice 3 matches state-of-the-art neural speech synthesis systems in naturalness while training ten times faster. We scale Deep Voice 3 to data set sizes unprecedented for TTS, training on more than eight hundred hours of audio from over two thousand speakers. In addition, we identify common error modes of attention-based speech synthesis networks, demonstrate how to mitigate them, and compare several different waveform synthesis methods. We also describe how to scale inference to ten million queries per day on one single-GPU server.
Replacing hand-engineered pipelines with end-to-end deep learning systems has enabled strong results in applications like speech and object recognition. However, the causality and latency constraints of production systems put end-to-end speech models back into the underfitting regime and expose biases in the model that we show cannot be overcome by "scaling up", i.e., training bigger models on more data. In this work we systematically identify and address sources of bias, reducing error rates by up to 20% while remaining practical for deployment. We achieve this by utilizing improved neural architectures for streaming inference, solving optimization issues, and employing strategies that increase audio and label modelling versatility.
This paper proposes sampling techniques to approximate the configuration space for optimal motion planning. We sample valid configurations in the workspace and construct path subconvex cells in the free configuration space. The radius of each cell is calculated using lower bounds on the robot’s minimum time to collision. Using theorems about path convexity, the shortest paths found between any two points in the decomposed space are guaranteed to be safe. Experimental results are provided for a planar arm
This paper presents a definition of convexity useful for describing local optimality in configuration spaces, proves that finding convex regions is relatively easy, and presents an algorithm for approximating the free configuration space using a set of such convex regions. The paper examines simple but interesting systems: serial planar arms with revolute joints, and a Reeds-Shepp car. The paper experimentally explores an approach for finding good (although not necessarily optimal) trajectories using the derived data structure.
Acute osmotic fluctuations in the brain occur during a number of clinical conditions and can result in a variety of adverse neurological symptoms. Osmotic perturbation can cause changes in the volumes of intra‐ and extracellular fluid and, due to the rigidity of the skull, can alter intracranial pressure thus making it difficult to analyze purely osmotic effects in vivo. The present study aims to determine the effects of changes in osmolarity on SH‐SY5Y human neuroblastoma cells in vitro, and the role of the actin–myosin network in regulating this response. Cells were exposed to hyper‐ or hypoosmotic media and morphological and cytoskeletal responses were recorded. Hyperosmotic shock resulted in a drop in cell body volume and planar area, a persisting shape deformation, and increases in cellular translocation. Hypoosmotic shock did not significantly alter planar area, but caused a transient increase in cell body volume and an increase in cellular translocation via the development of small protrusions rich in actin. Disruption of the actin–myosin network with latrunculin and blebbistatin resulted in changes to volume and shape regulation, and a decrease in cellular translocation. In both osmotic perturbations, no apparent disruptions to cytoskeletal integrity were observed by light microscopy. Overall, because osmotically induced changes persisted even after volume regulation occurred, it is possible that osmotic stress may play a larger role in neurological dysfunction than currently believed. © 2015 Wiley Periodicals, Inc.
In this work, we employ quantitative methods from the realm of statistics and machine learning to develop novel methodologies for author attribution and textual analysis. In particular, we develop techniques and software suitable for applications to Classical study, and we illustrate the efficacy of our approach in several interesting open questions in the field. We apply our numerical analysis techniques to questions of authorship attribution in the case of the Greek tragedian Euripides, to instances of intertextuality and influence in the poetry of the Roman statesman Seneca the Younger, and to cases of "interpolated" text with respect to the histories of Livy.
In this work, we employ quantitative methods from the realm of statistics and machine learning to de- velop novel methodologies for author attribution and tex- tual analysis. In particular, we develop techniques and software suitable for applications to Classical study, and we illustrate the ecacy of our approach in several inter- esting open questions in the eld. We apply our numer- ical analysis techniques to questions of authorship attri- bution in the case of the Greek tragedian Euripides, to instances of intertextuality and inuence in the poetry of the Roman statesman Seneca the Younger, and to cases of \interpolated" text with respect to the histories of Livy 1 .