We introduce the AI co-mathematician, a workbench for mathematicians to interactively leverage AI agents to pursue open-ended research. The AI co-mathematician is optimized to provide holistic support for the exploratory and iterative reality of mathematical workflows, including ideation, literature search, computational exploration, theorem proving and theory building. By providing an asynchronous, stateful workspace that manages uncertainty, refines user intent, tracks failed hypotheses, and outputs native mathematical artifacts, the system mirrors human collaborative workflows. In early tests, the AI co-mathematician helped researchers solve open problems, identify new research directions, and uncover overlooked literature references. Besides demonstrating a highly interactive paradigm for AI-assisted mathematical discovery, the AI co-mathematician also achieves state of the art results on hard problem-solving benchmarks, including scoring 48
Fault-tolerant quantum computing will require error rates far below those achievable with physical qubits. Quantum error correction (QEC) bridges this gap, but depends on decoders being simultaneously fast, accurate, and scalable. This combination of requirements has not yet been met by a machine-learning decoder, nor by any decoder for promising resource-efficient codes such as the colour code. Here we introduce AlphaQubit 2, a neural-network decoder that achieves near-optimal logical error rates for both surface and colour codes at large scales under realistic noise. For the colour code, it is orders of magnitude faster than other high-accuracy decoders. For the surface code, we demonstrate real-time decoding faster than 1 microsecond per cycle up to distance 11 on current commercial accelerators with better accuracy than leading real-time decoders. These results support the practical application of a wider class of promising QEC codes, and establish a credible path towards high-accuracy, real-time neural decoding at the scales required for fault-tolerant quantum computation.
The introduction of AlphaFold 21 has spurred a revolution in modelling the structure of proteins and their interactions, enabling a huge range of applications in protein modelling and design2-6. Here we describe our AlphaFold 3 model with a substantially updated diffusion-based architecture that is capable of predicting the joint structure of complexes including proteins, nucleic acids, small molecules, ions and modified residues. The new AlphaFold model demonstrates substantially improved accuracy over many previous specialized tools: far greater accuracy for protein-ligand interactions compared with state-of-the-art docking tools, much higher accuracy for protein-nucleic acid interactions compared with nucleic-acid-specific predictors and substantially higher antibody-antigen prediction accuracy compared with AlphaFold-Multimer v.2.37,8. Together, these results show that high-accuracy modelling across biomolecular space is possible within a single unified deep-learning framework.
The AlphaFold Protein Structure Database (AlphaFold DB, https://alphafold.ebi.ac.uk) is an openly accessible, extensive database of high-accuracy protein-structure predictions. Powered by AlphaFold v2.0 of DeepMind, it has enabled an unprecedented expansion of the structural coverage of the known protein-sequence space. AlphaFold DB provides programmatic access to and interactive visualization of predicted atomic coordinates, per-residue and pairwise model-confidence estimates and predicted aligned errors. The initial release of AlphaFold DB contains over 360,000 predicted structures across 21 model-organism proteomes, which will soon be expanded to cover most of the (over 100 million) representative sequences from the UniRef90 data set.
The chemotherapy resistance of esophageal adenocarcinomas (EACs) is underpinned by cancer cell extrinsic mechanisms of the tumor microenvironment (TME). We demonstrate that, by targeting the tumor-promoting functions of the predominant TME cell type, cancer-associated fibroblasts (CAFs) with phosphodiesterase type 5 inhibitors (PDE5i), we can enhance the efficacy of standard-of-care chemotherapy. In ex vivo conditions, PDE5i prevent the transdifferentiation of normal fibroblasts to CAF and abolish the tumor-promoting function of established EAC CAFs. Using shotgun proteomics and single-cell RNA-seq, we reveal PDE5i-specific regulation of pathways related to fibroblast activation and tumor promotion. Finally, we confirm the efficacy of PDE5i in combination with chemotherapy in close-to-patient and in vivo PDX-based model systems. These findings demonstrate that CAFs drive chemotherapy resistance in EACs and can be targeted by repurposing PDE5i, a safe and well-tolerated class of drug administered to millions of patients world-wide to treat erectile dysfunction.
We describe the operation and improvement of AlphaFold, the system that was entered by the team AlphaFold2 to the "human" category in the 14th Critical Assessment of Protein Structure Prediction (CASP14). The AlphaFold system entered in CASP14 is entirely different to the one entered in CASP13. It used a novel end-to-end deep neural network trained to produce protein structures from amino acid sequence, multiple sequence alignments, and homologous proteins. In the assessors' ranking by summed z scores (>2.0), AlphaFold scored 244.0 compared to 90.8 by the next best group. The predictions made by AlphaFold had a median domain GDT_TS of 92.4; this is the first time that this level of average accuracy has been achieved during CASP, especially on the more difficult Free Modeling targets, and represents a significant improvement in the state of the art in protein structure prediction. We reported how AlphaFold was run as a human team during CASP14 and improved such that it now achieves an equivalent level of performance without intervention, opening the door to highly accurate large-scale structure prediction.
These are whole-slide digital pathology images of esophageal adenocarcinoma (EAC) patient-derived xenograft models. EAC biopsy specimens were cultured in vitro before implanting into immuno-deficient mice. Mice were then divided into the following treatment groups and treated with combinations of chemotherapy and PDE5 inhibitors as follows: Untreated Epirubicin + Cisplatin + Capecitabine (ECX) ECX + Vardenafil ECX + Tadalafil All whole slide images are in Olympus .vsi format and can be opened using the BioFormats library in QuPath. We have also included classifiers and scripts for performing the digital pathology analysis in QuPath, and information to link each mouse with treatment groups and corresponding whole slide images. File metadata: classifiers.zip - Folder containing pixel classifiers used for segmentation of IHC stains. These classifiers are to be important into QuPath for the respective analysis of Periostin (POSTN) and alpha Smooth Muscle Actin (SMA) staining using the included Groovy scripts in the scripts folder. SMA.zip - Folder containing whole slide images of mouse patient-derived xenograft tumours in .vsi format (Olympus VS110). Immunohistochemistry with anti-alpha Smooth Muscle Actin antibody. POSTN.zip - Folder containing whole slide images of mouse patient-derived xenograft tumours in .vsi format (Olympus VS110). Immunohistochemistry with anti-Periostin antibody. scripts.zip - Folder containing all groovy scripts used in QuPath for tissue detection (Whole Section Tissue Detection.groovy), and segmentation of POSTN (POSTN Quantification Mice.groovy) and SMA staining tissue areas (SMA Quantification Mice.groovy). POSTN.xlsx - Data on the mice for which anti-POSTN IHC is available. Column names as follows: Image: file name (corresponding to .vsi file). Mouse_ID: Mouse identifier. Treatment: Treatment administered. ECX - Epirubicin, Cisplatin and Capecitabine. POSTN_Area: Total measured area of POSTN+ tissue in the tissue section (in um^2). Total_Area: Total area of the tissue section (in um^2). percentage_POSTN: Percentage of total tissue area that is stained with POSTN (%). SMA.xlsx - Data on the mice for which anti-SMA IHC is available. Column names as follows: Image: file name (corresponding to .vsi file). Mouse_ID: Mouse identifier. Treatment: Treatment administered. ECX - Epirubicin, Cisplatin and Capecitabine. SMA_Area: Total measured area of SMA+ tissue in the tissue section (in um^2). Total_Area: Total area of the tissue section (in um^2). percentage_SMA: Percentage of total tissue area that is stained with SMA (%).
Proteins are essential to life, and understanding their structure can facilitate a mechanistic understanding of their function. Through an enormous experimental effort 1–4 , the structures of around 100,000 unique proteins have been determined 5 , but this represents a small fraction of the billions of known protein sequences 6,7 . Structural coverage is bottlenecked by the months to years of painstaking effort required to determine a single protein structure. Accurate computational approaches are needed to address this gap and to enable large-scale structural bioinformatics. Predicting the three-dimensional structure that a protein will adopt based solely on its amino acid sequence—the structure prediction component of the ‘protein folding problem’ 8 —has been an important open research problem for more than 50 years 9 . Despite recent progress 10–14 , existing methods fall far short of atomic accuracy, especially when no homologous structure is available. Here we provide the first computational method that can regularly predict protein structures with atomic accuracy even in cases in which no similar structure is known. We validated an entirely redesigned version of our neural network-based model, AlphaFold, in the challenging 14th Critical Assessment of protein Structure Prediction (CASP14) 15 , demonstrating accuracy competitive with experimental structures in a majority of cases and greatly outperforming other methods. Underpinning the latest version of AlphaFold is a novel machine learning approach that incorporates physical and biological knowledge about protein structure, leveraging multi-sequence alignments, into the design of the deep learning algorithm.
While the vast majority of well-structured single protein chains can now be predicted to high accuracy due to the recent AlphaFold [1] model, the prediction of multi-chain protein complexes remains a challenge in many cases. In this work, we demonstrate that an AlphaFold model trained specifically for multimeric inputs of known stoichiometry, which we call AlphaFold-Multimer, significantly increases accuracy of predicted multimeric interfaces over input-adapted single-chain AlphaFold while maintaining high intra-chain accuracy. On a benchmark dataset of 17 heterodimer proteins without templates (introduced in [2]) we achieve at least medium accuracy (DockQ [3] ≥ 0.49) on 13 targets and high accuracy (DockQ ≥ 0.8) on 7 targets, compared to 9 targets of at least medium accuracy and 4 of high accuracy for the previous state of the art system (an AlphaFold-based system from [2]). We also predict structures for a large dataset of 4,446 recent protein complexes, from which we score all non-redundant interfaces with low template identity. For heteromeric interfaces we successfully predict the interface (DockQ ≥ 0.23) in 70% of cases, and produce high accuracy predictions (DockQ ≥ 0.8) in 26% of cases, an improvement of +27 and +14 percentage points over the flexible linker modification of AlphaFold [4] respectively. For homomeric inter-faces we successfully predict the interface in 72% of cases, and produce high accuracy predictions in 36% of cases, an improvement of +8 and +7 percentage points respectively.
Protein structures can provide invaluable information, both for reasoning about biological processes and for enabling interventions such as structure-based drug development or targeted mutagenesis. After decades of effort, 17% of the total residues in human protein sequences are covered by an experimentally determined structure 1 . Here we markedly expand the structural coverage of the proteome by applying the state-of-the-art machine learning method, AlphaFold 2 , at a scale that covers almost the entire human proteome (98.5% of human proteins). The resulting dataset covers 58% of residues with a confident prediction, of which a subset (36% of all residues) have very high confidence. We introduce several metrics developed by building on the AlphaFold model and use them to interpret the dataset, identifying strong multi-domain predictions as well as regions that are likely to be disordered. Finally, we provide some case studies to illustrate how high-quality predictions could be used to generate biological hypotheses. We are making our predictions freely available to the community and anticipate that routine large-scale and high-accuracy structure prediction will become an important tool that will allow new questions to be addressed from a structural perspective.
Deep reinforcement learning (RL) has led to many recent and groundbreaking advances. However, these advances have often come at the cost of both increased scale in the underlying architectures being trained as well as increased complexity of the RL algorithms used to train them. These increases have in turn made it more difficult for researchers to rapidly prototype new ideas or reproduce published RL algorithms. To address these concerns this work describes Acme, a framework for constructing novel RL algorithms that is specifically designed to enable agents that are built using simple, modular components that can be used at various scales of execution. While the primary goal of Acme is to provide a framework for algorithm development, a secondary goal is to provide simple reference implementations of important or state-of-the-art algorithms. These implementations serve both as a validation of our design decisions as well as an important contribution to reproducibility in RL research. In this work we describe the major design decisions made within Acme and give further details as to how its components can be used to implement various algorithms. Our experiments provide baselines for a number of common and state-of-the-art algorithms as well as showing how these algorithms can be scaled up for much larger and more complex environments. This highlights one of the primary advantages of Acme, namely that it can be used to implement large, distributed RL algorithms that can run at massive scales while still maintaining the inherent readability of that implementation. This work presents a second version of the paper which coincides with an increase in modularity, additional emphasis on offline, imitation and learning from demonstrations algorithms, as well as various new agents implemented as part of Acme.
Branched-chain α-keto acids (BCKAs) are downstream catabolites of branched-chain amino acids (BCAAs). Mitochondrial oxidation of BCKAs is catalyzed by branched-chain ketoacid dehydrogenase (BCKDH), an enzyme sensitive to inhibitory phosphorylation by BCKD kinase (BCKDK). Emerging studies show that defective BCAA catabolism and elevated BCKAs levels correlate with glucose intolerance and cardiac dysfunction. However, if/how BCKDH and BCKDK exert control on the availability and flux of intramyocellular BCKAs and if BCKA reprograms nutrient metabolism by influencing insulin action remains unexplored. We observed altered BCAA catabolizing enzyme expression in the murine heart and skeletal muscle during physiological fasting and diet-induced obesity and after ex vivo exposure of C2C12 cells to increasing concentration of saturated fatty acid, palmitate. BCKAs per se impaired insulin-induced AKT phosphorylation and AKT activity in skeletal myotubes and cardiomyocytes. In skeletal muscle cells, mTORC1 and protein translation signaling was enhanced by BCKA with concomitant suppression of mitochondrial respiration. Lowering intracellular BCKA levels by genetic and pharmacological activation of BCKDHA enhanced insulin signaling and activated pyruvate dehydrogenase, an effector of glucose oxidation and substrate metabolism. Our findings suggest that BCKAs profoundly influence muscle insulin function, providing new insight into the molecular nexus of BCAA metabolism and signaling with cellular insulin action and respiration.
New biological tools are required to understand the functional significance of genetic events revealed by whole genome sequencing (WGS) studies in oesophageal adenocarcinoma (OAC). The MFD-1 cell line was isolated from a 55-year-old male with OAC without recombinant-DNA transformation. Somatic genetic variations from MFD-1, tumour, normal oesophagus, and leucocytes were analysed with SNP6. WGS was performed in tumour and leucocytes. RNAseq was performed in MFD-1, and two classic OAC cell lines FLO1 and OE33. Transposase-accessible chromatin sequencing (ATAC-seq) was performed in MFD-1, OE33, and non-neoplastic HET1A cells. Functional studies were performed. MFD-1 had a high SNP genotype concordance with matched germline/tumour. Parental tumour and MFD-1 carried four somatically acquired mutations in three recurrent mutated genes in OAC: TP53 , ABCB1 and SEMA5A , not present in FLO-1 or OE33. MFD-1 displayed high expression of epithelial and glandular markers and a unique fingerprint of open chromatin. MFD-1 was tumorigenic in SCID mouse and proliferative and invasive in 3D cultures. The clinical utility of whole genome sequencing projects will be delivered using accurate model systems to develop molecular-phenotype therapeutics. We have described the first such system to arise from the oesophageal International Cancer Genome Consortium project.
Background: Defective branched-chain amino acid (BCAA) catabolism is central to the pathogenesis of obesity, insulin resistance (IR) and heart disease. Branched-chain α-keto acids (BCKAs), a catabolic product of BCAAs, is oxidized in the mitochondria by branched-chain ketoacid dehydrogenase (BCKDH), an enzyme sensitive to inhibitory phosphorylation by BCKD kinase (BCKDK). BCAAs activate mTORC1, which causes IR by inhibiting insulin signalling. However, the effect of BCKA on muscle insulin signalling is unexplored.
Background: Branched chain α-keto acids (BCKA) are intracellular catabolic product of branched-chain amino acids (BCAAs) catabolism. Mitochondrial oxidation of BCKA is catalyzed by branched chain ketoacid dehydrogenase (BCKDH), enzyme sensitive to inhibitory phosphorylation by branched chain ketoacid dehydrogenase kinase (BCKDK). Skeletal and cardiac muscle generate significant BCKA. Defective BCAA catabolism and elevated BCKA is central to the pathogenesis of obesity, insulin resistance and heart disease. However, the effect of BCKA on muscle insulin signalling is unexplored.
Background: Dysregulated branched amino acid (BCAA) metabolism is an important predictor of impaired insulin sensitivity and cardiovascular dysfunction. BCAA uptake is facilitated by branch chain aminotransferase (BCAT) to yield branched-chain α-keto acids (BCKA). Mitochondrial oxidation of BCKA is catalyzed by branched chain ketoacid dehydrogenase (BCKDH), enzyme sensitive to inhibitory phosphorylation by BCKD kinase. However, the clinical association between dysregulated BCAA catabolism and cardiovascular dysfunction merits investigation. Furthermore, the tissue-specific role of BCAA in inducing insulin resistance (IR) is unexamined.
Dieldrin is a legacy organochlorine pesticide that is persistent in the environment, despite being discontinued from use in North America since the 1970s. Some epidemiologic studies suggest that exposure to dieldrin is associated with increased risks of neurodegenerative disease and breast cancer by inducing inflammatory responses in tissues as well as oxidative stress. However, the direct effects of organochlorine pesticides on the heart have not been adequately addressed to date given that these chemicals are detectable in human serum and are environmentally persistent; thus, individuals may show latent adverse effects in the cardiovascular system due to long-term, low-dose exposure over time. Our objective was to determine whether low-level exposure to dieldrin at an environmentally relevant dose results in aberrant molecular signaling in the vertebrate heart. Using transcriptomic profiling and immunoblotting, we determined the global gene and targeted protein expression response to dieldrin treatment and show that dieldrin affects gene networks in the heart that are associated with processes related to cardiovascular disease, specifically cardiac arrest and ventricular fibrillation. We report that genes regulating inflammatory responses, a significant risk factor for cardiovascular disease, are upregulated by dieldrin whereas transcripts related to lysosomal function are significantly downregulated. To verify these findings, proteins in these pathways were examined with immunoblotting, and our results demonstrate that dieldrin constitutively activates Akt/mTOR signaling and downregulates lysosomal genes, participating in autophagy. Our data demonstrate that dieldrin induces genes associated with cardiovascular dysfunction and compromised lysosomal physiology, thereby identifying a novel mechanism for pesticide-induced cardiotoxicity.