Human inspection of potential drug compounds is crucial in the virtual drug screening pipeline. However, there is a pressing need to accelerate this process, as the number of molecules humans can realistically examine is extremely limited relative to the scale of virtual screens. Furthermore, computational medicinal chemists can evaluate different poses inconsistently, and there is no standard way of recording annotations. We propose Autoparty, a containerized tool to address these challenges. Autoparty leverages on-premises active learning for drug discovery to facilitate human-in-the-loop training of models that extrapolate human intuition. We leverage multiple uncertainty quantification metrics to query the user with informative examples for model training, limiting the number of human expert training labels. The collected annotations populate a persistent and exportable local database for broad downstream uses. Incorporating Autoparty resulted in a 40% increase in hit rate over shape similarity alone among 193 experimentally tested compounds in a real-world case study.
At sufficiently high resolution, x-ray crystallography and cryogenic electron microscopy are capable of resolving small spherical map features corresponding to either water or ions. Correct classification of these sites provides crucial insight for understanding structure and function as well as guiding downstream design tasks, including structure-based drug discovery and de novo biomolecule design. However, direct identification of these sites from experimental data can prove challenging, and existing empirical approaches leveraging the local environment can only characterize limited ion types. We present a representation of chemical environments using interaction fingerprints and develop a machine learning model to predict the identity of input water and ion sites. We validate the method, named Metric Ion Classification (MIC), on a wide variety of biomolecular examples to demonstrate its utility, identifying many probable mismodeled ions deposited in the PDB. Compared to existing methods, MIC achieves superior accuracy for uniquely classifying water/ion sites while expanding the set of potential site identities. Finally, we collect all steps of this approach into an easy-to-use open-source package that can integrate with existing structure determination pipelines, and we provide a ChimeraX implementation to further enable use of the tool.
Alzheimer's disease (AD) is a multifactorial neurodegenerative disorder characterized by heterogeneous molecular changes across diverse cell types, posing significant challenges for treatment development. To address this, we introduced a cell-type-specific, multi-target drug discovery strategy grounded in human data and real-world evidence. This approach integrates single-cell transcriptomics, drug perturbation databases, and clinical records. Using this framework, letrozole and irinotecan were identified as a potential combination therapy, each targeting AD-related gene expression changes in neurons and glial cells, respectively. In an AD mouse model with both Aβ and tau deposits, this combination therapy significantly improved memory performance and reduced AD-related pathologies compared with vehicle and single-drug treatments. Single-nucleus transcriptomic analysis confirmed that the therapy reversed disease-associated gene networks in a cell-type-specific manner. These results highlight the promise of cell-type-directed combination therapies in addressing multifactorial diseases like AD and lay the groundwork for precision medicine tailored to patient-specific transcriptomic and clinical profiles.
Transformative neuropathology is redefining human brain research by integrating foundational descriptive pathology with advanced methodologies. These approaches, spanning multi-omics studies and machine learning applications, will drive discovery for the identification of biomarkers, therapeutic targets, and complex disease patterns through comprehensive analyses of postmortem human brain tissue. Yet critical challenges remain, including the sustainability of brain banks, expanding donor participation, strengthening training pipelines, enabling rapid autopsies, supporting collaborative platforms, and integrating data across modalities. Innovations in digital pathology, tissue quality enhancement, harmonization of data standards, and machine learning integration offer opportunities to accelerate tissue-level "pathomics" research in brain health through cross-disciplinary collaborations. Lessons from neuroimaging, particularly in establishing common data frameworks and multi-site collaborations, offer a valuable roadmap for streamlining innovations. In this perspective, we outline actionable solutions for leveraging existing resources and strengthening collaboration -where we envision future opportunities to drive translational discoveries stemming from transformative neuropathology.
Accurate prediction of enzymatic activity from amino acid sequences could drastically accelerate enzyme engineering for applications such as bioremediation and therapeutics development. In recent years, Protein Language Model (PLM) embeddings have been increasingly leveraged as the input into sequence-to-function models. Here, we use consistently collected catalytic turnover observations for 175 orthologs of the enzyme Adenylate Kinase (ADK) as a test case to assess the use of PLMs and their embeddings in enzyme kinetic prediction tasks. In this study, we show that nonlinear probing of PLM embeddings outperforms baseline embeddings (one-hot-encoding) and the specialized k_cat (catalytic turnover number) prediction models DLKcat and CatPred. We also compared fixed and learnable aggregation of PLM embeddings for k_cat prediction and found that transformer-based learnable aggregation of amino-acid PLM embeddings is generally the most performant. Additionally, we found that ESMC 600M embeddings marginally outperform other PLM embeddings for k_cat prediction. We explored Low-Rank Adaptation (LoRA) masked language model fine-tuning and direct fine-tuning for sequence-to-k_cat mapping, where we found no difference or a drop in performance compared to zero-shot embeddings, respectively. And we investigated the distinct hidden representations in PLM encoders and found that earlier layer embeddings perform comparable to or worse than the final layer. Overall, this study assesses the state of the field for leveraging PLMs for sequence-to-k_cat prediction on a set of diverse ADK orthologs.
Nature uses structural variations on protein folds to fine-tune the geometries of proteins for diverse functions, yet deep learning-based de novo protein design methods generate highly regular, idealized protein fold geometries that fail to capture natural diversity. Here, using physics-based design methods, we generated and experimentally validated a dataset of 5,996 stable, de novo designed proteins with diverse non-ideal geometries. We show that deep learning-based structure prediction methods applied to this set have a systematic bias towards idealized geometries. To address this problem, we present a fine-tuned version of Alphafold2 that is capable of recapitulating geometric diversity and generalizes to a new dataset of thousands of geometrically diverse de novo proteins from 5 fold families unseen in fine-tuning. Our results suggest that current deep learning-based structure prediction methods do not capture some of the physics that underlie the specific conformational preferences of proteins designed de novo and observed in nature. Ultimately, approaches such as ours and further informative datasets should lead to improved models that reflect more of the physical principles of atomic packing and hydrogen bonding interactions and enable improved generalization to more challenging design problems.
Researchers are developing increasingly robust molecular representations, motivating the need for thorough methods to stress-test and validate them. Here, we use a variational auto-encoder (VAE), an unsupervised deep learning model, to generate anomalous examples of SELF-referencIng Embedded Strings (SELFIES), a popular molecular string format. These anomalies defy the assertion that all SELFIES convert into valid SMILES strings. Interestingly, we find specific regions within the VAE's internal landscape (latent space), whose decoding frequently generates inconvertible SELFIES anomalies. The model's internal landscape self-organization helps with exploring factors affecting molecular representation reliability. We show how VAEs and similar anomaly generation methods can empirically stress-test molecular representation robustness. Additionally, we investigate reasons for the invalidity of some discovered SELFIES strings (version 2.1.1) and suggest changes to improve them, aiming to spark ongoing molecular representation improvement.
The investigation of chromatin organization in single cells holds great promise for identifying causal relationships between genome structure and function. However, analysis of single-molecule data is hampered by extreme yet inherent heterogeneity, making it challenging to determine the contributions of individual chromatin fibers to bulk trends. To address this challenge, we propose ChromaFactor, a novel computational approach based on non-negative matrix factorization that deconvolves single-molecule chromatin organization datasets into their most salient primary components. ChromaFactor provides the ability to identify trends accounting for the maximum variance in the dataset while simultaneously describing the contribution of individual molecules to each component. Applying our approach to two single-molecule imaging datasets across different genomic scales, we find that these primary components demonstrate significant correlation with key functional phenotypes, including active transcription, enhancer-promoter distance, and genomic compartment. Also, we find that some bulk trends exist at the single-cell level, but only in a small fraction of cells, suggesting that critical changes in genome organization may be driven by specific rare subpopulations rather than occurring uniformly across all cells. ChromaFactor offers a robust tool for understanding the complex interplay between chromatin structure and function on individual DNA molecules, pinpointing which subpopulations drive functional changes and fostering new insights into cellular heterogeneity and its implications for bulk genomic phenomena.
High-throughput phenotypic screening has historically relied on manually selected features, limiting our ability to capture complex cellular processes, particularly neuronal activity dynamics. While recent advances in self-supervised learning have revolutionized the study of cellular morphology and transcriptomics, dynamic cellular processes remain challenging to phenotypically profile. To address this, we developed Plexus, a self-supervised model designed to capture and quantify network-level neuronal activity. Unlike existing tools that focus on static readouts, Plexus leverages a network-level cell encoding method, efficiently encoding dynamic neuronal activity into rich representational embeddings. In turn, Plexus achieves state-of-the-art performance in detecting phenotypic changes in neuronal activity. Here we validated Plexus using a comprehensive GCaMP6m simulation framework and demonstrated its ability to classify distinct phenotypes compared with traditional signal-processing approaches. To enable practical application, we integrated Plexus with a scalable experimental system using human induced pluripotent stem cell-derived neurons expressing the GCaMP6m calcium indicator and CRISPR interference machinery. This platform successfully identified nearly 17 times as many phenotypic changes in response to genetic perturbations compared with conventional methods, as demonstrated in a 52-gene CRISPR interference screen across multiple induced pluripotent stem cell lines. Using this framework, we identified potential genetic modifiers of aberrant neuronal activity in frontotemporal dementia, illustrating its utility for understanding complex neurological disorders. Grosjean et al. present a network-aware, self-supervised learning approach for screening neuronal activity dynamics. They demonstrate its applicability across a range of neural interventions.
Alzheimer's disease (AD) is a multifactorial neurodegenerative disorder characterized by heterogeneous molecular changes across diverse cell types, posing significant challenges for treatment development. To address this, we introduced a cell-type-specific, multi-target drug discovery strategy grounded in human data and real-world evidence. This approach integrates single-cell transcriptomics, drug perturbation databases, and clinical records. Using this framework, letrozole and irinotecan were identified as a potential combination therapy, each targeting AD-related gene expression changes in neurons and glial cells, respectively. In an AD mouse model, this combination therapy significantly improved memory function and reduced AD-related pathologies compared to vehicle and single-drug treatments. Single-nuclei transcriptomic analysis confirmed that the therapy reversed disease-associated gene networks in a cell-type-specific manner. These results highlight the promise of cell-type-directed combination therapies in addressing multifactorial diseases like AD and lay the groundwork for precision medicine tailored to patient-specific transcriptomic and clinical profiles.
Accumulation of abnormal tau protein into neurofibrillary tangles (NFTs) is a pathologic hallmark of Alzheimer disease (AD). Accurate detection of NFTs in tissue samples can reveal relationships with clinical, demographic, and genetic features through deep phenotyping. However, expert manual analysis is time-consuming, subject to observer variability, and cannot handle the data amounts generated by modern imaging. We present a scalable, open-source, deep-learning approach to quantify NFT burden in digital whole slide images (WSIs) of post-mortem human brain tissue. To achieve this, we developed a method to generate detailed NFT boundaries directly from single-point-per-NFT annotations. We then trained a semantic segmentation model on 45 annotated 2400 μm by 1200 μm regions of interest (ROIs) selected from 15 unique temporal cortex WSIs of AD cases from three institutions (University of California (UC)-Davis, UC-San Diego, and Columbia University). Segmenting NFTs at the single-pixel level, the model achieved an area under the receiver operating characteristic of 0.832 and an F1 of 0.527 (196-fold over random) on a held-out test set of 664 NFTs from 20 ROIs (7 WSIs). We compared this to deep object detection, which achieved comparable but coarser-grained performance that was 60% faster. The segmentation and object detection models correlated well with expert semi-quantitative scores at the whole-slide level (Spearman’s rho ρ = 0.654 (p = 6.50e-5) and ρ = 0.513 (p = 3.18e-3), respectively). We openly release this multi-institution deep-learning pipeline to provide detailed NFT spatial distribution and morphology analysis capability at a scale otherwise infeasible by manual assessment.
Abstract Behavioral larval zebrafish screens leverage a high-throughput small molecule discovery format to find neuroactive molecules relevant to mammalian physiology. We screen a library of 650 central nervous system active compounds in high replicate to train deep metric learning models on zebrafish behavioral profiles. The machine learning initially exploited subtle artifacts in the phenotypic screen, necessitating a complete experimental re-run with rigorous physical well-wise randomization. These large matched phenotypic screening datasets (initial and well-randomized) provide a unique opportunity to quantify and understand shortcut learning in a full-scale, real-world drug discovery dataset. The final deep metric learning model substantially outperforms correlation distance–the canonical way of computing distances between profiles–and generalizes to an orthogonal dataset of diverse drug-like compounds. We validate predictions by prospective in vitro radio-ligand binding assays against human protein targets, achieving a hit rate of 58% despite crossing species and chemical scaffold boundaries. These neuroactive compounds exhibit diverse chemical scaffolds, demonstrating that zebrafish phenotypic screens combined with metric learning achieve robust scaffold hopping capabilities.
At sufficiently high resolution, x-ray crystallography and cryogenic electron microscopy are capable of resolving small spherical map features corresponding to either water or ions. Correct classification of these sites provides crucial insight for understanding structure and function as well as guiding downstream design tasks, including structure-based drug discovery and de novo biomolecule design. However, direct identification of these sites from experimental data can prove extremely challenging, and existing empirical approaches leveraging the local environment can only characterize limited ion types. We present a novel representation of chemical environments using interaction fingerprints and develop a machine-learning model to predict the identity of input water and ion sites. We validate the method, named Metric Ion Classification (MIC), on a wide variety of biomolecular examples to demonstrate its utility, identifying many probable mismodeled ions deposited in the PDB. Finally, we collect all steps of this approach into an easy-to-use open-source package that can integrate with existing structure determination pipelines.### Competing Interest StatementL.S. is a consultant for Deep Apple Therapeutics. G.S. is a cofounder of and consultant for Deep Apple Therapeutics. M.J.K. is a consultant for Deep Apple Therapeutics.
In this work, we introduce AutoFragDiff, a fragment-based autoregressive diffusion model for generating 3D molecular structures conditioned on target protein structures. We employ geometric vector perceptrons to predict atom types and spatial coordinates of new molecular fragments conditioned on molecular scaffolds and protein pockets. Our approach improves the local geometry of the resulting 3D molecules while maintaining high predicted binding affinity to protein targets. The model can also perform scaffold extension from user-provided starting molecular scaffold.
Melanocytic atypia, ranging from benign to malignant, often leads to diagnostic discordance, complicating its prediction by machine learning models. To overcome this, we paired H&E-stained histology images with contiguous or serial sections immunohistochemically (IHC) stained for melanocytic cells via antibodies for MelanA, MelPro, or SOX10. We developed a deep-learning pipeline to identify melanocytic atypia by digitizing a real-world archival dataset of 122 paired whole slide images from 61 confirmed melanoma in situ (MIS) cases at two institutions. Only 37.7% of the cases contained tissue pairs that matched well enough for deep learning. Nonetheless, the MelanA+MelPro models achieved an average area under the receiver-operating characteristic (AUROC) of 0.948 and an average area under the precision-recall curve (AUPRC) of 0.611, while the SOX10 models had an average of 0.867 AUROC and 0.433 AUPRC. Despite learning from biologically different IHC stains, the convolutional neural network (CNN) models independently exhibited an intuitive convergent rationale by explainable AI saliency calculations. Different antibodies, with nuclear versus cytoplasmic staining, provided complementary yet consistent information, which the CNNs integrated effectively. The resulting multi-antibody virtual stains identified morphologic cytologic and small-scale architectural features directly from H&E-stained histology images, which can assist pathologists in assessing cutaneous MIS. ### Competing Interest Statement E.S.K. is the founder of SFZ Health, LLC DBA Pathfolio. G.G. is now an employee of Genentech. M.T. is now a student at Mount Sinai.
Chemical probes interrogate disease mechanisms at the molecular level by linking genetic changes to observable traits. However, comprehensive chemical screens in diverse biological models are impractical. To address this challenge, we developed ChemProbe, a model that predicts cellular sensitivity to hundreds of molecular probes and drugs by learning to combine transcriptomes and chemical structures. Using ChemProbe, we inferred the chemical sensitivity of cancer cell lines and tumor samples and analyzed how the model makes predictions. We retrospectively evaluated drug response predictions for precision breast cancer treatment and prospectively validated chemical sensitivity predictions in new cellular models, including a genetically modified cell line. Our model interpretation analysis identified transcriptome features reflecting compound targets and protein network modules, identifying genes that drive ferroptosis. ChemProbe is an interpretable in silico screening tool that allows researchers to measure cellular response to diverse compounds, facilitating research into molecular mechanisms of chemical sensitivity.
Message-passing neural networks (MPNNs) on molecular graphs generate continuous and differentiable encodings of small molecules with state-of-the-art performance on protein-ligand complex scoring tasks. Here, we describe the Protein-Graph Network (PGN) package, an open-source toolkit that constructs ligand-receptor graphs based on atom proximity and allows users to rapidly apply and evaluate MPNN architectures for a broad range of tasks. We demonstrate the utility of PGN by introducing benchmarks for affinity and docking score prediction tasks. Graph networks generalize better than fingerprint-based models and perform strongly for the docking score prediction task. Overall, MPNNs with Proximity Graph data structures augment the prediction of ligand-receptor complex properties when ligand-receptor data are available.
Quantitatively mapping enzyme sequence-catalysis landscapes remains a critical challenge in understanding enzyme function, evolution, and design. Here, we expand an emerging microfluidic platform to measure catalytic constants-k cat and K M-for hundreds of diverse naturally occurring sequences and mutants of the model enzyme Adenylate Kinase (ADK). This enables us to dissect the sequence-catalysis landscape's topology, navigability, and mechanistic underpinnings, revealing distinct catalytic peaks organized by structural motifs. These results challenge long-standing hypotheses in enzyme adaptation, demonstrating that thermophilic enzymes are not slower than their mesophilic counterparts. Combining the rich representations of protein sequences provided by deep-learning models with our custom high-throughput kinetic data yields semi-supervised models that significantly outperform existing models at predicting catalytic parameters of naturally occurring ADK sequences. Our work demonstrates a promising strategy for dissecting sequence-catalysis landscapes across enzymatic evolution and building family-specific models capable of accurately predicting catalytic constants, opening new avenues for enzyme engineering and functional prediction.