
Large language models (LLMs) increasingly write code, analyze data, and orchestrate scientific workflows. This can create reproducibility challenges because LLMs blur the boundary between how an analysis is built and what it depends on at run time. Guidelines for LLM-assisted science begin with a critical choice: whether the LLM sits on the data path of the published analysis, a live step results depend on, or off the data path, producing durable artifacts such as code. Reproducible analyses require preserving data, code, and runtime; an on-path LLM becomes part of the runtime, a dependency that may change, be deprecated, or become inaccessible. We derive six recommendations: (1) keep the LLM off the data path where possible, (2) preserve LLM-generated artifacts, (3) verify results by methods suited to the LLM’s role, (4) consider open-weight models, (5) record the model version, and (6) assess determinism.
It is common to build predictive models for complex phenotypes, such as diseases, where sample sizes are often limited by the cost of data collection. Machine learning and statistical methods are widely used in these cases. However, they may fail to identify high-quality features and lead to suboptimal models with high underfitting or overfitting risks. Here, we present mBoost to identify low-quality features and suboptimal models to enhance phenotype prediction. Importantly, it provides confidence on model performance, helping researchers avoid the frustrating cycle of model switching for possibly better performance. We demonstrate mBoost with simulations and real-data examples. Specifically, mBoost identifies optimal depth and width for a deep neural network in a drug example, detects suboptimal models in microbiome cohorts, and proposes a measure for telling whether the trained model is applicable to new data. Thus, mBoost will serve as a practical tool for model building and diagnosis.
Using crosslinking heat-activated purification (CLAP), Guo et al. recently concluded that polycomb repressive complex 2 (PRC2) is not an RNA-binding protein (RBP), attributing prior reports to UV-crosslinked immunoprecipitation (CLIP)-related artifacts. Here, our re-analysis of raw datasets yields different results. Using an in-house computational pipeline, we detect substantial PRC2 enrichment throughout the transcriptome, including XIST. Implementation of the authors’ published pipeline also reaffirms PRC2 as an RBP. Closer examination of the authors’ computational workflow reveals several non-standard analytical practices. First, reads originated from non-target species and other unmappable sequences were retained to derive normalization factors. Second, PCR duplicate removal was applied selectively to the mappable read fraction but not to the unmappable reads. Third, a stringent enrichment threshold was imposed during visualization, resulting in near-zero PRC2 profiles and contradicting the positive PRC2 enrichment in the underlying intermediate outputs and their GEO table. These analytical choices reduced the signal-to-background ratio of the PRC2 interactome and influenced their interpretation of PRC2’s RNA-binding activity.
Retrieval-augmented reasoning pipelines are increasingly used to structure how large language models (LLMs) incorporate external evidence in clinical question answering, but their reliability under model variability remains unclear. We evaluated 34 LLMs on 169 expert-curated, text-only, multiple-choice radiology questions, comparing zero-shot inference with a fixed shared-evidence condition in which all models received identical structured evidence reports before answer selection. Primary endpoints were inter-model decision dispersion and panel-based robustness of correctness. Shared evidence reduced decision dispersion (median entropy 0.48→0.13; p = 5.6 × 10−9) and increased robustness of correctness (mean 0.74→0.81; p = 5.6 × 10−9). Majority consensus increased, and consensus strength remained strongly correlated with robustness. However, high agreement did not guarantee correctness, and rare high-consensus failures persisted. These findings show that shared evidence can improve reproducibility under model variability, but accuracy and agreement alone are insufficient to characterize reliability.
Multimodal learning enables the integration of heterogeneous data, but value emerges only when data are transformed into stable, reusable evidence. Such evidence can reduce interpretive cost, support trust across contexts, and accumulate into data capital.
Embeddings, numerical vectors learned by deep-learning models, are increasingly used to represent complex molecular biology data and support predictive tasks and generative design. There is a growing need for systematic approaches to interpret and explain the information encoded in high-dimensional embedding spaces. Here, we introduce EmmaEmb, a quantitative, model-agnostic framework for geometric correction, direct analysis, and comparison of embedding spaces. Our framework encompasses local and global analysis methods to quantify data distribution within an embedding space and enable comparisons of representations across spaces in relation to known biological features. Through experiments with seven embedding models across six molecular biology tasks, we demonstrate that our methods reveal insights from embedding spaces that align with downstream predictive tasks, uncover misclassification patterns, and contextualize differences in biological information captured by ProtT5, AlphaFold2, and ESM C. We provide an open-source Python library implementing all analysis methods and a guided diagnostic workflow.
Artificial intelligence (AI) has greatly expanded the generative capacity of drug discovery, yet the ability to translate large candidate pools into structured, resource-aware decisions remains limited. This review addresses this emerging optimization bottleneck by examining quantum annealing as a potential decision-optimization layer within hybrid computational workflows. We focus on discrete, constraint-dominated tasks—such as compound subset selection, combinatorial design, and multi-objective prioritization—that can be formulated as quadratic unconstrained binary optimization (QUBO) problems. Rather than positioning quantum annealing as a predictive tool, we analyze its role as a complementary optimization interface integrated with AI-generated scores. We further discuss practical implementation considerations, including embedding overhead, noise, scalability, and benchmarking challenges. By emphasizing workflow-level design and rigorous evaluation, this review provides a pragmatic framework for assessing annealing-based optimization in drug discovery and related data-intensive scientific domains.
Single-cell sequencing enables detailed study of cell-state transitions, but extracting smooth, low-dimensional structures from noisy, high-dimensional data remains challenging. Neighbor embedding (NE) algorithms, such as t-distributed stochastic neighbor embedding (t-SNE) and uniform manifold approximation and projection (UMAP), are widely used to embed high-dimensional single-cell data into low dimensions, but they often introduce distortions that can lead to misleading interpretations. To address these challenges, we build on the predictability-computability-stability (PCS) framework for reliable and reproducible data-driven discoveries. First, we systematically evaluate popular NE algorithms through empirical and theoretical analyses, revealing their key limitations such as algorithmic artifacts and instability. We then introduce NESS, a principled and interpretable machine learning approach that improves NE representations by leveraging algorithmic stability, enabling more robust inference of smooth biological structures from single-cell data. Finally, we apply NESS to multiple datasets, including studies of pluripotent stem cell differentiation, organoid development, and diverse tissue-specific lineages. Across these settings, NESS consistently yields biologically meaningful insights.
Material-composition information extracted from spectral computed tomography (CT) images facilitates clinical diagnosis by characterizing pathological tissues. However, the complexity of human anatomy and the diversity of tissue types make precise material quantification from single-energy CT (SECT) highly challenging. Here, we propose an X-ray imaging induced multi-material decomposition (MMDX) framework. By integrating the physical principles of photoelectric effect and Compton scattering into a deep-learning network architecture, MMDX achieves superior performance on both phantom and clinical patient datasets, with a 12.92-dB improvement in peak signal-to-noise ratio (PSNR) and a 3.94% increase in volume fraction accuracy. To evaluate the clinical utility of MMDX, we constructed a large-scale clinical CT dataset with 1,637,738 CT images from 7,629 patients. MMDX demonstrates superior performance across 20 downstream tasks, including three diagnosis tasks, two prognosis tasks, and 15 biomarker prediction tasks, showing promise as an accessible AI tool for precision clinical diagnosis.
Biological systems comprise carefully coordinated processes. Challenges, e.g., pregnancy or surgery, can introduce major disruptions. Attempting to understand the corresponding complex adaptations, emerging technologies measure a multitude of biomarkers. However, current analytical tools often focus on changes in individual biomarkers and do not explicitly analyze the underlying, dynamically changing network of biomarker interactions. NeDis is an easy-to-use, highly customizable, open-source package that enables quantification of the functional disruption of biomarker networks across conditions and time. This allows the discovery of interconnected and coordinated subgroups of biomarkers with characteristic disruption profiles (e.g., increasing disruption over time) and provides a dynamic perspective on the coordination and adaptation of biological systems. Synthetic experiments demonstrate that NeDis captures more intricate signals than dimensionality reduction (e.g., principal-component analysis [PCA] or t-distributed stochastic neighbor embedding [t-SNE]). On high-dimensional single-cell mass cytometry (CyTOF) data, NeDis reveals coordinated functional disruptions of the immune system during human pregnancy, e.g., pointing to increased susceptibility to infection.
Partial information decomposition (PID) has emerged as a principled way to decompose the information carried by neural activity into components identifying whether interactions among neurons or brain areas generate synergistic or redundant information. Here, we demonstrate that empirical measures of synergy and redundancy based on either Gaussian or discrete probability estimators suffer from a substantial limited-sampling estimation bias. This bias is much larger for synergy than for redundancy. The gap between them increases with the number of parameters specifying the probability distributions. We develop procedures that effectively correct for the bias and provide rules of thumb for the sample sizes required to obtain unbiased estimates. We show that, when used on empirical brain datasets, they successfully remove large synergy biases across species, recording modalities, and experimental designs. Our bias corrections extend the range of neuroscience questions and experimental designs addressable with PID and allow accurate comparisons between synergy and redundancy.
The suprachiasmatic nucleus (SCN) is the master circadian clock in mammals, comprising ∼20,000 neurons organized into a bilaterally symmetric oval structure. System-level time computations in the SCN depend on coordinated spatiotemporal patterns of neuronal activity, yet most prevailing methods perform time-series analyses while disregarding neural spatiotemporal organization. Here, we developed an interpretable machine learning framework for the integrative analysis of large-scale spatiotemporal calcium signals from SCN neurons, with built-in validation and biological interpretability. Applying this framework, we identified distinct neural spatiotemporal states and subtypes whose spatial mapping across hemispheres revealed hemispheric asymmetry, particularly during the subjective day. Circadian timekeeping ability also displayed side specificity, with each hemisphere encoding a full yet unique time feature representation. Attribution analysis further indicated spectral features of calcium signals as the primary discriminative elements underlying this asymmetry. Overall, we demonstrate that hemispheric asymmetry of state switching is a fundamental property of the brain’s circadian clock.
Trained on massive datasets, foundation models produce embeddings used for many applications. We created a gene set foundation model (GSFM) trained on a massive collection of unlabeled gene sets from Rummagene and RummaGEO. Rummagene extracts gene sets from supplemental materials of publications, and RummaGEO hosts gene sets computed from published transcriptomics studies. Several GSFM architectures were benchmarked for their ability to predict gene function, gene-disease associations, and protein-protein interactions as well as to perform gene set enrichment analysis. Gene function predictions were compared with other models and evaluated using labeled gene sets from the Gene Ontology and KEGG pathways, the GWAS Catalog, and ChEA. The best GSFM architecture is a denoising autoencoder trained on multi-hot-encoded gene sets. This GSFM model achieves better performance compared with the other models. Gene-focused landing pages were created to serve GSFM gene function predictions for all human genes. These landing pages are served on a dedicated platform that also provide GSFM gene set augmentation and GSFM gene set enrichment analysis.
The Hes family, basic-helix-loop-helix transcription factors and downstream effectors of Notch signaling, regulate the fate choices of pancreatic progenitors, muscle stem cells, neuronal progenitors, and presomitic mesoderm cells. Bioluminescence imaging (BLI) has revealed ultradian oscillatory dynamics of Hes-family members Hes1, Hes5, and Hes7. However, identifying which of the Hes target genes also oscillate remains challenging due to the time-consuming and costly nature of tracking individual target genes using BLI. Here, we propose OscillomeR, a computational framework that reconstructs ultradian oscillations from RNA-sequencing data to identify oscillatory target genes at high throughput. OscillomeR predicts thousands of oscillatory genes in synchronized or unsynchronized cell types, identifying both known and novel Hes-family targets. It also captures the dynamic rewiring of gene-regulatory networks during cell differentiation. Overall, OscillomeR is an effective tool for elucidating the functions of oscillatory transcription factors at the genomic scale.
Complex medical reasoning requires integrating heterogeneous clinical evidence across multiple inference steps. Large language models (LLMs) now approach this through two routes: internalized reasoning and externalized agent scaffolding (frameworks that decompose problems collaboratively among multiple LLMs). To determine whether these routes are exclusive or complementary, we introduce MedicalAgentsBench, a filtered benchmark of 862 complex clinical questions drawn from the union of eight medical datasets via difficulty-aware curation and contamination screening. Evaluating three internalized reasoning models (DeepSeek-R1, o1-mini, and o3-mini), seven base models, and nine externalized agent-based methods, we find that internalized and externalized approaches each independently improve performance and that their benefits compound: the highest accuracy is achieved by layering agent workflows onto an internalized reasoning model (i.e., o3-mini + MDAgents with 35.1%). Pareto analysis shows this combination dominates the cost-performance frontier; moreover, lightweight optimization on inexpensive models offers an entry point for resource-constrained settings.
Recent advances in general-purpose AI provide new insights into how the neocortex and cerebellum, despite their uniform circuit architectures, support diverse functions and human intelligence. Beyond the traditional focus on visual processing, this paper offers a cross-domain comparison of the brain and AI through the lens of world-model-based computation. We argue that both the neocortex and cerebellum predict future world states from past inputs and construct predictive world models through prediction-error learning. These predictive world models are repurposed for sensory comprehension and motor output generation, thereby supporting multi-domain capabilities. Underlying these capabilities and even human-like adaptive intelligence, the neocortex implements hierarchical attention-based processing. Interestingly, autoregressive generative models in transformer-based AI have independently converged on similar computational principles. Together, these shared mechanisms suggest a common computational foundation through which uniform circuits can support cross-domain high-level intelligence in both biological and artificial systems.
Tracking neurons across days in high-density extracellular recordings is essential for investigating the mechanisms of learning and representational drift. However, in weeks-long recordings, identifying matches across sessions is hindered by changes in spike waveforms and unit turnover. We introduce DANT (density-based across-day neuron tracking), a framework that iterates between density-based clustering in feature space and probe-motion estimation inferred from provisional matches. The estimated motion is then used to reregister spike waveforms across sessions before clustering is recomputed in the next iteration. Within this loop, DANT learns a decision boundary from match and non-match labels and uses it in post hoc curation. Applied to weeks-long Neuropixels recordings from cortex and striatum in freely moving rats during reaction-time and self-timing tasks, DANT substantially increases match yield while maintaining a low false-positive rate relative to existing approaches. These results establish DANT as a general unsupervised solution for longitudinal tracking in chronic recordings.
Link prediction is important across biological, social, and technological networks, but many methods either require domain-specific node attributes, do not scale to large graphs, or miss higher-order topology. We present BLANT-Predict, a topology-only framework that uses sampled graphlets and orbit-pair frequencies to rank likely missing edges. Across 12 real-world networks (up to about 1 million nodes and 3 million edges), we compare against 13 baseline methods and observe higher precision with strong scalability. Beyond standard k-fold cross-validation, we evaluate predictions against future out-of-sample network snapshots to better reflect real deployment. BLANT-Predict maintains superior precision in these future tests, indicating practical value for predicting previously unobserved links.
Yaacov, Levi, and Peled argue that AI-powered clinical trial matching has achieved near-expert technical accuracy yet fails to increase patient enrollment because the dominant bottleneck is systemic—involving logistics, workflow misalignment, and consent burden—not informational. They propose reframing AI’s role from matching engine to trial facilitation system.