Many of the most important problems in science and engineering are inverse problems: given a desired outcome, find a design that achieves it. Evaluating whether a candidate meets the spec is often routine; a binding energy can be computed, a reactor yield simulated, a pharmacokinetic profile predicted. But searching a combinatorial design space for inputs that satisfy those targets is fundamentally harder. We introduce SciDesignBench, a benchmark of 520 simulator-grounded tasks across 14 scientific domains and five settings spanning single-shot design, short-horizon feedback, long-horizon refinement, and seed-design optimization. On the 10-domain shared-core subset, the best zero-shot model reaches only 29.0
The finite symmetric group S_n provides a natural domain for permutations, yet learning probability distributions on S_n is challenging due to its factorially growing size and discrete, non-Euclidean structure. Recent permutation diffusion methods define forward noising via shuffle-based random walks (e.g., riffle shuffles) and learn reverse transitions with Plackett-Luce (PL) variants, but the resulting trajectories can be abrupt and increasingly hard to denoise as n grows. We propose Soft-Rank Diffusion, a discrete diffusion framework that replaces shuffle-based corruption with a structured soft-rank forward process: we lift permutations to a continuous latent representation of order by relaxing discrete ranks into soft ranks, yielding smoother and more tractable trajectories. For the reverse process, we introduce contextualized generalized Plackett-Luce (cGPL) denoisers that generalize prior PL-style parameterizations and improve expressivity for sequential decision structures. Experiments on sorting and combinatorial optimization benchmarks show that Soft-Rank Diffusion consistently outperforms prior diffusion baselines, with particularly strong gains in long-sequence and intrinsically sequential settings.
Non-monotonic sequence generation methods, such as masked diffusion models, provide a flexible alternative to left-to-right autoregressive modeling by allowing tokens to be generated in non-fixed and prescribed orders. Despite their practical advantages, most existing non-monotonic models are order-agnostic and rely on a fixed-length masked token grid, limiting their ability to support variable-length generation and adaptive insertion order. In this work, we introduce a probabilistic framework for learning insertion order in variable-length insertion models. We formalize a bijective correspondence between insertion trajectories and permutations, which enables an exact reparameterization of the data likelihood as a sum over permutations. Building on this result, we propose the , a stochastic generative model that jointly learns to insert, to insert, and to terminate, trained via permutation-based variational inference. Unlike prior masked or fixed-canvas approaches, IP natively supports variable-length generation and learns data-driven preferences over insertion orders. Experiments on planning benchmarks and molecular SMILES generation demonstrate that learning insertion order improves both modeling quality and generalization in domains without a canonical left-to-right structure.
Discrete biological sequence optimization demands iterative refinement while satisfying strict syntactic constraints. Diffusion-based approaches provide strong progressive refinement but are not naturally aligned with discrete, grammar-constrained edit operations, whereas autoregressive LLMs readily produce valid sequences yet often lack explicit long-horizon planning. To close this gap, we introduce (Sequence Trajectory Refinement via Internalized Denoising Emulation), a post-training framework that recasts optimization as an intrinsic reasoning problem in edit space. Rather than relying on external agentic search loops, trains an LLM to emit a full trajectory of atomic edits as explicit Chain-of-Thought, effectively internalizing a trajectory-based refinement policy under discrete constraints. We instantiate with a curriculum that combines supervised fine-tuning on Levenshtein-aligned shortest-edit demonstrations with GRPO-style reinforcement learning (and variants) to align edit trajectories with task rewards. Across protein and molecule optimization benchmarks, consistently outperforms a diverse set of baselines, while producing candidates that maintain high structural validity and achieve improved target properties.
DNA methylation research has vastly expanded over the past decade, producing a wealth of epigenome-wide association studies, biomarker algorithms such as epigenetic clocks, technical performance analyses, and functional annotations for CpG sites. However, these resources remain fragmented across dozens of databases and supplementary files within manuscripts, forcing researchers to spend time and effort on data cleaning and integration prior to meaningful analyses. No single resource currently unifies this information into a centralized, easy-to-query framework. Here, we present CpG Atlas, a curated relational database that integrates 18 distinct annotation layers encompassing over 1.2 million CpG sites across all four generations of Illumina methylation arrays (HM450K, EPIC v1, EPIC v2, and MSA). Built on a snowflake schema with a canonical probe identifier hub implemented in SQL, CpG Atlas consolidates over 800,000 CpG-trait associations, results from Mendelian randomization analyses, CpG membership across 81 epigenetic clocks, array manifest information, and probe reliability data. It further includes specialized layers such as solo-WCGW, CoRSIVs, PRC2 binding, transposon and retroelement annotations, tissue-specific differentially methylated positions across 17 tissues, and hallmarks of aging and cancer. To maximize utility and ease of use, the database is paired with an interactive web tool and a natural language-to-SQL query interface, enabling users to quickly perform complex multi-dimensional queries. Detailed documentation about every data source and table is also provided, facilitating the identification and interpretation of relevant studies. We demonstrate the utility of CpG Atlas through two case studies: a systematic enrichment analysis revealing distinct functional signatures across 16 epigenetic clocks, and an iterative biomarker discovery workflow for IBD that leverages cross-layer integration. Because it is readily scalable simply by adding or updating tables in the database, CpG Atlas provides a continuously evolving and extensible infrastructure for the epigenetics community that supports collaborative research, interpretable biomarker development, and integrative analyses across the growing landscape of epigenetic data.
Single-cell RNA sequencing has transformed our understanding of cellular diversity, yet current single-cell foundation models (scFMs) remain limited in their scalability, flexibility across diverse tasks, and ability to natively integrate textual information. In this work, we build upon the Cell2Sentence (C2S) framework, which represents scRNA-seq profiles as textual "cell sentences," to train Large Language Models (LLMs) on a corpus comprising over one billion tokens of transcriptomic data, biological text, and metadata. Scaling the model to 27 billion parameters yields consistent improvements in predictive and generative capabilities and supports advanced downstream tasks that require synthesis of information across multi-cellular contexts. Targeted fine-tuning with modern reinforcement learning techniques produces strong performance in perturbation response prediction, natural language interpretation, and complex biological reasoning. This predictive strength enabled a dual-context virtual screen that nominated the kinase inhibitor silmitasertib (CX-4945) as a candidate for context-selective upregulation of antigen presentation. Experimental assessment in human cell models unseen during training supported this prediction, demonstrating that C2S-Scale can effectively guide the discovery of context-conditioned biology. C2S-Scale unifies transcriptomic and textual data at unprecedented scales, surpassing both specialized single-cell models and general-purpose LLMs to provide a platform for next-generation single-cell analysis and the development of "virtual cells."
Operator learning for time-dependent partial differential equations (PDEs) has seen rapid progress in recent years, enabling efficient approximation of complex spatiotemporal dynamics. However, most existing methods rely on fixed time step sizes during rollout, which limits their ability to adapt to varying temporal complexity and often leads to error accumulation. Here, we propose the Time-Adaptive Transformer with Neural Taylor Expansion (TANTE), a novel operator-learning framework that produces continuous-time predictions with adaptive step sizes. TANTE predicts future states by performing a Taylor expansion at the current state, where neural networks learn both the higher-order temporal derivatives and the local radius of convergence. This allows the model to dynamically adjust its rollout based on the local behavior of the solution, thereby reducing cumulative error and improving computational efficiency. We demonstrate the effectiveness of TANTE across a wide range of PDE benchmarks, achieving superior accuracy and adaptability compared to fixed-step baselines, delivering accuracy gains of 60-80% and speed-ups of 30-40% at inference time. The code is publicly available at https://github.com/zwu88/TANTE for transparency and reproducibility.
Background: Patients with advanced-stage pancreatic ductal adenocarcinoma (PDAC) are regularly treated with FOLFIRINOX, a chemotherapy regimen based on 5-fluorouracil, irinotecan and oxaliplatin, which is associated with high toxicity. Dosing of FOLFIRINOX is based on body surface area, risking under- or overdosing caused by altered pharmacokinetics due to interindividual differences in body composition. This study aimed to investigate the relationship between body composition and treatment toxicity in advanced stage PDAC patients treated with FOLFIRINOX. Methods: Data from patients treated at the Maastricht University Medical Centre + between 2012 and 2020 were collected retrospectively (n = 65). Skeletal muscle-, visceral adipose tissue, subcutaneous adipose tissue-, (SM-Index, VAT-Index, SAT-Index resp.) and Skeletal Muscle Radiation Attenuation (SMRA) were calculated after segmentation of computed tomography (CT) images at the third lumbar level using a validated deep learning method. Lean body mass (LBM) was estimated using SM-Index. Toxicities were scored and grade 3-4 adverse events were considered dose-limiting toxicities (DLTs). Results: Sixty-seven DLTs were reported during the median follow-up of 51.4 (95%CI 39.2-63.7) weeks. Patients who experienced at least one DLT had significantly higher dose intensity per LBM for all separate cytotoxics of FOLFIRINOX. Independent prognostic factors for the number of DLTs per cycle were: sarcopenia ((3 = 0.292; 95%CI 0.013 to 0.065; p = 0.013), SM-Index change (% per 30 days, (3 = -0.045; 95%CI -0.079 to -0.011; p = 0.011), VAT-Index change (% per 30 days, (3 = -0.006; 95%CI -0.012 to 0.000; p = 0.040) between diagnosis and the first follow-up CT scan, and cumulative relative dose intensity >80 % ((3 = -0.315; 95 % CI -0.543 to -0.087; p = 0.008). Conclusion: Sarcopenia and early muscle and fat wasting during FOLFIRINOX treatment were associated with treatment-related toxicity, warranting exploration of body composition guided personalized dosing of chemotherapeutics to limit DLTs. (c) 2024 The Author(s). Published by Elsevier Ltd on behalf of European Society for Clinical Nutrition and Metabolism. This is an open access article under the CC BY license (http://creativecommons.org/licenses/ by/4.0/).
Discrete diffusion models offer a flexible, controllable approach to structured sequence generation, yet they still lag behind causal language models in expressive power. A key limitation lies in their reliance on the Markovian assumption, which restricts each step to condition only on the current state, leading to potential uncorrectable error accumulation. In this paper, We introduce CaDDi, a discrete diffusion model that conditions on the entire generative trajectory, thereby lifting the Markov constraint and allowing the model to revisit and improve past states. By unifying sequential (causal) and temporal (diffusion) reasoning in a single non‑Markovian transformer, CaDDi also treats standard causal language models as a special case and permits the direct reuse of pretrained LLM weights with no architectural changes. Empirically, CaDDi outperforms state‑of‑the‑art discrete diffusion baselines on natural‑language benchmarks, substantially narrowing the remaining gap to large autoregressive transformers.
This files contains all supplementary data supporting the development and validation of AAnet
Despite promise in preclinical models, most immuno-oncology drug candidates fail in clinical trials. These failures reflect limitations in our ability to directly model the response of human tumor and immune cells to immunotherapies. To address this gap and test the effect of innate immune agonists, we developed PERCEPT, an approach that uses ex vivo perturbational single-cell RNA sequencing to compare the response of immunomodulatory treatments with unstimulated controls directly in patient samples. Using PERCEPT, we tested cytokines and innate immune agonists in melanoma and Merkel cell carcinoma (MCC) and identified the dsRNA mimetic, RIG-I agonist, Stem Loop RNA (SLR) 14 as a powerful inducer of anti-viral states and enhancer of T cell activation. We compared transcriptional responder and non-responder patient samples and identified midkine (MDK), a multifunctional cytokine, as a potent repressor of IFN signaling in both tumor and immune cells. MDK expression dampened MHC-I presentation in human tumor cells and reduced activation of antigen-presenting cells, disrupting tumor immunity at multiple levels. In contrast to prior studies, we identified MDK as specifically enriched in neuroendocrine cancers such as MCC and small cell lung cancer compared with melanoma, suggesting the importance of lineage- and context-specific targeting. Our results demonstrate the utility of high-dimensional controlled perturbation of patient samples to identify mechanisms of innate immune response and resistance and demonstrate an actionable path towards clinical development of MDK-inhibiting therapies including FDA-approved ALK inhibitors in neuroendocrine cancers. ### Competing Interest Statement A.I. co-founded RIGImmune, Xanadu Bio and PanV, and is a member of the Board of Directors of Roche Holding and Genentech. A.I. is a co-inventor on patents related to Stem-Loop RNA (SLR) compounds. NIH Common Fund, https://ror.org/001d55x84, 1R37CA279834, 5P50CA121974, 3P50DE030707, T32CA233414, 5K12CA215110, T32AI07019 NIH Common Fund, https://ror.org/001d55x84, T32GM136651 AstraZeneca (Switzerland), https://ror.org/034rhks82 Conquer Cancer Foundation, https://ror.org/027327e78, Endowed Young Investigator Award, in Memory of John R. Durant, MD
Background: Indications for cardiac magnetic resonance imaging (CMR) are often stored in heterogenous, unstructured reports. Manual adjudication of indications is time-consuming and requires domain expertise. Recent large language models (LLMs) have shown promise in complex clinical interpretation and categorization tasks. No prior study has systematically evaluated the ability of state-of-the-art (SOTA) LLMs to extract indications from raw CMR reports. Research question: How well do SOTA open-source and commercial LLMs adjudicate clinical indications from real-world CMR reports? Methods: We analyzed 486 CMR reports from a large academic center. Reports were de-identified using the Stanford-Penn-MIDRC deidentification tool, and ground-truth indications were annotated by a physician expert. 18 LLMs varying in accessibility (8 open-source, 10 commercial), parameter size (4 to 70 billion), and training corpus (general vs medical) were evaluated. For each report, LLMs were instructed to extract the top two possible indications (correct if either matched the ground-truth indication)—reflecting the fact that real-world indications can fall into more than 1 category—from ten possible categories: oncologic therapy toxicity, cardiomyopathy/elevated troponin, chest pain/dyspnea, arrythmia/abnormal ECG, cardiac mass/metastasis, thrombus, structural evaluation, pericarditis, risk stratification, or viability evaluation (ischemic). Results: Higher-cost commercial models (Spearman’s rank r = 0.683, p = 0.03) and larger-parameter open-source models ( r = 0.307) exhibited better adjudication ability, Fig 1A, 1B . The best performing commercial LLMs performed markedly better than the top open-source LLMs (90% vs ~78% accuracy [acc]), Fig 2 . Grok 3 (91% acc, 0.94 F1-score) and OpenAI o3 (90% acc, 0.93 F1) were the best models overall, and Gemma 3 27B was the best open-source LLM (80% acc, 0.86 F1), Fig 2 . Reasoning models performed comparably to non-reasoning models, with Grok 3 mini having the best relative cost-vs-performance, Fig 1A, 2 . Interestingly, medical LLMs performed worse than their generally pretrained counterparts (e.g., MedGemma 27B vs Gemma 3 27B), suggesting domain-specific pretraining may negatively affect adjudication ability, Fig 2 . Conclusion: Open-source and commercial LLMs demonstrate promise in automated, accurate extraction of indications from CMR reports. Our findings help clinician-researchers decide between LLMs for use-cases involving CMR reports.
We explore the emergence of intelligent behavior in artificial systems by investigating how the complexity of rule-based systems influences the capabilities of models trained to predict these rules. Our study focuses on elementary cellular automata (ECA), simple yet powerful one-dimensional systems that generate behaviors ranging from trivial to highly complex. By training distinct Large Language Models (LLMs) on different ECAs, we evaluated the relationship between the complexity of the rules' behavior and the intelligence exhibited by the LLMs, as reflected in their performance on downstream tasks. Our findings reveal that rules with higher complexity lead to models exhibiting greater intelligence, as demonstrated by their performance on reasoning and chess move prediction tasks. Both uniform and periodic systems, and often also highly chaotic systems, resulted in poorer downstream performance, highlighting a sweet spot of complexity conducive to intelligence. We conjecture that intelligence arises from the ability to predict complexity and that creating intelligence may require only exposure to complexity.
Identifying functionally important cell states and structure within heterogeneous tumors remains a significant biological and computational challenge. Current clustering- or trajectory-based models are ill-equipped to address the notion that cancer cells reside along a phenotypic continuum. We present Archetypal Analysis network (AAnet), a neural network that learns archetypal states within a phenotypic continuum in single-cell data. Unlike traditional archetypal analysis, AAnet learns archetypes (AT) in a simplex-shaped neural network latent space. Using preclinical and clinical models of breast cancer, AAnet resolves distinct cell states and processes, including cell proliferation, hypoxia, metabolism, and immune interactions. Primary tumor ATs are recapitulated in matched liver, lung, and lymph node metastases. Spatial transcriptomics reveals archetypal organization within the tumor and intra-archetypal mirroring between cancer and adjacent stromal cells. AAnet identifies GLUT3 within the hypoxic AT that proves critical for tumor growth and metastasis. AAnet is a powerful tool, capturing complex, functional cell states from multimodal data. SIGNIFICANCE:Defining critical cell states among cells that reside along a phenotypic continuum is a current biological and computational challenge. In this study, we present AAnet, a neural network that learns archetypal cell states of cancer cells. AAnet defines discrete spatially localized ATs that resolve intratumoral heterogeneity.
Inverse problems governed by partial differential equations (PDEs) are crucial in science and engineering. They are particularly challenging due to ill-posedness, data sparsity, and the added complexity of irregular geometries. Classical PDE-constrained optimization methods are computationally expensive, especially when repeated posterior sampling is required. Learning-based approaches improve efficiency and scalability, yet most are designed for regular domains or focus on forward modeling. Here, we introduce GeoFunFlow, a geometric diffusion model framework for inverse problems on complex geometries. GeoFunFlow combines a novel geometric function autoencoder (GeoFAE) and a latent diffusion model trained via rectified flow. GeoFAE employs a Perceiver module to process unstructured meshes of varying sizes and produces continuous reconstructions of physical fields, while the diffusion model enables posterior sampling from sparse and noisy data. Across five benchmarks, GeoFunFlow achieves state-of-the-art reconstruction accuracy over complex geometries, provides calibrated uncertainty quantification, and delivers efficient inference compared to operator-learning and diffusion model baselines.