The performance of organic electronic devices is closely tied to their nanoscale donor-acceptor microstructure, yet quantifying these features remains challenging using conventional characterization tools. X-ray scattering and electron microscopy techniques provide high-fidelity structural information, but they are slow, expensive, and challenging to deploy in high-throughput or autonomous processing environments. Here, we propose a complementary, proxy-based route to microstructure inference that leverages the device's transient short-circuit current response under modulated illumination. Using a microstructure-aware excitonic drift-diffusion (EDD) framework that incorporates arbitrary time-dependent generation profiles, we compute current (J(t)) responses for 500 computationally generated donor-acceptor morphologies subjected to full-wave-rectified sinusoidal excitation. From each response, we extract a suite of physically motivated time-and frequency-domain features and pair them with key microstructural descriptors, specifically interfacial area, characteristic domain size, and connectivity. SHAP-based feature selection reveals that a compact subset of transient-response features captures most of the microstructure dependence. Linear surrogate models trained on these subsets achieve good test R2 across all descriptors, with interfacial area and domain-size predictions exceeding R2 = 0.9, using fewer than 10% of the simulated morphologies for training. These results demonstrate that high-frequency, amplitude-modulated electrical measurements can recover essential microstructural fingerprints, suggesting a pathway toward compact, non-destructive, and automation-ready characterization tools for high-throughput research in organic electronics.
Developing neural surrogates for unsteady flow that generalize across complex geometries and maintain accuracy over extended rollouts remains a critical challenge for accelerating flow simulations. We present a time-dependent, geometry-aware Deep Operator Network that predicts velocity fields for moderate-Re flows around parametric and non-parametric shapes. The model encodes geometry via a signed distance field (SDF) trunk and flow history via a CNN branch, and is trained using 841 high-fidelity simulations from the FlowBench dataset. The model attains similar to 5% relative L2 single-step error and up to 1000 & times; speedups over CFD on 262 test geometries not used for training. We provide physics-centric rollout diagnostics, including probe-based phase lag and Strouhal frequency analysis, to quantify long-horizon fidelity. These reveal accurate near-term transients but systematic error accumulation in fine-scale wakes, most pronounced for sharp-cornered geometries. We analyze failure modes and outline practical mitigation strategies including physics-informed regularization and diffusion-based refinement. Code, splits, and scripts are openly released here to support reproducibility and benchmarking.
Accurate in-season prediction of seed yield and seed composition traits such as oil and protein are useful for gaining accuracy and efficiency in soybean breeding. These predictions can also inform farmers, enabling them to improve their field management practices, and guide their market decisions. We report a Transformer-based deep learning framework built on 30 years of multi-environment performance data from the Northern and Southern Uniform Soybean Tests (UST) across North America. Unlike earlier studies on seed yield, oil and protein prediction that focus on limited years, regions, single modalities, we utilized a comprehensive dataset that includes weather, genotype, and management factors, ensuring a more holistic approach to soybean yield, oil, and protein prediction. Our model integrates multivariate time-series weather data with genotypic relationship information, maturity group, and geographic location, to predict variety performance in diverse environments. Our model captures complex temporal patterns associated with trait variability; showing high predictive accuracy (R2) of 77.6 ± 0.2%, 63.9 ± 4.7%, and 79.3 ± 2.3% for seed yield, oil, and protein, respectively. Additionally, for seed yield, we also evaluated multiple interpretability methods to assess feature importance for predictor variables and critical growing timepoints, and solar radiation and temperature were noted as the key predictors. Overall, these results demonstrate the usefulness of a Transformer-based model in trait predictions, and the utility of large cooperative datasets from breeding programs.
Materials used in polymer-based additive manufacturing processes, such as Digital Light Processing (DLP) and direct ink writing (DIW), typically exhibit non-Newtonian rheology. Carreau–Yasuda and power-law models describe basic shear-thinning and shear-thickening behavior well, but applying them to a new material requires choosing a functional form, deriving it, and re-implementing it inside the flow solver. We present a deployment workflow in which a neural network trained on experimental rheometry data serves as the viscosity closure inside a Cahn–Hilliard–Navier–Stokes (CHNS) finite element solver. Lipschitz regularization during training produces smooth viscosity predictions, and the trained network is exported in the Open Neural Network Exchange (ONNX) format and queried by the solver at runtime via the ONNX runtime, without solver modification or network reimplementation. The framework is built on a parallel octree-based adaptive mesh refinement infrastructure that concentrates resolution at the fluid interface. We validate the CHNS solver against benchmark shear-thinning bubble-rise cases from the literature, reproducing reported bubble shapes across varying power-law indices and Weber numbers. We characterized two silicone ink formulations, recorded their rise dynamics in perfluorodecalin on high-speed video, and used the resulting data to test the full workflow. Simulated rise velocities fall within the experimentally measured spread, and the simulated steady-state droplet shape agrees with the observed one. This work contributes to a growing body of literature on integrating neural constitutive closures into multiphysics simulations, and demonstrates a practical path for deploying experimentally trained rheological surrogates inside finite element solvers.
We extend the shifted boundary method (SBM) to the simulation of incompressible fluid flow using immersed octree meshes. Previous work on SBM for fluid flow primarily utilized two- or three-dimensional unstructured tetrahedral grids. Recently, octree grids have become an essential component of immersed CFD solvers, and this work addresses this gap and the associated computational challenges. We leverage an optimal (approximate) surrogate boundary constructed efficiently on incomplete and adaptive octree meshes. The resulting framework enables the simulation of the incompressible Navier-Stokes equations in complex geometries without requiring boundary-fitted grids. Simulations of benchmark tests in two and three dimensions demonstrate that the Octree-SBM framework is a robust, accurate, and efficient approach to simulating fluid dynamics problems with complex geometries.
Doubled haploid (DH) technology can fast-track crop breeding. Haploid induction yields haploids with only one set of genomes, which are usually sterile. Haploid fertility (HF) is the ability of haploid plants to set seed, and it is a critical bottleneck in DH pipelines. Genetic mechanisms to restore HF hold immense potential in DH crop breeding, yet its phenotyping remains manual, destructive, and inconsistent. While recent advances in imaging and machine learning have improved throughput for general plant traits, no curated image dataset exists for Arabidopsis thaliana that explicitly represents HF. Here, we present AutoSiQ, a dataset and baseline deep learning pipeline for automated HF quantification. AutoSiQ includes high-resolution scanned inflorescences annotated with a seven-class ontology encompassing green siliques, green fertile siliques, mature siliques, fertile siliques, cracked fertile siliques, cracked siliques, and flowers. This multi-class annotation scheme preserves biologically meaningful information beyond binary fertile/non-fertile distinctions, enabling reliable fertility estimation and future phenotyping applications. We release baseline object detection models (YOLOv5), trained using the AutoSiQ dataset, and evaluate their performance across confidence thresholds. Model predictions strongly correlate with manual counts, achieving R² up to 0.94 for total silique number estimation. We further demonstrate AutoSiQ’s utility for automated haploid fertility rate (HFR) estimation and genotype discrimination between two contrasting genotypes (WT and bmf2 mutant). A longitudinal analysis identifies ~60 days after sowing (DAS) as the optimal harvest time for maximizing mature silique counts by balancing between the number of immature buds and silique shattering. By releasing both the dataset and baseline code, AutoSiQ provides a reproducible and extensible foundation for high-throughput fertility phenotyping in haploid Arabidopsis.
Global food security depends on predicting crop responses to climate variability, yet process based crop models remain too computationally expensive for large scale exploration of genotype and environment interactions. Here we develop a probabilistic neural emulator of APSIM that reproduces key maize growth processes across 13 outputs with high fidelity (with R^2 of 0.93) while reducing simulation time by several orders of magnitude. Trained on two million simulations spanning diverse genetic, soil, and management conditions, and augmented with a convolutional synthetic weather generator that produces physically consistent climate sequences, the framework enables scalable exploration of crop responses under realistic and diverse environmental inputs while providing calibrated predictive uncertainty without costly Bayesian inference. Applying this framework across 100,000 trait configurations, six soil environments in Iowa and Illinois, and climate projections through the year 2100 under two emissions scenarios, we identify 181 maize trait combinations that consistently maintain high yield across all tested conditionsan analysis infeasible with the mechanistic model alone. We further show that radiation use efficiency and temperature driven root dynamics are dominant drivers of yield resilience. Notably, projected yield distributions vary substantially across locations, with some lower productivity sites exhibiting yield increases under future climate scenarios, indicating that climate change may reshape regional yield potential in nonintuitive ways. These results demonstrate how uncertainty aware emulation transforms mechanistic crop simulation from a computational bottleneck into an on demand discovery engine, one capable of interrogating the full genotype, environment and management space at a scale no process-based model can match.
Optimizing the processing window for organic photovoltaics (OPVs) can deliver high efficiency and yield. In practice, small drifts in solution concentration, additive fraction, drying rate, or ambient humidity readily perturb the evolving morphology, making champion devices hard to reproduce. We combine a complementary Taguchi and space-filling design-of-experiments strategy with machine-learning surrogates to map processing robustness in PM6:Y6 films using acetone and 1-chloronaphthalene as the processing additives. Robustness is quantified via parameter-space fraction (share of conditions exceeding a power conversion efficiency threshold), persistence curves (how that share contracts as the threshold tightens), and fill fraction (compactness of high-performing regions). The parameter space landscape when using acetone as a solvent additive exhibits a broad, robust regime, with around 73% of the explored space exceeding 9% PCE; this is consistent with our hypothesis that faster evaporation yields morphology tolerant to routine variability. This ML-enabled framework disentangles manufacturing-relevant robustness from peak PCE and offers a general, data-efficient way to compare processing-window robustness with reduced experimental effort. Such tools integrate naturally with self-driving laboratories and can identify new processing opportunities for OPVs.
Plant phenotyping in precision agriculture increasingly requires high-fidelity three-dimensional reconstruction and accessible visualization methods. This study presents an integrated pipeline combining Neural Radiance Fields (NeRF), 3D Gaussian Splatting (G-Splat), and Virtual Reality (VR) visualization for comprehensive plant analysis across developmental stages. We collected multi-view imagery of finger millet, proso millet, mungbean, and field pea under controlled greenhouse conditions, aligning data acquisition with standardized BBCH phenological scales. Camera pose estimation was performed using GLOMAP, followed by reconstruction via both Nerfacto and G-Splat implementations. Quantitative evaluation using PSNR, SSIM, and LPIPS metrics revealed complementary strengths of the two approaches: G-Splat achieved superior structural fidelity, while NeRF provided enhanced perceptual realism. Both reconstruction methods were successfully integrated into an immersive VR greenhouse environment deployed on Meta Quest headsets, maintaining consistently high framerates. This framework establishes a practical foundation for incorporating neural reconstruction and immersive technologies into agricultural phenotyping workflows, supporting both research applications and educational engagement.
Maize ear geometry (length, width, curvature, and volume) is closely tied to yield and grain-filling outcomes, but existing high-throughput phenotyping pipelines remain constrained by the cost, labor, and specialized hardware they require. We developed and validated a low-cost pipeline that reconstructs a watertight 3-D mesh of a maize ear from a single 20-second video captured with a consumer-grade DSLR on a motorized turntable under uniform LED illumination. Camera poses from a multi-seed COLMAP procedure initialize a Neural Radiance Field (NeRF), and a cylindrical holder of known diameter, visible in every frame, provides automatic metric scaling with downstream geometric quality control. Applied to 300 ears spanning a diverse maize inbred panel, 250 (83.3
Plant disease diagnosis is critical for food security, yet training disease-recognition models that generalize across crops, pathogens, and field conditions remains challenging because labeled disease images are far less abundant and standardized than data for other biotic stresses such as insects or weeds. Frontier vision-language models offer new opportunities through improved visual reasoning, but they still struggle with fine-grained disease identification due to the lack of structured, crop-specific symptom knowledge. To address this gap, we curate the largest plant disease image–symptom dataset to date, covering 335 crops, 1,251 disease classes, and approximately 839K images, designed to support training-free, agentic disease prediction. A scalable automated pipeline generates source-grounded symptom descriptions in which each claim is linked to a verbatim web quote; domain experts validate sampled crops and reconcile disease-name variants across sources. As a baseline, we introduce an autonomous visual reasoning agent that identifies anatomical context, narrows candidate diseases using symptom knowledge, sequentially compares reference images, and produces a fully explainable reasoning trace. Incorporating symptom knowledge improves accuracy by 16.2 percentage points on average at the full reference budget, with consistent gains across all four evaluation crops. Because the framework only requires crop-specific reference images and symptom knowledge, it can be extended to new crops without retraining, while the agentic baseline can directly benefit from future improvements in foundation model capabilities. Dataset and code are available at:https://sage-dataset.github.io/.
Implicit Neural Representations (INRs) provide compact models of geometry, but it is unclear when their learned shapes can be edited without retraining. We show that the Gram operator induced by the INR's penultimate features admits deformation eigenmodes that parameterize a family of realizable edits of the SDF zero level set. A key finding is that these modes are not intrinsic to the geometry alone: they are reliably recoverable only when the Gram operator is estimated from sufficiently rich sampling distributions. We derive a single closed-form update that performs geometric edits to the INR without optimization by leveraging the deformation modes. We characterize theoretically the precise set of deformations that are feasible under this one-shot update, and show that editing is well-posed exactly within the span of these deformation modes.
Soft and conducting organic materials are ideal candidates for stretchable bioelectronics and wearable devices. Despite recent advances, our understanding of conducting polymer nanostructures and how they arise remains incomplete, given the limited high-resolution studies and molecular-level descriptions of these systems. Here, we employ cryogenic transmission electron microscopy (cryo-EM) to investigate the evolution of poly(3,4-ethylenedioxythiophene):polystyrene sulfonate (PEDOT:PSS) morphology in solution and the resulting solid state structure in the presence of ionic and molecular additives. Our results reveal the formation of heterostructural elongated fibers consisting of PEDOT:PSS micelles in solution. Cryo-EM further reveals that additives increase the number of fibrils, in addition to inducing the formation of crystalline domains. We observe that fibril and crystalline phases in solutions act as a template for the growth of these nanostructures in the solid state. Furthermore, exploiting cryo-EM reveals the role of solid-liquid interactions in PEDOT:PSS through the imaging of PEDOT:PSS nanostructures after the hydration of thin films. Hydration leads to the swelling of heterostructural fibers while reducing the crystalline domain size. Such behavior explains the mechanical robustness of PEDOT:PSS thin films processed with various additives as well as the high electrical conductivity of PEDOT:PSS in applications such as organic electorchemical transistors.
Generative artificial intelligence (AI) has emerged as a powerful paradigm for molecular discovery. Current molecular language models (MLMs) face several limitations, including restricted training coverage that is biased toward drug-like molecules, weak alignment between latent space and chemical structure, and incomplete reconstruction fidelity. To address these issues, we introduce MolGen-Transformer, a transformer-based generative model trained on a 198 million-molecule dataset that expands beyond typical pharmaceutical chemistries. Using the SELFIES representation, the model achieves 100% reconstruction accuracy, ensuring chemically valid outputs via the learned latent space. Three sampling strategies are developed and deployed to explore the latent space: (1) diverse molecule generation via normal distribution sampling, yielding a Tanimoto diversity of 0.93; (2) chemically similar molecule generation via neighborhood search with Pareto frontier selection to support multi-metric similarity control; and, (3) interpolation between molecule pairs to identify structural intermediates, enabling latent geometry analysis and scaffold transitions. Importantly, these strategies can support approaches such as scaffold hopping, optimization, and multi-objective molecular design. MolGen-Transformer demonstrates a dense and chemically structured latent space, capable of both broad exploration and fine-grained molecule editing.
This study evaluated high-resolution multispectral (MS) imagery from the Pléiades Neo satellite constellation for seed yield (SY) prediction and phenomic-assisted selection (PAS) in a soybean cultivar development program. Data were collected during the 2022–2024 growing seasons in Iowa from more than 54,000 progeny-row (PR) and yield-trial (YT) plots at three time points per season. Random forest models were trained using spectral plot features derived from raw bands (RBs), RGB band-based vegetation indices (RGB VIs), and multispectral vegetation indices (MS VIs). MS VIs provided the most consistent predictive performance, with average R² values of 0.53 for PR and 0.65 for YT. While RBs performed comparably, RGB VIs were weaker. Later-season imagery had the greatest predictive value, and two well-timed acquisitions may capture most of the useful information. A leave-one-trial-out method showed that at a 30% selection threshold, MS VI models achieved mean sensitivity, accuracy, and specificity of 0.54, 0.72, and 0.80, respectively, in both PR and YT datasets. In independent 2024 YT-2 trials, satellite-based PAS showed moderate-to-strong agreement (50-70%) with breeder selections across maturity-group ranges. These results demonstrate that high-resolution satellite imagery can predict SY at the breeding-plot scale and support scalable advancement and culling decisions in soybean cultivar development.
Computational design of high-efficiency organic photovoltaics requires clear links between three-dimensional active-layer morphology and device performance. We present a data-driven workflow that first uses coreset selection to distill a large library of simulated morphologies into a small, representative subset, thereby focusing expensive morphology-aware exciton drift–diffusion simulations where they matter most. Using two device performance metrics, short-circuit current density, J_SC , and fill factor, FF, from these simulations, we then apply feature-selection strategies to identify a handful of interpretable morphological descriptors that accurately predict both quantities. Sample-size ablations show that model accuracy, and the identity and rankings of the selected descriptors, remain stable with as few as 50 samples with the device performance. The descriptors most predictive of J_SC differ from those for FF, reflecting distinct morphological bases for these two performance metrics. Moreover, cross-system comparisons (P3HT:PCBM vs. PM6:Y6) reveal shifts in the most influential descriptors, indicating that morphology–performance relationships are material-specific.
Mungbean (Vigna radiata (L.) R. Wilczek) is a vital source of digestible proteins and is well-suited for the plant-based protein industry. In this study, we analyzed pod morphological traits in the Iowa Mungbean Diversity (IMD) panel of 372 genotypes (2022-2023) using image-analysis-based phenotyping on 2,418 pod images. Pod morphological traits were extracted using deep learning image analysis, achieving excellent agreement with manual measurements (r > 0.96 for pod length (PL) and seed-per-pod (SPP)). Four complementary genome-wide association studies models identified 65 significant SNPs (-log10(P) ≥ 5.56) associated with pod curvature, length, width, and SPP traits. A significant SNP (5_35265704) on chromosome 4 was linked to pod dimensional traits, length, width, and curvature. A candidate gene, Virad04G0076900, located 15.6 kb from this SNP, is part of the GH3 gene family and has an Arabidopsis ortholog (AT4G27260) known for influencing organ elongation, pod, and seed development. Another SNP, 5_210437 on chromosome 6, has been found to be significantly associated with both PL and SPP. A candidate gene, Virad06G0002400 (36.5 kb from this SNP), encodes a potassium transporter and shares homology with the Arabidopsis gene HAK5 (AT4G13420), known to influence pod growth. Image-based measurements achieved genomic prediction accuracies ranging from 0.61 to 0.85 across various traits, demonstrating comparable accuracy to manual methods for linear traits and up to 22% improvement for complex shape traits. These results highlight the potential of deep learning-assisted phenomics integrated with genomic tools to accelerate selection for improved pod architecture in mungbean breeding programs across the Midwestern United States and globally.
Organic photovoltaics (OPVs) require joint optimization of materials parameters and active-layer microstructure to maximize device performance, including short-circuit current, 𝐽 𝑠𝑐 . We present a material property and microstructure-aware surrogate that predicts across diverse donor-acceptor systems and microstructures while providing calibrated uncertainties and interpretable design rules. We curate a 25k-sample, physics-informed dataset by sweeping electron and hole mobilities (𝜇 𝑛 , 𝜇 𝑝 ) and exciton lifetime (𝜏 𝑥 ) across multiple microstructure classes, and derive a compact feature set that couples materials parameters with graph-based microstructure descriptors (interfacial area, and connectivity/tortuosity) to predict 𝐽 𝑠𝑐. A lightweight random-forest model achieves 𝑅 2 ≥ 0.98 with ≤ 40% of the data used for training, and maintains accuracy under material-wise holdouts (i.e., material property holdouts). Partial-dependence analyses reveal regime transitions in the coupled design space: at low mobilities and short 𝐿 𝑑 , performance is dual-transport-limited; beyond a mobility threshold, 𝜏 𝑥 (hence 𝐿 𝑑 ) and interfacial proximity dominate 𝐽 𝑠𝑐 . The framework accelerates exploration of material-microstructure trade-offs, identifies whether to intervene via material (mobilities, and 𝜏 𝑥 ) or processing (phase-separation length scale, connectivity, interfacial area), and is extensible to other organic optoelectronic systems. Code, features, and data are released to enable reproducibility and reuse.
Highly conducting polymers are essential for next-generation wearable electronics. However, achieving high conductivity remains an art form owing to complex intermolecular dopant-polymer interactions. In this study, we use AI-guided high-throughput experimentation combined with quantum chemical calculations to explore samples of diverse polymer order and polaron delocalization to reveal hidden correlations between charge transport, polymer order, carrier delocalization, and dopant location in F4TCNQ-doped pBTTT. We find that undoped aggregation benefits polaron delocalization and conductivity after doping, and lamellar stacking order correlates with two orders of magnitude variation in carrier mobility and highly influences polaron delocalization. Using quantum chemical theory, we deduce that increased mobility originates from highly delocalized polarons formed by "peripheral" counterions located at distances (approximate to 1.3-1.8 nm) much greater than those of the lamellar intercalated counterions (approximate to 0.4-0.8 nm). We find that achieving high conductivity (sigma > 100 S/cm) in F4TCNQ-doped pBTTT requires processing conditions promoting ordered domains decorated by peripheral counter ions.
Advances in hyperspectral imaging (HSI) and 3D reconstruction have enabled accurate, high-throughput characterization of agricultural produce quality and plant phenotypes, both essential for advancing agricultural sustainability and breeding programs. HSI captures detailed biochemical features of produce, while 3D geometric data substantially improves morphological analysis. However, integrating these two modalities at scale remains challenging, as conventional approaches involve complex hardware setups incompatible with automated phenotyping systems. Recent advances in neural radiance fields (NeRF) offer computationally efficient 3D reconstruction but typically require moving-camera setups, limiting throughput and reproducibility in standard indoor agricultural environments. To address these challenges, we introduce HSI-SC-NeRF, a stationary-camera multi-channel NeRF framework for high-throughput hyperspectral 3D reconstruction targeting postharvest inspection of agricultural produce. Multi-view hyperspectral data is captured using a stationary camera while the object rotates within a custom-built Teflon imaging chamber providing diffuse, uniform illumination. Object poses are estimated via ArUco calibration markers and transformed to the camera frame of reference through simulated pose transformations, enabling standard NeRF training on stationary-camera data. A multi-channel NeRF formulation optimizes reconstruction across all hyperspectral bands jointly using a composite spectral loss, supported by a two-stage training protocol that decouples geometric initialization from radiometric refinement. Experiments on three agricultural produce samples demonstrate high spatial reconstruction accuracy and strong spectral fidelity across the visible and near-infrared spectrum, confirming the suitability of HSI-SC-NeRF for integration into automated agricultural workflows.