Accurate 3D dental model segmentation is critical for digital dental treatment, as it provides valuable clinical references. Existing methods fail to adaptively evaluate the importance or contribution of different geometric attributes during heterogeneous features fusion, hindering the accuracy of end-to-end segmentation. In this paper, we pioneer the description of geometric attributes of 3D dental model as views. A multi-view geometry-adaptive fusion network (MGAFNet) is proposed to dynamically seek the optimal combination of views through distinctive and sharable features exploration for fine-grained 3D dental model segmentation. Specifically, during distinctive features extraction, we design geometry-aware enhancement module (GAE) to improve topological variations learning in teeth. After that, a multivariate sharable cross-interaction module (SCIM) is developed to facilitate the flow of information and capture sharable features among views. Subsequently, a multivariate adaptive representation fusion module (MARF) is implemented to adaptively balance the importance or contribution of views by constructing weight matrices for distinctive and sharable features from different feature sources. Compared to eight advanced methods, our MGAFNet achieves state-of-the-art performance on both a public benchmark and a private clinical dataset. It demonstrates robustness in handling various dental conditions (e.g., misaligned, missing and supernumerary teeth), avoiding category confusion and blurry boundary segmentation.
Engineering enzymes with enhanced activity and stability is a central goal of biotechnology, yet the inherent trade-off between optimizing global protein fitness and specific substrate binding affinity poses a significant challenge. Here, we present ESM-FEP, a computational framework that synergistically integrates a fine-tuned protein language model with alchemical free energy perturbation (FEP) to overcome this limitation. Our workflow employs a parameter-efficient fine-tuned ESM-2 model to perform high-throughput saturation mutagenesis, rapidly identifying mutations that preserve protein fitness. Top-ranking candidates are then subjected to rigorous FEP simulations to precisely quantify changes in substrate binding affinity. When applied to engineer the Zea mays dioxygenase ZmHSL1B for improved detoxification of the herbicide mesotrione, ESM-FEP efficiently navigated the mutational landscape and identified a quadruple mutant M5 (Q140H/Y205F/L332R/K336F). This variant demonstrated a catalytic efficiency approximately 7-fold higher than that of the wild-type enzyme, which was corroborated by in vitro assays and a detailed kinetic analysis. Furthermore, transgenic Arabidopsis thaliana expressing the engineered mutant M5 exhibited significantly enhanced herbicide tolerance, validating its functional efficacy in a biological context. The ESM-FEP framework establishes a generalizable and efficient strategy for the rational design of gain-of-function enzymes, with broad applications in biocatalysis, bioremediation, and precision agriculture.
Accurate segmentation of brain tissues from MRI scans is critical for neuroscience and clinical applications, but achieving consistent performance across the human lifespan remains challenging due to dynamic, age-related changes in brain appearance and morphology. While prior work has sought to mitigate these shifts by using self-supervised regularization with paired longitudinal data, such data are often unavailable in practice. To address this, we propose DuMeta++, a dual meta-learning framework that operates without paired longitudinal data. Our approach integrates: (1) meta-feature learning to extract age-agnostic semantic representations of spatiotemporally evolving brain structures, and (2) meta-initialization learning to enable data-efficient adaptation of the segmentation model. Furthermore, we propose a memory-bank-based class-aware regularization strategy to enforce longitudinal consistency without explicit longitudinal supervision. We theoretically prove the convergence of our DuMeta++, ensuring stability. Experiments on diverse datasets (iSeg-2019, IBIS, OASIS, ADNI) under few-shot settings demonstrate that DuMeta++ outperforms existing methods in cross-age generalization. Code will be available at https://github.com/ladderlab-xjtu/DuMeta++.
Existing BraTS-GLI datasets provide a widely used benchmark for adult glioma MRI segmentation, but their task definition focuses on tumor subregions and does not systematically represent coexisting white matter hyperintensities (WMH). In joint segmentation settings, such unlabeled abnormalities introduce task-specific label noise by treating pathological regions as normal tissue. To address this limitation, we introduce BraTS-GLI Anatomy-Lesion, a controlled-access, labels-only derived resource built from the BraTS 2023-GLI training cohort. The resource provides 1,251 unified eight-class anatomy-lesion label sets aligned with the original four-modal MRI cases, including image-repair labels for 116 cases requiring repaired imaging inputs. The cohort is organized into a 394-case purified subset and an 857-case extended subset, with case-level metadata covering label source, image-repair requirements, quality-control status, access conditions, checksums, and release boundaries. Compared with the original BraTS-GLI annotations, the resource substantially expands foreground supervision by incorporating healthy brain tissues and previously unlabeled coexisting abnormalities within a unified label space. A validation study using MedNeXt and T1/FLAIR inputs suggests that WMH-aware supervision preserves healthy-tissue segmentation performance across both in-domain GLI and external WMH datasets, while improving sensitivity to coexisting lesions relative to noisy-control training. The resource is intended for scientific research and supports joint anatomy-lesion supervision, label-noise analysis, and reproducible evaluation. Data are available at https://www.synapse.org/Synapse:syn75210889/wiki/, and code is available at https://github.com/xyx200/brats-gli-anatomy-lesion-code. The data resource DOI is https://doi.org/10.7303/SYN75210889.
Positron emission tomography (PET) is pivotal in medicine and healthcare, but requires high radiation doses for quality imaging. Although deep learning has shown promise for low-dose PET reconstruction/enhancement, existing methods primarily rely on fully supervised learning, which requires paired low- and full-dose scans that are clinically impractical to obtain. To overcome this fundamental limitation, we propose DuSS-PET, a novel self-supervised PET reconstruction framework that does not require paired low-/full-dose images as supervised training targets in the main reconstruction stage. Our method is built upon a computational Noisier2Noise (Nr2N) paradigm, featuring a Bernoulli corruption strategy that serves as an image-space approximation to emulate low-count noise patterns in image-derived pseudo-sinogram representations. The DuSS-PET framework integrates three core components: an adaptive Swin transformer for pseudo-sinogram restoration, a plug-and-play module for image enhancement, and a consistency-constrained SSL strategy that enforces agreement between the pseudo-sinogram and image representations. Extensive experiments on multi-center datasets demonstrate that DuSS-PET achieves superior reconstruction performance, outperforming state-of-the-art SSL methods and achieving comparable or even better results than fully supervised approaches across various dose levels. In particular, it exhibits exceptional generalization across different scanner vendors. This work is among the first self-supervised low-dose PET reconstruction frameworks that enforce consistency between reconstructed images and pseudo-sinograms, where the pseudo-sinogram generated in the image space serves as a computational consistency constraint for enhanced imaging without paired full-dose training targets. The source code will be publicly available at https://github.com/ladderlab-xjtu/DuSS-PET.
This study aims to develop a two-stage 3D denoising diffusion implicit model (DDIM) framework for CT-free attenuation correction in cardiac PET imaging, enabling direct generation of attenuation-corrected (AC) images from non-attenuation-corrected (NAC) PET scans. The method is comprehensively validated using both [18F]FDG PET and [13N]ammonia cardiac PET datasets to demonstrate its clinical applicability across different perfusion and metabolic imaging protocols. The framework employs a two-stage approach: (1) a noise-to-image DDIM was first pretrained on all available AC images (i.e., no need of paired NACs) to learn a diverse AC distributions, enabling the high-fidelity generation of AC images with varying appearances; (2) the pretrained model was fine-tuned with a limited set of paired NAC-AC images to form a conditional DDIM, ensuring anatomically aligned, controllable generation. The model architecture uses a 3D U-Net, trained on 224 paired NAC-AC and 396 unpaired AC images for [18F]FDG, and 608 paired NAC-AC images and 885 unpaired AC images for [13N]ammonia. Performance was evaluated through quantitative metrics (including NMAE, NRMSE, SSIM and PSNR) and visual assessment. The proposed two-stage DDIM framework achieved excellent agreement with clinical CT-based attenuation correction (CT-AC), demonstrating superior correlation (slope = 0.78, R^2 = 0.95 for [18F]FDG; slope = 0.99, R^2 = 0.91 for [13N]ammonia) and lower errors compared to existing approaches. Ablation studies confirmed the benefits of both the two-stage training strategy and the incorporation of unpaired AC images, as evidenced by narrower confidence intervals in Bland-Altman analysis and reduced percentage errors. The two-stage 3D DDIM framework achieves performance comparable to clinical CT-AC while effectively leveraging unpaired data, demonstrating significant potential for robust cardiac PET attenuation correction.
Forensic histopathology, essential for determining cause of death and disease diagnosis, is severely impeded by postmortem autolysis, i.e., an irreversible, stochastic degradation process that distorts tissue morphology and introduces diagnostic subjectivity, thereby underscoring the value of restoring autolyzed images to a diagnostically plausible, pre-autolysis state for improving objectivity in forensic practice. This restoration task is fundamentally challenging due to the large, non-deterministic morphological changes caused by autolysis and the infeasibility of pixel-wise paired data, which invalidates assumptions underlying supervised and cycle/structure-consistent unpaired translation methods. To address this, we formalize forensic histopathology autolysis restoration as a new task: under unpaired supervision, transform postmortem images with severe autolysis into diagnostically meaningful “antemortem” representations. We contribute AutoPath, the first homologous yet unpaired dataset for this problem, constructed by splitting specimens into adjacent tissue blocks—one processed immediately, the other exposed to induce autolysis—yielding nearly ten thousand 10× patches from 69 cases with varying liver conditions. We further frame the problem as a Schrödinger Bridge between the autolyzed and non-autolyzed distributions, offering a principled approach to modeling stochastic, severe morphological degradation. Critically, we demonstrate the misalignment of generic image-level generative metrics (e.g., FID) with diagnostic utility and propose a forensically grounded, slide-level diagnostic distribution consistency evaluation. Overall, this work establishes a reproducible benchmark (encompassing task definition, a real-world dataset, and an evaluation methodology) toward rigorous and practically meaningful progress in autolysis restoration for forensic pathology.
Ecotoxicity assessments, which rely on animal testing, face serious challenges, including high costs and ethical concerns. Computational toxicology presents a promising alternative; nevertheless, existing predictive models encounter difficulties such as limited datasets and pronounced overfitting. To address these issues, we propose a framework for predicting pesticide ecotoxicity using graph contrastive learning (PE-GCL). By pre-training on large-scale unlabeled compounds, the PE-GCL captured the intrinsic regulation of molecules. This knowledge is then transferred to specific downstream tasks, thereby enhancing the model generalization in scenarios with small sample sizes. Performance evaluation showed that the PE-GCL outperformed traditional supervised models across most prediction tasks, whereas independent external validation confirmed its superior predictive accuracy for unseen data. Furthermore, interpretability was incorporated to elucidate potential correlations between ecotoxicity and molecular substructures. The trained models were deployed on a publicly accessible web server (https://dpai.ccnu.edu.cn/PERA/) to facilitate the use of the proposed framework.
BACKGROUND:Pesticides are crucial for protecting crops from pests and diseases to meet the growing global food demand. In the era of artificial intelligence (AI), computer-aided approaches have the potential to significantly enhance the efficiency and safety of pesticide design. The application of these methods depends on comprehensive and accurate pesticide data, essential for developing effective pesticides. However, there remains a lack of integrated and user-friendly databases to display complete information on pesticides. RESULTS:Digital Pesticide is a comprehensive online platform, containing information on over 2000 pesticides, each with nearly 200 curated attributes. These include identifiers, chemical properties, agrochemical data, regulatory guidelines, environmental impact assessments, and potential human health effects. A web-based user interface has been developed to facilitate efficient data management and access. This open-access platform allows users to browse, search, and download information according to their specific needs. Digital Pesticide is free to use on https://dpai.ccnu.edu.cn/digpesticide/. CONCLUSION:Digital pesticide is a comprehensive pesticide database with a dynamic web platform and it is the cornerstone of the drive for sustainable agricultural development and management. It bridges a critical gap in pesticide research and is expected to serve as a key resource for global pesticide and pest management. © 2025 Society of Chemical Industry.
Cortical parcellation delineates the cerebral cortex into distinct regions according to their distinctiveness in anatomy and/or function, which is a fundamental preprocess in brain cortex analysis and can influence the accuracy and specificity of subsequent neuroscientific research and clinical diagnosis. Conventional methods for cortical parcellation involve spherical mapping and multiple morphological feature computation, which are time-consuming and prone to error due to the spherical mapping process. Recent geometric learning approaches have attempted to automate this process by replacing the registration-based parcellation with deep learning-based methods. However, they have not fully addressed spherical mapping and cortical features quantification, making them sensitive to variations in mesh structures. In this work, to directly parcellate original surfaces in individual space with minimal preprocessing, we present a full-band spectral-accelerated spatial diffusion strategy for stable information propagation on highly folded cortical surfaces, contributing to adaptive learning of fine-grained geometric representations and the construction of a compact deep network (termed Cortex-Diffusion) for fully automatic parcellation. Using only raw 3D vertex coordinates and having merely 0.49 MB of learnable parameters, it demonstrates state-of-the-art parcellation accuracy, efficiency, and superior robustness to mesh resolutions and discretization patterns in both the cases of infant and adult brain imaging datasets.
BackgroundAccording to the 2021 WHO classification of tumors of the central nervous system, isocitrate dehydrogenase (IDH) status serve an independent prognostic biomarker and is closely associated with tumor diagnosis and treatment response. At present, the determination of IDH status still relies on invasive surgical procedures.MethodA total of 345 patients with pathologically confirmed gliomas diagnosed at the First Affiliated Hospital of Xi’an Jiaotong University between October 2019 and October 2024 were retrospectively included, comprising 148 (42.9%) IDH-wild and 197 (57.1%) IDH-mutant. An additional 495 glioma patients were obtained from the public TCIA dataset. Patients were randomly split into training, validation, and test cohorts 6:2:2. A Hierarchical Attention-Based Multiple Instance Learning (HAB-MIL) framework was developed, integrating auxiliary positional encoding into feature maps to capture spatially specific information and generate refined 3D lesion representations. Model performance was evaluated using five-fold cross-validation, with receiver operating characteristic (ROC) curves, area under the curve (AUC), sensitivity, and specificity as assessment metrics.ResultHAB-MIL achieved competitive performance, with AUCs of 0.917 and 0.892 on the glioma datasets from TCIA and the First Affiliated Hospital of Xi’an Jiaotong University. Additionally, our work achieves results that are comparable to the state-of-the-art methods in TCIA dataset and demonstrates that multiple instance learning has great potential for IDH prediction.ConclusionThe proposed HAB-MIL achieved IDH classification based on conventional preoperative MRI images, eliminating the need for pixel-level annotations and significantly reducing the annotation burden for doctors.
Longitudinal prediction of infant brain MRIs is crucial for individualized neurodevelopment tracking and disorder forecasting. However, existing methods, such as diffusion-based generative models, often struggle to capture the complex spatiotemporal dynamics of developing brains, leading to unreliable predictions that lack subject-specific, anatomically consistent growth patterns. To address this, we propose a Flexibly Distilled 3D Rectified Flow (FDRF) framework, which integrates anatomical constraints for dual-stream predictions of volumetric images and tissue maps along developmental trajectories. Our framework features an age-conditioned feature fusion module for controllable prediction with targeted age appearances and employs anatomical constraints derived from segmentation labels and high-frequency image details to ensure subject-level spatiotemporal consistency. Additionally, we introduce a flexible distillation of rectified flow, enabling a unified one-step generative model for high-fidelity cross-time predictions while preserving individualized anatomical details. Given 6-month MRIs and tissue maps as the input, our model reliably predicts their spatiotemporal growths at 12 and 24 months, outperforming existing diffusion-based baselines by relatively large margins. Our codes can be found at https:// github.com/ladderlab-xjtu/FDRF.
Forensic pathology is critical in determining the cause and manner of death through post-mortem examinations, both macroscopic and microscopic. The field, however, grapples with issues such as outcome variability, laborious processes, and a scarcity of trained professionals. This paper presents SongCi, an innovative visual-language model (VLM) designed specifically for forensic pathology. SongCi utilizes advanced prototypical cross-modal self-supervised contrastive learning to enhance the accuracy, efficiency, and generalizability of forensic analyses. It was pre-trained and evaluated on a comprehensive multi-center dataset, which includes over 16 million high-resolution image patches, 2,228 vision-language pairs of post-mortem whole slide images (WSIs), and corresponding gross key findings, along with 471 distinct diagnostic outcomes. Our findings indicate that SongCi surpasses existing multi-modal AI models in many forensic pathology tasks, performs comparably to experienced forensic pathologists and significantly better than less experienced ones, and provides detailed multi-modal explainability, offering critical assistance in forensic investigations. To the best of our knowledge, SongCi is the first VLM specifically developed for forensic pathological analysis and the first large-vocabulary computational pathology (CPath) model that directly processes gigapixel WSIs in forensic science.
We introduce RynnEC, a video multimodal large language model designed for embodied cognition. Built upon a general-purpose vision-language foundation model, RynnEC incorporates a region encoder and a mask decoder, enabling flexible region-level video interaction. Despite its compact architecture, RynnEC achieves state-of-the-art performance in object property understanding, object segmentation, and spatial reasoning. Conceptually, it offers a region-centric video paradigm for the brain of embodied agents, providing fine-grained perception of the physical world and enabling more precise interactions. To mitigate the scarcity of annotated 3D datasets, we propose an egocentric video based pipeline for generating embodied cognition data. Furthermore, we introduce RynnEC-Bench, a region-centered benchmark for evaluating embodied cognitive capabilities. We anticipate that RynnEC will advance the development of general-purpose cognitive cores for embodied agents and facilitate generalization across diverse embodied tasks. The code, model checkpoints, and benchmark are available at: https://github.com/alibaba-damo-academy/RynnEC
Recent advancements in generative learning have enabled PET image synthesis from relatively more accessible MRI scans, offering a safer, cost-effective, and scalable alternative to traditional PET imaging, e.g., for Alzheimer's disease (AD) diagnosis. However, current MRI-to-PET translation methods face limitations in controllability and fidelity, often failing to capture personalized metabolic activations and fine-grained structural details in critical regions. To address these challenges, we propose a novel controllable MRI-to-PET translation framework, termed DIReCT, which leverages rectified flow to generate high-fidelity PET images tailored to downstream diagnostic and analytical needs. By injecting cross-modal guidance from a pretrained visionlanguage model (BiomedCLIP), DIReCT incorporates both common imaging knowledge and individualized clinical information to enhance the personalization of PET synthesis. Extensive experiments on the ADNI dataset demonstrate that DIReCT significantly outperforms existing methods across various image quality metrics. Notably, the synthesized FDG-PET images by DIReCT achieve analytical performance comparable to real FDG-PET scans, excelling in capturing AD-related pathological features for reliable group comparisons and personalized diagnosis.
BACKGROUND:Rapid advances in generative artificial intelligence (AI) are accelerating the process of pesticide development. However, transfer learning-based de novo design focuses on generating molecules that are highly similar to existing inhibitors, which may limit the exploration of novel scaffolds and thereby constrain innovative breakthroughs in pesticide development. RESULTS:This study proposes a new strategy for fungicide design using antibiotics. First, by combining pre-training and transfer learning, a character-level recurrent neural network model was able to generate antibiotic-like molecules that retained key features while avoiding excessive similarity to existing fungicides. Fungicide-like molecules were then further identified by training graph neural network models that could discriminate between fungicides and antibiotics. Interestingly, two of the generated molecules were found to share the same scaffold as florylpicoxamid, a recently approved fungicide that was not included in the training set. As a proof-of-concept, the inhibitory activity of the screened molecules against cytochrome bc1 complex was determined. Compound cp461 displayed enzyme inhibition comparable to that of florylpicoxamid, with a median inhibitory concentration of 17.9 ± 1.1 nm. CONCLUSION:Overall, the pioneering work leveraged AI and antibiotic data to facilitate the design of novel fungicides and provided a viable research idea for pesticide development. © 2025 Society of Chemical Industry.
Introduction: Pesticides play a pivotal role in ensuring food security, and the development of green pesticides is an inevitable trend in global agricultural progress. Although deep learning-based generative models have revolutionized de novo drug design in pharmaceutical research, their application in pesticide research and development remains unexplored. Objectives: This study aims to pioneer the application of generative artificial intelligence to pesticide design by proposing a reinforcement learning-based framework for obtaining pesticide-like molecules with high binding affinity. Methods: This framework comprises two key components: PestiGen-G, which systematically explores the pesticide-like chemical space using a character-based generative model coupled with the REINFORCE algorithm; and PestiGen-S, which combines a fragment-based generative model with the Monte Carlo Tree Search algorithm to generate molecules that stably bind to the specific target protein. Results: Experimental results show that the molecules generated by PestiGen have superior pesticide-likeness and binding affinity compared to those generated by existing methods. In addition, we employ an active learning strategy to reduce the false-positive rate of the generated molecules. Finally, through collaboration with domain experts, we successfully designed a novel 4-hydroxyphenylpyruvate dioxygenase inhibitor (YH23768) with favorable enzyme inhibition and herbicidal potency. Conclusion: This proof-of-concept study highlights the utility of PestiGen as a valuable tool for pesticide design. The web server based on the model is freely available at https://dpai.ccnu.edu.cn/PestiGen/.
4-Hydroxyphenylpyruvate dioxygenase (HPPD) represents a pivotal enzyme in the metabolic processes of plants, exerting a decisive influence on the catabolism of the amino acid tyrosine. It is located within the plastids of plant cells, where it functions as a catalyst for the conversion of 4-hydroxyphenylpyruvate to homogentisate (HGA; Raspail et al., 2011). This reaction represents the initial committed step in the biosynthetic pathway that culminates in the formation of plastoquinone and tocopherols (Ren et al., 2011). These compounds are indispensable for the optimal functioning of photosynthesis and the protection of plants from oxidative stress. Given its pivotal function in these processes, the regulation of HPPD is rigorously controlled at the transcriptional, posttranscriptional, and posttranslational levels. The HPPD protein has become the subject of extensive research, with the aim of understanding its structure and function, its regulation in different plant species, and its evolutionary history. Furthermore, studies on HPPD contribute to the broader understanding of plant physiology, particularly with regard to the mechanisms of stress tolerance (Ji et al., 2016; Frelin et al., 2017), the regulation of secondary metabolism (Dreesen et al., 2018; Y. Yang et al., 2024), and the interplay between metabolic pathways that are essential for plant growth and survival (Diaz-Tielas et al., 2019). Overall, HPPD represents a vital node in plant metabolic networks, with implications ranging from basic botanical research to practical applications in crop management and improvement. In recent years, research on the role of HPPD in plant metabolism and physiology has seen a notable expansion (Concepcion et al., 2021; Kim et al., 2021; Park et al., 2022; Wang et al., 2025). For example, the CRISPR-Cas9-mediated gene editing of the OsHPPD 3′ UTR resulted in the creation of new rice lines exhibiting resistance to HPPD-inhibiting herbicides (Wu et al., 2023). A fluorescent biosensor targeting HPPD was developed for noninvasive and real-time imaging of plant abiotic stress responses, thereby facilitating the detection of HPPD expression in plants under various stresses, including drought, salinity, and extreme temperatures (Fu et al., 2022). An efficient germline-specific evolution system was employed to generate herbicide-resistant HPPD variants, which could be utilized in crop breeding (Wang et al., 2024). Despite recent advances, challenges remain in understanding the molecular mechanisms underlying HPPD in plant stress tolerance, as well as in developing artificial evolution and functional interpretation of HPPD. A specialized database on the HPPD family could integrate information from multiple dimensions, thereby offering novel insights and perspectives. Here, we present the HPPD Targetome Database (HTD), a comprehensive database that covers the molecular, genomic, and biological aspects of the HPPD family. The database contains more than 18 164 enzyme entries across 11 037 species, including plants, animals, fungi, and bacteria. HTD provides comprehensive data on proteomics and genomics, including gene ontology annotations and gene expression profiles. Furthermore, the database contains information on metabolic pathway analysis and protein interaction networks, thereby providing an integrated view of the biological functions and interactions. Additionally, the HTD platform has been expanded to include a collection of chemical inhibitors for HPPD, along with their associated physicochemical properties and thermodynamic parameters. The database is freely available online at https://chemyang.ccnu.edu.cn/ccb/database/HTD/. The transcript and protein sequences were obtained from the National Center for Biotechnology Information (NCBI) GenBank and UniProt databases, respectively. The sequence data are extracted by querying the aforementioned databases with the relevant keywords. In this instance, the protein or gene name was employed as the search term (4-hydroxyphenylpyruvate dioxygenase, hppd, hpd). A deep protein language mode ESM-1b was employed to predict the potential missense variants (Rives et al., 2021). The protein structures are retrieved from the RCSB PDB, our research group, or modeled by AlphaFold2 (Jumper et al., 2021). Gene functions are annotated using the Gene Ontology database. The gene expression profiles of the microarray data were obtained from NCBI GEO database using the advanced R package geoquery. The KEGG REST API was used to retrieve specific types of pathway data. The protein interaction network was obtained from the STRING database. The chemical inhibitors were sourced from a combination of literature in Web of Science, our research group, and the BindingMOAD database. The literature was retrieved by combining keywords, including (('4-Hydroxyphenylpyruvate dioxygenase' OR HPPD) and inhibit*). The chemical structures and thermodynamic parameters were collated from the literature. The backend of the HTD platform was developed using LNMP environment (Linux, Nginx, MySQL, and PHP). We used PHP for server-side scripting and built the interactive interface using Bootstrap and jQuery. The MySQL server was used as data management, and NGL Viewer as the molecular viewer (Rose et al., 2018). Gene and protein sequence alignments were performed using NCBI Blast (v.2.2.28+) (Camacho et al., 2009). HTD is freely available online to users. We recommend using a modern web browser that supports the HTML5 standard, such as Firefox, Chrome, Safari, Edge. There is no login requirement to access data in HTD. HTD is dedicated to the development of a comprehensive and user-friendly multi-omics data platform centered on the HPPD family (Fig. 1). To this end, we have developed dedicated modules dedicated to HPPD-related proteins, encompassing domains, such as the transcriptome, proteome, gene function and expression, metabolic pathways, and interaction networks. At present, the HTD platform contains 18 164 protein entries from 11 037 species. Genetic variations have been identified from transcriptome data or predicted using a deep protein language model (Rives et al., 2021). Furthermore, 223 crystal structures have been curated, complemented by additional structures generated through AlphaFold2 modeling. The database also includes over 2000 chemical interactions with experimentally determined thermodynamic parameters. A total of 17 277 genes have been assigned gene ontology annotations, while 2953 gene expression profiles have been collated from the NCBI GEO database. Moreover, 8476 KEGG pathways have been identified, underscoring the role of HPPD in tyrosine catabolism and its interconnections with broader metabolic networks. In order to construct protein interaction networks, 8009 interaction networks were retrieved from the STRING database. The HTD platform incorporates a suite of diverse tools designed to facilitate comprehensive analysis of protein sequences and structures. A web-based user interface has been developed for the purpose of facilitating the retrieval and presentation of the data (Fig. 2). The HTD platform offers users the opportunity to examine the evolutionary patterns of the HPPD protein across species, tracing its conservation, divergence, and potential functional adaptations (Fig. 2a–c). For example, a multiple sequence alignment demonstrated that Medicago sativa HPPD (MsHPPD) exhibits a high degree of similarity with HPPD sequences from other species, including Medicago truncatula, Lactuca sativa, and Arabidopsis thaliana (Jiang et al., 2017). Furthermore, phylogenetic analysis showed that MsHPPD clusters with HPPD sequences from dicotyledonous plants, indicating its evolutionary proximity to other dicot HPPDs. The diversity of isozymes across species provides a valuable opportunity to investigate evolutionary divergence in stress resistance. The comprehensive coverage of HPPD sequences across 11 037 species offered by HTD also presents the potential to identify novel isozymes and functional variants, thereby paving the way for new research directions in enzyme evolution and ecological adaptation. The application of gene ontology annotations to HPPD-related proteins, in conjunction with gene expression profiles, enables comprehensive investigations into the mechanisms governing HPPD regulation at the transcriptional and posttranscriptional levels. Users may examine the factors that regulate HPPD expression across species and contexts, including developmental stages or environmental stress (Fig. 2d). An innovative bioorthogonal fluorescence labeling and imaging technique was developed specifically for the specific purpose of visualizing the distribution and cellular uptake of HPPD in plants without interfering with normal plant growth (Zeng et al., 2023). The technique has been employed to monitor the real-time expression and localization of HPPD in Arabidopsis thaliana under a range of abiotic stress conditions, including high temperatures, drought, salinity, and heavy metal exposure. The findings indicated a significant correlation between HPPD expression levels and stress exposure, thereby substantiating the enzyme's involvement in plant stress resistance. To elucidate the molecular mechanisms underlying the stress response, we sought to investigate the interaction network of HPPD during this process and identified dozens of key interacting proteins. This discovery provides a foundation for elucidating the mechanisms underlying the stress response associated with HPPD. HTD facilitates multi-level investigations into the impact of alterations in HPPD expression, genetic variations, or chemical interactions affect the overall metabolic equilibrium, rendering it a valuable instrument for systems biology research (Fig. 2e–f). For example, the rational redesign of the Arabidopsis thaliana HPPD enzyme results in the production of hydroxyphenylacetate (HPA) rather than the native product, HGA by altering its catalytic pathway (Lin et al., 2021). A number of HPPD mutants, including the S267W variant, were created with the objective of selectively producing HPA. Moreover, the crystal structure of the S267W mutant furnished crucial structural insights, corroborating the alternative product pathway. These findings establish a robust platform for enzyme engineering, thereby enabling the design of biocatalysts with novel functions. By integrating transcriptomic, proteomic, and metabolic pathway data, HTD facilitates the exploration of the role of HPPD within broader metabolic networks. The incorporation of KEGG pathways, particularly those pertaining to tyrosine metabolism, allows users to visualize the role of HPPD within pivotal metabolic pathways, including its connections to the biosynthesis of vital metabolites, such as plastoquinone and tocopherols. HPPD has gained significant attention as the target of a class of herbicides known as HPPD inhibitors (Santucci et al., 2017; Lin et al., 2019; T-L. Yang et al., 2024), which cause a characteristic bleaching of plant tissues, ultimately leading to plant death. The selectivity of HPPD inhibitors is of paramount importance for ensuring the safety of these inhibitors in crops. Targetome structure-based design represents a promising approach for the development of highly selective inhibitors, with a foundation in the structures from diverse species (Lin et al., 2023). HTD's collection of three-dimensional structures provides indispensable data for investigating the structure–function relationship of HPPD enzymes (Fig. 2g,h). These data allow users to investigate the impact of structural variations on inhibitor binding, which is crucial for the discovery of herbicide-resistant crops. The mesotrione-resistant variants of the HPPD enzyme were developed through directed evolution (Qian & Shi, 2024). Four critical mutations (P329S, T331A, K335E, and G412S) in the cotton HPPD gene were identified as conferring resistance to the herbicide mesotrione while preserving the enzyme's native activity. The combination of three or four of these mutations exhibited significantly higher resistance than the single mutations, demonstrating a synergistic effect. These findings establish a rationale for employing gene-editing techniques to incorporate these advantageous mutations into crops, thus enhancing herbicide resistance and optimizing weed management strategies. The HTD platform represents a transformative resource for the study of the evolution, function, and regulation of the HPPD family in diverse range of species. HTD offers unparalleled opportunities for understanding HPPD's role in evolutionary biology, metabolic engineering, and enzyme regulation by integrating a wide range of omics data. Although the HTD provides a comprehensive platform, there are some potential limitations that could be addressed to enhance its utility and robustness in the future. The current version of the HTD database is deficient in robust analytical workflows. The integration of commonly used bioinformatics tools would facilitate the discovery of evolutionarily conserved domains and pathways specific to certain species. The implementation of regular updates that integrate new data would ensure the currency and relevance of the database to ongoing research. We hope that HTD will serve as the central repository for functional research in HPPD family, thereby facilitating advancements in plant stress response, crop breeding, and herbicide discovery. This work was supported by the National Key Research and Development Program of China (no.: 2023YFD1700500), the National Natural Science Foundation of China (nos.: 22377031 and 22007035), self-determined research funds of CCNU from the colleges' basic research and operation of MOE (no.: CCNU24JCPT023), the Hubei Provincial Department of Education Science and Technology Plan Project (no.: 2024CSA063), the Postdoctoral Fellowship Program (Grade B) of China Postdoctoral Science Foundation (no.: GZB20230250). None declared. G-FY and H-YL contributed to the study conception and design. L-CM, H-YL, X-HY, L-JC, J-HM, and Y-TX contributed to the data collection. FW constructed the online website. L-CM wrote the draft manuscript, and G-FY and H-YL revised the manuscript. All authors read and approved the final manuscript. L-CM and FW contributed equally to this work. HTD database is available online at https://chemyang.ccnu.edu.cn/ccb/database/HTD/. The New Phytologist Foundation remains neutral with regard to jurisdictional claims in maps and in any institutional affiliations.
Geometric deep learning has shown great potential for cortical surface analysis, but its performance often depends on a largescale training set of cortical surfaces, which are traditionally derived from MRI scans through complex and time-consuming preprocessing pipelines. Although deep learning-based surface reconstruction methods have streamlined this process, they still rely on MRI data, limiting the availability of training data. To address this, we propose CortexGen, a geometric generative framework that synthesizes highly realistic cortical surfaces without requiring MRI scans. CortexGen employs geometric variational encoders to map cortical surfaces into a latent space, where latent flow matching models efficiently learn the true data distribution. This enables a two-stage cortical surface synthesis process: first, deforming an icosahedron-discretized sphere into a coarse cortical surface, and second, refining it into a high-resolution surface. Experiments show that CortexGen generates diverse, realistic cortical surfaces with 163,842 vertices in just 1.4 seconds per surface. Using these synthetic surfaces as augmented training data significantly improved learning-based cortical surface parcellation in few-shot settings. Our code and pretrained models are available at https://github.com/ladderlab-xjtu/CortexGen.
The utility of Magnetic Resonance Imaging (MRI) in anomaly detection and disease diagnosis is well recognized. However, the current imaging protocol is often hindered by long scanning durations and a misalignment between the scanning process and the specific requirements of subsequent clinical assessments. While recent studies have actively explored accelerated MRI techniques, the majority have concentrated on improving overall image quality across all voxel locations, overlooking the attention to specific abnormalities that hold clinical significance. To address this discrepancy, we propose a model-unrolled deep-learning method, guided by weakly supervised lesion attention, for accelerated MRI oriented by downstream clinical needs. In particular, we construct a lesion-focused MRI reconstruction model, which incorporates customized learnable regularizations that can be learned efficiently by using only image-level labels to improve potential lesion reconstruction but preserve overall image quality. We then design a dedicated iterative algorithm to solve this task-driven reconstruction model, which is further unfolded as a cascaded deep network for lesion-focused fast imaging. Comprehensive experiments on two public datasets, i.e., fastMRI and Stanford Knee MRI Multi-Task Evaluation (SKM-TEA), demonstrate that our approach, referred to as Lesion-Focused MRI (LF-MRI), surpassed existing accelerated MRI methods by relatively large margins. Remarkably, LF-MRI led to substantial improvements in areas showing pathology. The source code and pretrained models will be publicly available at https://github.com/ladderlab-xjtu/LF-MRI.