In manufacturing engineering, process planning transforms engineering drawings into manufacturing process plans by determining appropriate machining operations and their sequences, thereby directly influencing production efficiency and product quality. Traditional process planning relies heavily on expert experience, resulting in low efficiency and limited knowledge standardization. Although computer-aided process planning (CAPP) methods have been developed, they typically depend on handcrafted rules or structured inputs, limiting their ability to handle complex semantics, diverse drawing styles, and cross-task generalization. Vision-language models (VLMs) provide a promising paradigm for drawing understanding. However, general-purpose VLMs often lack specialized manufacturing knowledge and struggle to align high-level representations with pixel-level details in high-resolution drawings. This leads to the loss of critical information and degrades the quality of generated process text. We propose ProcessMM, a domain-specialized multimodal model for manufacturing process planning. ProcessMM adopts hierarchical vision-text alignment with dual-channel features. By incorporating OCR text into visual representations via an OCR-guided attention (OGA) module, ProcessMM enhances fine-grained detail perception. We also develop a sentence-level attention attribution method that provides process planners with visual evidence for generated process text. Evaluations on three engineering datasets collected from real-world manufacturing processes show that ProcessMM consistently outperforms dominant multimodal large language models (MLLMs) in manufacturing process-text generation, demonstrating effective and efficient automated process planning from engineering drawings.
Deep neural networks are highly vulnerable to adversarial examples, i.e.,small perturbations that can significantly degrade model performance. While adversarial training has become the primary defense strategy, most studies focus on balanced datasets, overlooking the challenges posed by real-world long-tail data. Motivated by the fact that perturbations in adversarial examples inherently alter the training distribution, we theoretically investigate their impact. We first revisit adversarial training for long-tail data and identify two key limitations: (i) a skewed training objective caused by class imbalance, and (ii) unstable evolution of adversarial distributions. Furthermore, we show that perturbations can simultaneously address both adversarial vulnerability and class imbalance. Based on these insights, we propose Rebalanced Adversarial Intensity for Long-Tailed Data (RAIL), a plug-and-play framework that adaptively adjusts perturbations during adversarial training. Extensive experiments demonstrate that RAIL consistently enhances adversarial robustness and class-balance on long-tailed datasets.
Neural Predictors (NPs) are pivotal for efficient Neural Architecture Search (NAS) but suffer from limited accuracy due to the high cost of acquiring labeled training data. While active learning (AL) can mitigate this by prioritizing informative samples, existing methods suffer from selection bias under imbalanced data distributions, often sacrificing representativeness for diversity. To address this, we redefine the sample selection mechanism in AL and propose a Distribution-aware Active Learning framework for Neural Predictor (called DARE). The goal is to select samples that not only ensure diversity but also exhibit a high degree of generalizability, making them more representative of the underlying data distribution. Our approach first extracts architecture representations via a graph-based encoder enhanced with a consistency-driven objective. Then, a two-stage selection strategy identifies both globally diverse and locally reliable samples through progressive representation learning and refinement. For non-uniform data distributions, we further introduce an adaptive mechanism that anchors sampling to key regions with high similarity density, avoiding performance degradation caused by outliers. Extensive experiments on NAS-Bench-101, NAS-Bench-201, NAS-Bench-301, and DARTS demonstrate that DARE achieves state-of-the-art prediction performance with extremely scarce labeled data. For instance, on NAS-Bench-101, it outperforms the second-best baseline by 4.29% in Kendall’s Tau with only 424 samples. Furthermore, DARE efficiently identifies optimal architectures using as few as 100 queries on DARTS.
Lead optimization, the systematic refinement of therapeutic compounds through iterative structural modification, faces a dual challenge in modern drug discovery: navigating astronomically vast molecular design spaces while balancing conflicting demands on potency, pharmacokinetics, and safety. We present MASCOT (Multi-Agent SearCh for molecular OpTimization), a role-specialized multi-agent framework for molecular optimization. Integrated with a chemically constrained graph-editing search, MASCOT coordinates three specialized agents: a trade-off agent that reprioritizes competing objectives, a strategy agent that adapts how molecular edits are proposed, and a reflection agent that distills lessons from previous decisions. Computational experiments showed that MASCOT achieved the best performance over competing methods on six benchmark settings. On the SARS-CoV-2 main protease task, its mean docking-score improvement was 3.6 times that of the strongest baseline. Applied to the clinically used anesthetic remimazolam (RM), MASCOT prioritized RM-1, which showed a shorter liver microsomal half-life, higher brain exposure, and a larger therapeutic index than RM. Subsequent derivative design yielded RM-7. Extensive animal studies established RM-7 as a rapid-recovery intravenous anesthetic candidate with greater potency, faster functional recovery, a wider safety margin, and preserved flumazenil reversibility. These results demonstrate that multi-agent coordination can link adaptive molecular search to medicinal chemistry and experimental pharmacology.
Biomedical event trigger detection remains a foundational yet challenging task due to the pronounced long-tail distribution of trigger types, where rare event classes suffer from severe underrepresentation. To address this class imbalance, we propose DCA-Net, a novel framework integrating semantic Markov Chain Monte Carlo (MCMC) sampling with dynamic context association (DCA) mechanisms. Our approach first constructs a trigger transition graph to perform semantic MCMC sampling on neighboring triggers, enabling large language models (LLMs) to generate class-discriminative synthetic instances through semantic guidance. The framework further incorporates a gated attention-based DCA module that dynamically captures multi-granularity trigger-context dependencies using adaptive receptive fields. Complemented by a class-weighted focal loss emphasizing hard-to-learn rare triggers, DCA-Net achieves state-of-the-art performance across MLEE, BioNLP2011-Genia, and BioNLP2013-Genia benchmarks with F1-scores of 90.63 %, 87.52 %, and 75.39 % respectively, outperforming existing methods by up to 8.6 %. Ablation studies systematically validate the synergistic benefits of our graph-guided augmentation and adaptive context modeling components.
reasoning is a type of thinking that involves inducing rules (or knowledge) from examples and generalizing them to new instances. Developing neural networks with reasoning ability is an important step towards human-like intelligence. Although neural networks have achieved impressive performances in multiple tasks such as image recognition, the essence of reasoning, i.e. the capacity to learn and apply abstract rules, remains unsolved. In this work, we propose SIMAR, a novel Self-Inference Mechanism for Abstract Reasoning. In particular, the self-inference mechanism regularizes the rule representations for abstract reasoning to be more robust, thereby enhancing the generalization capacity. This mechanism mimics the "hypothesis testing" of humans' cognitive process, where we internally generate proposals for the abstract rules, and then utilize the rules in different instances to check whether they can explain the instances. Both theoretical analysis and experimental results show that the self-inference mechanism can enforce our model to learn robust representations for abstract rules. Based on these representations, our method outperforms the state-of-theart models across a majority of the reasoning tasks, encompassing abstract reasoning in visual form and mathematical reasoning in textual form. Notably, SIMAR exhibits remarkable superiority in generalization capacity. The key idea of self-inference is general and useful for learning rule representations, providing new perspectives for future research on abstract reasoning.
Adversarial training breaks down in long-tailed settings, exhibiting severe robustness degradation on worst-performing (often tail) classes. We identify a key cause of this failure as a posterior mismatch: coarse-grained absolute labels collapse class posteriors into point estimates, leading to biased class-frequency estimation and an enlarged robust generalization gap, which ultimately amplifies worst-class vulnerability. To address this issue, we propose Posterior-driven Adversarial Training (PAT), which learns a posterior surrogate to provide fine-grained probabilistic supervision for adversarial training, and integrates weight perturbations to encourage a flatter loss landscape. Our theory shows that accurate posterior approximation simultaneously tightens class-frequency estimation error and robust generalization bounds, while a flat weight loss landscape stabilizes sensitivity to posterior approximation errors. Extensive experiments on long-tailed benchmarks confirm that PAT consistently improves robustness, with especially large gains on worst-class.
Loose parts in nuclear power plant reactor primary circuits pose significant structural risks to pressure vessels, which may result in severe operational consequences. Vibration monitoring systems like Loose Parts Monitoring Systems (LPMS) primarily capture normal operational data, with fault-related or abnormal signals being exceptionally scarce. Neural network-based fault recognition methods are prone to overfitting and limited generalizability under condition of limited samples, adversely affecting vibration signal classification accuracy. To address these issues, this paper proposes an intelligent diagnosis method for loose parts faults, combining feature enhancement, data augmentation, and model enhancement. The method employs Short-Time Fourier Transform (STFT) and Wavelet Transform (WT) to transform one-dimensional time-series signals into twodimensional representations. A self-attention deep Variational Autoencoder (SA-DVAE) is developed by integrating self-attention mechanisms into the standard VAE framework, and the Frechet Inception Distance (FID) value is adopted to assess and iteratively refine the quality of augmented samples. A two-stream deep convolutional neural network (TS-DCNN) is subsequently constructed to process STFT- and WT-transformed twodimensional data through separate streams using combined original-augmented datasets, demonstrating enhanced performance relative to conventional diagnosis models. Through experimental validation with actual loose parts fault data and the analysis of ablation, comparative and lightweight model experiments, the proposed method is verified to achieve superior classification accuracy. Finally, the relationship between classification accuracy and the quantity of generated samples is systematically evaluated. The proposed method overcomes data scarcity limitations and achieves robust fault diagnosis where conventional AI methods typically falter.
Efficient identification of high-value molecules under limited experimental budgets remains a central challenge in automated chemical design. In Level 4 automation settings, where chemists define candidate spaces and machine learning models guide experimental selection, pool-based active learning has emerged as a practical framework for prioritizing compounds. However, conventional approaches primarily optimize predictive accuracy of surrogate structure-function models, which may not directly align with the objective of maximizing search efficiency toward identifying the molecule with the most favorable property value within a predefined candidate pool. We propose a policy-based active learning framework that reformulates molecular pool selection as a sequential decision-making problem. The iterative selection process is modeled as a Markov decision process, and a policy network is trained to optimize cumulative search performance under constrained evaluation budgets. Property prediction models are incorporated as contextual signals rather than optimization targets, enabling direct optimization of search efficiency. We construct 1409 molecular identification tasks derived from MoleculeNet and ChEMBL. In addition, six literature-curated in vivo tasks are constructed to assess performance. Across both benchmark and in vivo settings, the framework demonstrates improved efficiency in identifying optimal molecules within limited evaluation cycles. Case studies further illustrate the strengths and limitations of the framework. These results highlight the potential of policy-driven active learning to enhance molecular identification efficiency in predefined candidate pools, offering a generalizable strategy for budget-constrained chemical discovery.
Generating molecules with desired chemical properties is a crucial and promising area of research in drug discovery, as it has the potential to accelerate the identification of novel therapeutic compounds. Recent developments in diffusion models have showcased their remarkable generative capabilities, effectively handling continuous data modalities such as images and audio. However, when it comes to generating discrete data, particularly molecular representations like SMILES strings and molecular graphs, these models encounter significant challenges, especially in few-shot learning scenarios where only a limited number of samples are available. In this paper, we explore the potential of diffusion models for generating continuous representations of molecules-molecular images. Specifically, we propose ProtoDiff, a diffusion-based method that incorporates few-shot learning for molecular image generation. We frame molecular image generation as a few-shot controllable generation problem that extracts prototypes from a limited set of molecules to guide the generation process and introduces a novel sparsity regularization in the objective function of diffusion to emphasize the meaningful pixels of molecules, i.e., the limited pixels of the chemical bonds. We train and evaluate ProtoDiff on the ChEMBL dataset, achieving new state-of-the-art results on the majority of molecular generation tasks.
Molecular property prediction is a critical task in computational chemistry and drug discovery. While deep learning has advanced this field, the increasing complexity of models contrasts with the scarcity of labeled data, leading to severe overfitting and limited generalization. In this paper, we propose TasProp, a task-specific pre-training strategy for molecular property prediction, particularly for the scenarios with small labeled datasets. To learn a robust molecular representation, TasProp first projects both labeled and unlabeled data into a unified latent space. Then, we introduce a task-specific contrastive loss that aligns closely with the final prediction task and apply it to the labeled data. This contrastive loss encourages the model to learn more cohesive and distinguishable molecular representations corresponding to property categories, which in turn, enhances the model's performance on downstream property prediction tasks. Additionally, we propose a novel data augmentation method, accompanied by a theoretical analysis, to mitigate the challenge of labeled data scarcity. With the task-specific pre-training and augmented data, TasProp outperforms the state-of-the-art methods on many molecular property prediction tasks, including three publicly available datasets and two curated datasets related to anesthesiology. Furthermore, we provide an interactive web resource to facilitate model exploration and application, allowing users to easily predict the properties of input molecules online.
Facing the urgent demand for process design automation in the context of intelligent manufacturing, this study identifies significant limitations in current research approaches.Existing vision-based process generation models tend to overemphasize image feature extraction while generally neglecting critical textual annotations in drawings (such as dimensional tolerances and material specifications), leading to insufficient multimodal information coordination.The underlying issues stem from the inadequate modeling of the multimodal characteristics (visual-textual correlations) inherent in engineering drawings by existing methods, and the inability of traditional fusion strategies to achieve progressive cross-modal alignment, thereby constraining the accuracy and logical consistency of generated process descriptions.To address these shortcomings, this study proposes a generative process design algorithm based on multimodal information fusion.By deeply integrating OCR-extracted textual features with visual features, the algorithm effectively addresses key challenges such as incomplete extraction of technical parameters and inaccurate semantic representation from process drawings.Our algorithm demonstrates superior performance over traditional vision-only models on metrics including BLEU and ROUGE, without significantly increasing computational overhead.It provides reliable technical support for the intelligent transformation of process design and holds significant practical value for advancing the digital transformation of the manufacturing industry.
Spatial Transcriptomics (ST) reveals gene expression patterns within the context of tissue microenvironment by in situ measuring gene expression at native spatial locations, enabling the deciphering of spatial domains. However, existing methods often simplify cellular dependencies to pairwise k-nearest neighbor graphs, overlooking the contextual and higher-order dependencies of cellular networks. To address this issue, we propose stHGNN, a dual-view hypergraph enhanced framework for identifying spatial domains in ST data. stHGNN captures contextual and higher-order cellular dependencies using hypergraphs, adopts hypergraph representation learning from both spatial and gene expression views, and then facilitates cross-view knowledge transfer through a cross-attention mechanism and self-correlation reorganization. Additionally, a multi-view self-supervised clustering strategy is adopted to learn clustering-friendly representations, while hypergraph smoothing is applied to enhance consistency with biological priors. Finally, a zero-inflated negative binomial reconstruction is adopted to handle the inherent over-dispersion and dropouts in ST data. Comprehensive experiments and downstream analysis demonstrate the effectiveness of our stHGNN.
Large language models (LLMs) have revolutionized machine learning with their few-shot learning and reasoning capabilities, demonstrating impressive results in fields like natural language processing and computer vision. However, when applied to the domains of biology and chemistry, current LLMs face substantial limitations, particularly in capturing the nuanced relationships between the molecular structure and pharmacochemical properties. This challenge has constrained the application of few-shot learning for small-molecule generation and optimization in drug discovery. Here, we introduce DrugLLM, a novel LLM tailored specifically for molecular optimization. DrugLLM leverages Functional Group Tokenization (FGT), which effectively tokenizes molecules for LLM learning, achieving over 53% token compression compared to SMILES. Besides, we propose a new pre-training strategy that enables DrugLLM to iteratively predict and modify molecular structures based on a few prior modifications, aligning each modification toward optimizing a specified pharmacological property. In multiple computational experiments, DrugLLM achieved state-of-the-art performance in few-shot molecular generation, surpassing all the mainstream LLMs including GPT-4. Furthermore, by applying DrugLLM to optimize HCN2 inhibitors, two bioactive compounds were obtained and successfully validated through wet-lab experiments. These results highlight the robust potential of DrugLLM in accelerating the optimization of molecules and AI-driven drug discovery.
Structure-based drug design aims to generate molecules that fill the cavity of the protein pocket with a high binding affinity. Many contemporary studies employ sequential generative models. Their standard training method is to sequentialize molecular graphs into ordered sequences and then maximize the likelihood of the resulting sequences. However, the exact likelihood is computationally intractable, which involves a sum over all possible sequential orders. Molecular graphs lack an inherent order and the number of orders is factorial in the graph size. To avoid the intractable full space of factorially-many orders, existing works pre-define a fixed node ordering scheme such as depth-first search to sequentialize the 3D molecular graphs. In these cases, the training objectives are loose lower bounds of the exact likelihoods which are suboptimal for generation. To address the challenges, we propose a unified generative framework named MolEM to learn the 3D molecular graphs and corresponding sequential orders jointly. We derive a tight lower bound of the likelihood and maximize it via variational expectation-maximization algorithm, opening a new line of research in learning-based ordering schemes for 3D molecular graph generation. Besides, we first incorporate the molecular docking method QuickVina 2 to manipulate the binding poses, leading to accurate and flexible ligand conformations. Experimental results demonstrate that MolEM significantly outperforms baseline models in generating molecules with high binding affinities and realistic structures. Our approach efficiently approximates the true marginal graph likelihood and identifies reasonable orderings for 3D molecular graphs, aligning well with relevant chemical priors.
Pharmacophores are abstractions of essential chemical interaction patterns, holding an irreplaceable position in drug discovery. Despite the availability of many pharmacophore tools, the adoption of deep learning for pharmacophore-guided drug discovery remains relatively rare. We herein propose a knowledge-guided diffusion framework for 'on-the-fly' 3D ligand-pharmacophore mapping, named DiffPhore. It leverages ligand-pharmacophore matching knowledge to guide ligand conformation generation, meanwhile utilizing calibrated sampling to mitigate the exposure bias of the iterative conformation search process. By training on two self-established datasets of 3D ligand-pharmacophore pairs, DiffPhore achieves state-of-the-art performance in predicting ligand binding conformations, surpassing traditional pharmacophore tools and several advanced docking methods. It also manifests superior virtual screening power for lead discovery and target fishing. Using DiffPhore, we successfully identify structurally distinct inhibitors for human glutaminyl cyclases, and their binding modes are further validated through co-crystallographic analysis. We believe this work will advance the AI-enabled pharmacophore-guided drug discovery techniques.
Deep generative models, such as diffusion models, have shown promising progress in image generation and audio generation via simplified continuity assumptions. However, the development of generative modeling techniques for generating multi-modal data, such as parametric CAD sequences, still lags behind due to the challenges in addressing long-range constraints and parameter sensitivity. In this work, we propose a novel framework for quantitatively constrained CAD generation, termed Target-Guided Bayesian Flow Network (TGBFN). For the first time, TGBFN handles the multi-modality of CAD sequences (i.e., discrete commands and continuous parameters) in a unified continuous and differentiable parameter space rather than in the discrete data space. In addition, TGBFN penetrates the parameter update kernel and introduces a guided Bayesian flow to control the CAD properties. To evaluate TGBFN, we construct a new dataset for quantitatively constrained CAD generation. Extensive comparisons across single-condition and multi-condition constrained generation tasks demonstrate that TGBFN achieves state-of-the-art performance in generating high-fidelity, condition-aware CAD sequences. The code is available at https://github.com/scu-zwh/TGBFN.
The pursuit of optimal neural network architectures is foundational to the progression of Neural Architecture Search (NAS). However, the existing NAS methods suffer from the following problem using traditional search strategies, i.e., when facing a large and complex search space, it is difficult to mine more effective architectures within a reasonable time, resulting in inferior search results. This research introduces the Generative Pre-trained Transformer NAS (GPT-NAS), an innovative approach designed to overcome the limitations which are inherent in traditional NAS strategies. This approach improves search efficiency and obtains better architectures by integrating GPT model into the search process. Specifically, we design a reconstruction strategy that utilizes the trained GPT to reorganize the architectures obtained from the search. In addition, to equip the GPT model with the design capabilities of neural architecture, we propose the use of the GPT model for training on a neural architecture dataset. For each architecture, the structural information of its previous layers is utilized to predict the next layer of structure, iteratively traversing the entire architecture. In this way, the GPT model can efficiently learn the key features required for neural architectures. Extensive experimental validation shows that our GPT-NAS approach beats both manually constructed neural architectures and automatically generated architectures by NAS. In addition, we validate the superiority of introducing the GPT model in several ways, and find that the accuracy of the neural architecture on the image dataset obtained from the search after introducing the GPT model is improved by up to about 9%.
Multifunctional therapeutics have emerged as a solution to the constraints imposed by drugs with singular or insufficient therapeutic effects. The primary challenge is to integrate diverse pharmacophores within a single-molecule framework. To address this, we introduced DeepSA, a novel edit-based generative framework that utilizes deep simulated annealing for the modification of articaine, a well-known local anesthetic. DeepSA integrates deep neural networks into metaheuristics, effectively constraining molecular space during compound generation. This framework employs a sophisticated objective function that accounts for scaffold preservation, anti-inflammatory properties, and covalent constraints. Through a sequence of local editing to navigate the molecular space, DeepSA successfully identified AT-17, a derivative exhibiting potent analgesic properties and significant anti-inflammatory activity in various animal models. Mechanistic insights into AT-17 revealed its dual mode of action: selective inhibition of NaV1.7 and 1.8 channels, contributing to its prolonged local anesthetic effects, and suppression of inflammatory mediators via modulation of the NLRP3 inflammasome pathway. These findings not only highlight the efficacy of AT-17 as a multifunctional drug candidate but also highlight the potential of DeepSA in facilitating AI-enhanced drug discovery, particularly within stringent chemical constraints.