While transient perioperative side effects of intravenous anesthetics are often tolerated, the persistent postoperative sequelae resulting from drug accumulation pose a critical threat to patient safety. Etomidate, introduced in the 1970s, remains favored for its minimal hemodynamic impact but is severely limited by sustained adrenal suppression, leading to higher mortality and poorer outcomes in critically ill patients. To address these challenges, we reframed our strategy from solely optimizing receptor specificity to enhancing metabolic efficiency, thereby reducing prolonged postoperative exposure and mitigating sustained adverse effects. Using a deep-learning based molecule optimization algorithm, we identified metabolically favorable lead compounds and synthesized 31 novel imidazole-based etomidate derivatives. Among these, ETO-4 emerged as the most promising candidate, retaining potent anesthetic activity while accelerating metabolic clearance and significantly diminishing adrenal suppression. Plasma cortisol assays confirmed the effect of ETO-4 on adrenal function is greatly reduced. These findings underscore a paradigm shift in anesthetic drug design, demonstrating that prioritizing enhanced metabolic profiles can yield safer, more effective agents that improve postoperative outcomes.
Accurately reconstructing point clouds from a single-view image remains inherently challenging due to the limited input information, which often leads to discrepancies in shape and scale between the reconstructed results and the target object. To tackle this issue, we propose the Consistency Diffusion Model, which, for the first time, jointly explores both 2D and 3D priors within a unified Bayesian framework. Specifically, to ensure precise reconstruction, the effectiveness of various 2D priors is first evaluated, and an associative aggregation mechanism is thus explored to enhance the training supervision. In parallel, we innovate the explicit introduction of 3D priors into the training process, i.e. object-level 3D priors are incorporated as a bound term within the variational Bayesian framework. This term is theoretically shown to tighten the variational bound, thereby enhancing the learning of the shape and scale consistency during the training. Overall, this explored joint learning not only provides robust training guidance but also reinforces reconstruction consistency even under limited visual cues. Extensive experimental evaluations demonstrate that our method consistently achieves state-of-the-art performance, exhibiting exceptional robustness and generalization across both synthetic and real-world datasets.
Text-conditioned diffusion models have revolutionized the field of controllable real image editing, enabling high-fidelity and precise image manipulation. Recent methods target specific editing tasks, using internal representations from reconstruction to ensure consistency. Although effective for single tasks, they fail to balance precision and consistency across diverse image editing tasks. In this work, we propose a novel inference-time real-image editing framework that enables executing multiple editing tasks by tuning editing operators. Our key insight is to treat real image editing as a multi-objective optimization problem, optimizing editing operators for a Pareto optimal solution that balances editing accuracy and consistency at each denoising iteration. Additionally, we design a benchmark for operator-guided real-image editing that covers various local and global editing tasks. Extensive experimental evaluations demonstrate the method’s effectiveness in executing precise edits while preserving image fidelity across all tasks, thereby establishing it as the new state-of-the-art.
Semantic image editing methods employing large-scale diffusion models have made significant strides in precise and controlled image editing with text prompts as guidance. However, these models struggle to handle complex images containing hard-described objects and/or multiple objects. In this work, we introduce a novel inference-time multi-object image editing strategy, Point2pix-Zero, editing a single object with the simple guidance of clicked points and the text of target objects. We employ an interactive methodology, point-discovery, as text-free guidance to identify the semantic information of intended edited objects and generate text prompts automatically. Instead of exploiting internal cross-attention maps of diffusion models as a guide, we inject external attention maps to rectify the visual-and-semantic pairing mismatches in cross-attention maps during the denoising process. Extensive empirical evaluations demonstrate the effectiveness of our proposed inference-time method in ensuring precise editing while maintaining image fidelity. Our method showcases superior performance in single-and multi-object image editing, positioning it as a new state-of-the-art.
Adversarial training has emerged as a leading strategy for enhancing the robustness of machine learning models against adversarial attacks. Its effectiveness often wanes when faced with unseen adversarial examples, resulting in suboptimal robust generalization. To address this issue, we introduce a novel energy-based optimization strategy to improve the robust generalization by incorporating the principles of energy-based models. Our framework models the energy of natural and adversarial examples, where natural samples are assigned to lower energy and adversarial samples to higher energy. During the inference phase, the influence of adversarial perturbation can be alleviated by energy minimization. Theoretically, we show that the proposed energy-based optimization strategy yields a tighter robust-generalization bound through an explicit energy-discrepancy term; this analysis provides an explanatory bound and should not be interpreted as certified robustness. Empirically, a series of evaluations provide evidence for the efficacy of the proposed methodology under the specified threat models and evaluation protocols, showing strong and competitive performance across three extensively utilized datasets. Specifically, EM-AT achieves 77.71% standard-AA robustness on CIFAR-10 and remains highly competitive under comparable lightweight settings. The source codes are available at https://github.com/LitterQ/EM-AT.
While multimodal medical image segmentation improves accuracy via complementary information, real-world constraints often result in incomplete modality inputs, posing a major challenge to robust segmentation. This work addresses the most constrained single-modality setting via a UNet-based distillation framework, which prunes skip connections in the teacher network and adaptively modulates distillation strength to guide compact, informative student network representations. First, entropy-based pruning is applied to skip connections to reduce low-level redundancy and promote semantic abstraction in the teacher network. This enhances the bottleneck and retained skip connection, yielding more informative features for effective distillation. Second, an entropy- and depth-aware temperature schedule is introduced to adaptively control distillation strength across critical semantic routes. Such modulation guides the student to focus on informative signals, enhancing representation under limited capacity. Experimental results on benchmark medical imaging datasets demonstrate that our method outperforms existing single-modality approaches.
Brain tumor segmentation in multi-modal MRIs poses significant challenges when one or more modalities are missing. Recent approaches commonly employ parallel fusion strategies; however, these methods often risk losing crucial shared information across modalities, which can degrade segmentation performance. In this paper, we advocate leveraging sequential information bottleneck fusion to effectively preserve shared information across modalities. From an information-theoretic perspective, sequential fusion not only produces more robust fused representations in missing-data scenarios but also achieves a tighter generalization upper bound compared to parallel fusion approaches. Building on this principle, we propose the Sequential Multi-modal Segmentation Network (SMSN), which integrates an Information-Bottleneck Fusion Module (IBFM). The IBFM sequentially extracts modality-common features while reconstructing modality-specific features through a dedicated feature extraction module. Extensive experiments on the BRATS18 and BRATS20 glioma datasets demonstrate that SMSN consistently outperforms traditional parallel fusion-based baselines, achieving exceptional robustness in diverse missing-modality settings. Furthermore, SMSN exhibits superior cross-domain generalization, as evidenced by its ability to transfer a trained model from BRATS20 to a brain metastasis dataset without fine-tuning. To ensure reproducibility, the code of the SMSN is provided in the supplementary file.
Visual classifiers are expected to generalize under data shifts, target shifts, and their combinations, yet most existing methods focus on domain invariance while failing to address intra-image predictive sufficiency. We investigate the structural hypothesis that each image contains a sample-adaptive oracle intra-image predictive subset sufficient for label prediction, while the remaining patches form non-essential complementary context that may correlate with the label. The theoretical analysis shows that restricting prediction to this oracle subset preserves the Bayes risk achievable by the full-patch representation while admitting a complexity bound that tightens with the oracle-subset size. Based on this view, we propose PatchGen, a text-free module that learns a sample-dependent soft predictive-subset mask as a task-driven proxy for the unobserved oracle subset mask. Specifically, histopathology visualizations suggest that PatchGen assigns higher scores to tumor-consistent regions than to some frequently co-occurring inflammatory context. Extensive experiments on natural and histopathological image benchmarks spanning all three shift settings show that PatchGen improves average performance over matched-backbone baselines in most evaluated configurations, enhances generalization to unknown classes, and remains competitive with vision-language methods without text supervision.
Predicting stock price movements is a longstanding challenge in financial research due to the market’s inherent volatility and the interplay of both quantitative indicators and qualitative sentiment. Recent advances in generative artificial intelligence (AI), particularly large language models (LLMs), have opened new avenues for integrating textual information into predictive models. In this study, we propose a novel attention-based framework that leverages LLM-generated text embeddings as predictive features, combining them with numerical financial indicators to forecast stock prices. Unlike conventional cross-attention approaches, our unified self-attention strategy, which processes concatenated multi-modal features as a single sequence, achieves superior predictive accuracy by enabling richer intra-modal interactions. To further capture the temporal dependencies within financial indicators, we incorporate a long short-term memory (LSTM) module into the framework. Extensive experiments on real-world financial datasets demonstrate that our model could achieve promising performance, highlighting the effectiveness of generative AI-driven textual representations in enhancing financial forecasting. This work underscores the potential of combining structured data with LLM-derived features through attention-based architectures for more robust and interpretable stock price prediction.
It is unknown whether the glucagon-like peptide-1 (GLP-1) receptor agonists have a significant protective effect against acute islet injury. High mobility group box 1 (HMGB1) is a damage-associated molecular pattern (DAMP) molecule released from stressed or injured pancreatic β-cells, which triggers inflammatory responses through toll-like receptor 4 (TLR4) signaling. This study investigated the protective effect and mechanism of liraglutide on acute islet injury induced by low doses of streptozotocin (STZ). The results showed that liraglutide pretreatment preserved the structural integrity of pancreatic islets, improved insulin levels and glucose tolerance, and significantly reduced the incidence of diabetes in STZ-treated mice. Liraglutide was also found to inhibit STZ-induced release of HMGB1 and reduce the expression of TLR4 and inflammatory factors IFN-γ, IL-1β, and CXCL10. Moreover, administration of exogenous HMGB1 or antagonism of the GLP-1 receptor diminished liraglutide’s protective effects. These findings suggest that liraglutide has a strong protective effect on STZ-induced acute islet injury, most likely through the inhibition of HMGB1 release, which provides an experimental basis for the application of liraglutide as a protective agent for acute islet injury.
Uncertainty in medical image segmentation is inherently non-uniform, with boundary regions exhibiting substantially higher ambiguity than interior areas. Conventional training treats all pixels equally, leading to unstable optimization during early epochs when predictions are unreliable. We argue that this instability hinders convergence toward Pareto-optimal solutions and propose a region-wise curriculum strategy that prioritizes learning from certain regions and gradually incorporates uncertain ones, reducing gradient variance. Methodologically, we introduce a Pareto-consistent loss that balances trade-offs between regional uncertainties by adaptively reshaping the loss landscape and constraining convergence dynamics between interior and boundary regions; this guides the model toward Pareto-approximate solutions. To address boundary ambiguity, we further develop a fuzzy labeling mechanism that maintains binary confidence in non-boundary areas while enabling smooth transitions near boundaries, stabilizing gradients, and expanding flat regions in the loss surface. Experiments on brain metastasis and non-metastatic tumor segmentation show consistent improvements across multiple configurations, with our method outperforming traditional crisp-set approaches in all tumor subregions.
Brain tumor segmentation is often based on multiple magnetic resonance imaging (MRI). However, in clinical practice, certain modalities of MRI may be missing, which presents a more difficult scenario. To cope with this challenge, Knowledge Distillation, Domain Adaption, and Shared Latent Space have emerged as commonly promising strategies. However, recent efforts to address the missing modality problem in brain tumor segmentation typically overlook the modality gaps and thus fail to learn important invariant feature representations across different modalities. Such drawback consequently leads to limited performance for missing modality models. To ameliorate these problems, pre-trained models are used in natural visual segmentation tasks to minimize the gaps. However, promising pre-trained models are difficult to obtain in the brain tumor segmentation task due to the lack of sufficient data. Along this line, in this paper, we propose a novel paradigm that aligns latent features of involved modalities to a well-defined distribution anchor as the substitution of the pre-trained model. As a major contribution, we prove that our novel training paradigm ensures a tight evidence lower bound, thus theoretically certifying its effectiveness. Extensive experiments on different backbones validate that the proposed paradigm can enable invariant feature representations and produce models with narrowed modality gaps. Models with our alignment paradigm show their superior performance on both BraTS2018, BraTS2020 and Brain Metastasis datasets.
Generalization remains a significant challenge in visual classification tasks, particularly in handling unknown classes in real-world applications. Existing research focuses on the class discovery paradigm, which tends to favor known classes, and the incremental learning paradigm, which suffers from catastrophic forgetting. Recent approaches such as the L-Reg technique employ logic-based regularization to enhance generalization but are bound by the necessity of fully defined logical formulas, limiting flexibility for unknown classes. This paper introduces PL-Reg, a novel partial-logic regularization term that allows models to reserve space for undefined logic formulas, improving adaptability to unknown classes. Specifically, we formally demonstrate that tasks involving unknown classes can be effectively explained using partial logic. We also prove that methods based on partial logic lead to improved generalization. We validate PL-Reg through extensive experiments on Generalized Category Discovery, Multi-Domain Generalized Category Discovery, and long-tailed Class Incremental Learning tasks, demonstrating consistent performance improvements. Our results highlight the effectiveness of partial logic in tackling challenges related to unknown classes.
The differences among medical imaging modalities, driven by distinct underlying principles, pose significant challenges for generalization in multi-modal medical tasks. Beyond modality gaps, individual variations, such as differences in organ size and metabolic rate, further impede a model's ability to generalize effectively across both modalities and diverse populations. Despite the importance of personalization, existing approaches to multi-modal generalization often neglect individual differences, focusing solely on common anatomical features. This limitation may result in weakened generalization in various medical tasks. In this paper, we unveil that personalization is critical for multi-modal generalization. Specifically, we propose an approach to achieve personalized generalization through approximating the underlying personalized invariant representation X_h across various modalities by leveraging individual-level constraints and a learnable biological prior. We validate the feasibility and benefits of learning a personalized X_h, showing that this representation is highly generalizable and transferable across various multi-modal medical tasks. Extensive experimental results consistently show that the additionally incorporated personalization significantly improves performance and generalization across diverse scenarios, confirming its effectiveness.
3D point cloud analysis has recently garnered significant attention due to its capacity to provide more comprehensive information compared to 2D images. To confront the inherent irregular and unstructured properties of point clouds, recent research efforts have introduced numerous well-designed set abstraction blocks. However, few of them address the issues of information loss and feature mismatch during the sampling process. To address these problems, we have explored the Markov process to revisit point clouds analysis, wherein different-scale point sets are treated as states, and information updating between these point sets is modeled as the probability transition. In the framework of Markov analysis, our encoder can be shown to effectively mitigate information loss in downsampled point sets, while our decoder can accurately recover corresponding features for the upsampled point sets. Furthermore, we introduce a difference-wise attention mechanism to specifically extract discriminative point features, focusing on informative point feature distillation within the states. Extensive experiments demonstrate that our method equipped with Markov process consistently achieves superior performance across a range of tasks including object classification, pose estimation, shape completion, part segmentation, and semantic segmentation. The code is publicly available at https://github.com/ssr0512/Markov-Process-Analysis-on-Point-Cloud.git.
LayerNorm is pivotal in Vision Transformers (ViTs), yet its fine-tuning dynamics under data scarcity and domain shifts remain underexplored. This paper shows that shifts in LayerNorm parameters after fine-tuning (LayerNorm shifts) are indicative of the transitions between source and target domains; its efficacy is contingent upon the degree to which the target training samples accurately represent the target domain, as quantified by our proposed Fine-tuning Shift Ratio (FSR). Building on this, we propose a simple yet effective rescaling mechanism using a scalar λ that is negatively correlated to FSR to align learned LayerNorm shifts with those ideal shifts achieved under fully representative data, combined with a cyclic framework that further enhances the LayerNorm fine-tuning. Extensive experiments across natural and pathological images, in both in-distribution (ID) and out-of-distribution (OOD) settings, and various target training sample regimes validate our framework. Notably, OOD tasks tend to yield lower FSR and higher λ in comparison to ID cases, especially with scarce data, indicating under-represented target training samples. Moreover, ViTFs fine-tuned on pathological data behave more like ID settings, favoring conservative LayerNorm updates. Our findings illuminate the underexplored dynamics of LayerNorm in transfer learning and provide practical strategies for LayerNorm fine-tuning.
Tabular anomaly detection under the one-class classification setting poses a significant challenge, as it involves accurately conceptualizing "normal" derived exclusively from a single category to discern anomalies from normal data variations. Capturing the intrinsic correlation among attributes within normal samples presents one promising method for learning the concept. To do so, the most recent effort relies on a learnable mask strategy with a reconstruction task. However, this wisdom may suffer from the risk of producing uniform masks, i.e., essentially nothing is masked, leading to less effective correlation learning. To address this issue, we presume that attributes related to others in normal samples can be divided into two non-overlapping and correlated subsets, defined as CorrSets, to capture the intrinsic correlation effectively. Accordingly, we introduce an innovative method that disentangles CorrSets from normal tabular data. To our knowledge, this is a pioneering effort to apply the concept of disentanglement for one-class anomaly detection on tabular data. Extensive experiments on 20 tabular datasets show that our method substantially outperforms the state-of-the-art methods and leads to an average performance improvement of 6.1% on AUC-PR and 2.1% on AUC-ROC.
3D open-world classification is a challenging yet essential task in dynamic and unstructured real-world scenarios, requiring robust subsequent knowledge adaptation capabilities. While current approaches predominantly rely on 2D pre-trained models through 3D-to-2D projection, their performance degrades severely under arbitrary object orientations. Unlike these present efforts, this work makes a pioneering exploration of 3D generative models for 3D open-world classification-specifically, leverageing the accumulated prior knowledge from these models to provide anchors for novel categories, while integrating a rotation-invariant feature extractor. This innovative synergy endows our pipeline with the advantages of being training-free and pose-invariant, thus well suited to adapt novel categories in 3D open-world classification. Extensive experiments on benchmark datasets demonstrate the potential of this pipeline, achieving state-of-the-art performance on ModelNet10(double dagger) and McGill(double dagger) with 32.7% and 8.7% overall accuracy improvement, respectively. The code is available in the supplementary materials.
Acne is a prevalent skin disorder causing significant physical and psychological distress. Effective treatment relies on accurate diagnosis and monitoring, but conventional analysis using 2D facial images is limited by issues like self-occlusion and lacks quantitative depth, failing to capture the true extent of the condition. To address this problem, we propose a novel system that integrates deep learning-based acne segmentation with 3D facial reconstruction from a few non-simultaneously captured 2D images. Our methodology first employs a state-of-the-art segmentation model to accurately identify acne lesions from multi-view facial photographs. It then uses a novel framework to reconstruct a detailed 3D facial model, even from unconstrained images, and maps the segmented acne onto this model. Experimental results demonstrate that our system effectively generates accurate 3D face models with integrated acne distributions. The TransUNet model achieved superior segmentation performance with an F1-Score of 0.7765, and our 3D reconstruction method surpassed existing techniques with an average PSNR of 24.36 and SSIM of 0.92. This approach provides clinicians with a comprehensive tool for improved diagnosis, personalized treatment planning, and objective monitoring of skin lesions.
Open compound domain adaptation (OCDA) aims to transfer knowledge from a labeled source domain to a mix of unlabeled homogeneous compound target domains while generalizing to open unseen domains. Existing OCDA methods solve the intradomain gaps by a divide-and-conquer strategy, which decomposes the problem into several individual and parallel domain adaptation (DA) tasks. In this work, starting from the general DA theory, we establish a novel generalization bound for the setting of OCDA. Built upon this, we argue that conventional OCDA approaches may substantially underestimate the inherent variance inside the compound target domains for model generalization, constraining the model’s performance. We subsequently present stochastic compound mixing (SCMix), an augmentation strategy with the primary objective of mitigating the divergence between the source and mixed target distributions. Theoretical analyses are conducted to substantiate the superiority of SCMix, proving that single-target mixing is a subgroup of our method. Extensive experiments show that our method attains a lower empirical risk on OCDA semantic segmentation tasks, thus supporting our theories. In particular, combining the transformer architecture, SCMix achieves a notable performance boost compared to SoTA results.