Diffusion large language models (DLLMs) have emerged as an alternative to autoregressive (AR) decoding with appealing efficiency and modeling properties, yet their implications for agentic multi-step decision making remain underexplored. We ask a concrete question: when the generation paradigm is changed but the agent framework and supervision are held fixed, do diffusion backbones induce systematically different planning and tool-use behaviors, and do these differences translate into end-to-end efficiency gains? We study this in a controlled setting by instantiating DLLM and AR backbones within the same agent workflow (DeepDiver) and performing matched agent-oriented fine-tuning on the same trajectory data, yielding diffusion-backed DLLM Agents and directly comparable AR agents. Across benchmarks and case studies, we find that, at comparable accuracy, DLLM Agents are on average over 30
Although interpretable deep learning achieved significant progress in recent years, their data-driven nature often leads to interpretation results that are inconsistent with established polarimetric scattering priors when applied to polarimetric synthetic aperture radar (PolSAR) imagery analysis. Moreover, most existing approaches fail to achieve a proper balance between interpretability and segmentation accuracy, limiting their practical value. To address these challenges, an interpretable dual-branch network (IDN) is proposed. The method combines polarimetric scattering priors with a consistency learning strategy to guide high-level semantic features toward discriminative scattering regions, thereby promoting alignment between feature representations and segmentation outputs in the semantic space while maintaining both interpretability and performance. The proposed IDN consists of three components. First, prior knowledge of polarimetric scattering is integrated with a class activation mapping, which improves the localization capability of the class activation maps by leveraging the distinctive polarimetric scattering characteristics of the targets, while also suppressing background interference. Second, an interpretable segmentation consistency learning module is designed to enhance spatial consistency between extracted features and actual object regions. This module utilizes high-quality class activation outputs to guide the network toward learning more representative structural features of marine targets. Third, to address the challenge of segmenting densely distributed and irregularly shaped objects in complex environments, a refined class activation and edge-guided network is designed. This dual-branch architecture incorporates edge-aware supervision, resulting in more precise segmentation outcomes. The proposed IDN is evaluated using GF-3 PolSAR imagery, with a typical marine aquaculture raft dataset employed as a case study to demonstrate the method's effectiveness. To enhance transparency and reproducibility, the implementation of the proposed method has been released as open-source on GitHub.
Recently, personalized influence maximization, which generalizes classic influence maximization by wisely allocating personalized discounts to users, showed its superiority in achieving cost-effective marketing plans over traditional influence maximization techniques. However, under continuous cost reaction scenarios, it is hard to obtain the approximation guarantee. To this end, we first propose a personalized lattice influence maximization (PLIM) problem, which can be approximately guaranteed through a greedy algorithm from an innovative perspective, i.e., match item selection with a knapsack constraint (MIS-K). Furthermore, to address the practical challenges associated with accurately learning user purchase probabilities, we extend the proposed PLIM framework in the form of robust optimization. We develop a robust algorithm that achieves a solution-dependent approximation guarantee and analyze the computational hardness of approximating RPLIM. By introducing an innovative Re-arrangement Sampling technique, we demonstrate that sampling can effectively mitigate uncertainty in purchase probabilities, enabling a strong robust ratio of (1-epsilon)(1-1/epsilon) with high probability. Experimental results on multiple benchmark datasets confirm the effectiveness and efficiency of our proposed methods.
Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain. We introduce RetailBench, a data-grounded simulation benchmark for evaluating tool-using LLM agents in single-store supermarket operation. RetailBench models retail management as a partially observable decision process and is designed to support thousand-day-scale simulations. In this environment, agents must manage pricing, replenishment, supplier selection, shelf assortment, inventory aging, customer feedback, external events, and cash-flow constraints. We evaluate seven contemporary LLMs under representative agent frameworks over a 180-day evaluation horizon and compare them with a privileged oracle policy. Results show substantial variation across models: only a small subset survives the full evaluation horizon, and even the strongest LLM runs remain substantially behind the oracle policy in final net worth and sales outcomes. Behavioral analysis attributes these gaps to incomplete evidence acquisition, surface-level decision making, and the lack of a consistent long-horizon policy. RetailBench provides a controlled testbed for studying reliable autonomy in economically grounded long-horizon decision-making.
Compressed sensing formulations that target an $\ell _{0}$ -norm objective are inherently nonconvex and discontinuous. In this article, a global optimization problem with a power-mean function is first formulated for compressed sensing. To mitigate the numerical instability of minimizing the power-mean function with a large negative exponent, the problem is reformulated as a sequential majorization-minimization (MM) problem with iteratively reweighted convex surrogate functions at different anchor points. To eliminate the dependency of the solution quality on anchor points, multiple neurodynamic optimization models are employed to seek global optimal solutions collaboratively through repeated reinitialization using a particle swarm optimization rule. Extensive experiments demonstrate that the proposed method achieves superior performance compared to 14 state-of-the-art algorithms in terms of signal sparsity and reconstruction accuracy.
Reinforcement learning has shown significant potential for dynamic Heating, Ventilation, and Air Conditioning control in buildings. However, existing research predominantly focuses on online or off-policy methods that require active environment interactions, leaving the practical utility of purely offline historical data largely unexplored. This research gap hinders the real-world deployment of reinforcement-learning-based controllers despite the abundance of data in Building Management Systems. To address this, this paper provides a comprehensive evaluation of state-of-the-art offline reinforcement learning algorithms, analyzed through the dual lenses of architectural design and dataset characteristics. We propose an end-to-end framework that integrates sequential observation modeling to resolve the partially observable Markov decision process inherent in building dynamics. Crucially, we investigate the quantitative impact of dataset quality — measured by a novel Regret Ratio (δτ) — and quantity on control efficacy. Our findings reveal a counter-intuitive yet vital insight: datasets with a certain level of sub-optimality and diversity are more effective for training robust controllers than expert-only data, as they provide essential transitions for “trajectory stitching”. Experimental results demonstrate that the proposed offline controller can reduce indoor temperature violation ratios by up to 28.5% and achieve energy savings of 12.1% compared to baseline methods. Furthermore, the framework exhibits superior data efficiency, requiring an order of magnitude less data than traditional off-policy approaches. This study establishes both a technical framework and a theoretical benchmark for deploying data-driven Heating, Ventilation, and Air Conditioning control in real-world buildings.
A theoretical analysis of CRC-aided successive cancellation list (CA-SCL) decoding for polar codes remains an open problem, despite its widespread practical adoption. While low-density parity-check (LDPC) codes benefit from mature analytical tools, such as density evolution (DE), for predicting the performance of belief-propagation (BP) decoding, similar techniques are not directly applicable to CA-SCL decoding. This limitation stems from the complex path-pruning mechanism inherent in CA-SCL decoding. In this paper, we propose an analytical framework based on a novel path-survival model that captures the evolution of the correct path's rank during decoding. The proposed framework enables efficient prediction of CA-SCL decoding performance without requiring exhaustive list-specific Monte Carlo simulations. Extensive numerical evaluations demonstrate its effectiveness across a wide range of code lengths, code rates, list sizes, and channel models.
This review focuses on deep reinforcement learning-based protein-ligand docking methods.Comprehensively summarized recent advances, covering existing methods, model architectures, training strategies, and feature extraction technologies. The precise prediction methods such as spatial pose prediction, alternative direct prediction, and affinity prediction are detailed, along with the application of deep learning approaches like 3D Convolutional Neural Networks (3D-CNN), Spatial Graph Convolutional Neural Networks (SG-CNN), and the Transformer Docking Algorithm. Discussed scoring function optimization, traditional limitations, and advances in machine-learning-based scoring function development. The article proposes future directions, including multimodal data fusion and novel architectures, enhancing protein-ligand docking efficiency and drug discovery success.
This article addresses a collective heterogeneous multiagent pursuit-evasion (MPE) game problem where pursuers cooperatively capture escaping evaders. The analytical challenge of the present design lies in solving the associated coupled Hamilton-Jacobi-Isaacs (HJI) equations induced by the additional interacting roles in the MPE game while ensuring the achievement of the Nash equilibrium. To tackle this issue, a gaming framework is accordingly proposed to solve the coupled HJI equations. Sufficient conditions are derived to guarantee both the capturability and Nash equilibrium of the proposed collective MPE gaming scheme. Finally, numerical simulations are conducted to verify the effectiveness of the present MPE gaming strategy.
A hybrid Koopman deep learning algorithm is developed to predict the final synchronization of networked nonlinear dynamics with different topologies merely using neighboring state information. This algorithm introduces a nonlinear encoder as an observable function that maps the nonlinear state into a high-dimensional Hilbert space. By this means, a networked linear model is established to predict the future state of multiple transformed linear systems in the lifted space. Meanwhile, a nonlinear decoder is constructed, as the inverse of the lifting function, to retrieve the original nonlinear states. The virtue of the present algorithm lies in distilling and merging the linear features of multiple different topologies solely from the individual and/or neighboring state series. Therefore, the final synchronization states are calculated within the encoded linear space and subsequently decoded to recover the synchronization of the original nonlinear systems. Compared to most existing relevant algorithms that could only predict consensus values for linear networks, the present method could predict the final synchronization state of networked nonlinear dynamics with varying backbones. Sufficient conditions are derived to guarantee the prediction capability of the distributed final synchronization prediction (DFSP). Extensive numerical simulations verify its effectiveness.
Discrete-time neurodynamic approaches (also called recurrent neural networks (RNNs)) are easily implemented on software and simulated on digital circuits. First, this paper proposes two modified discrete-time RNNs for quickly dealing with constrained $l_{1}$-norm minimization problems. Next, the two modified discrete-time RNNs are proven to be globally convergent to an optimal solution under a large step size. Finally, we apply the obtained results for image recovery. Two convergent discrete-time RNN based algorithms for non-blind image restoration are presented. Due to having a low complexity, the two discrete-time RNNs are more computationally efficient than the existing discrete-time RNN for image restoration. Computed results with application examples show that the two discrete-time RNN-based algorithms are indeed superior to the existing discrete-time RNN-based algorithms with regards to computation time.
Hybrid AC/DC microgrids, which integrate the respective benefits of AC and DC topologies, have attracted increasing attention as an effective system configuration. Nevertheless, the widespread use of nonlinear loads in hybrid microgrid (HM) may severely deteriorate power quality. To overcome this problem, this paper presents a model predictive control strategy for the interlinking converter (ILC) connecting the AC and DC microgrids, in which virtual harmonic impedance is incorporated. According to the mechanism of harmonic distortion, a harmonic-sequence-based virtual impedance is introduced into the inverter model to establish a virtualized control framework. Based on this framework, predictive control is applied to govern the harmonic voltage compensation process of the ILC. Under the proposed method, the ILC not only mitigates voltage distortion at the point of common coupling (PCC) by identifying and compensating harmonic voltages caused by nonlinear loads, but also realizes active power sharing between the DC distributed generators (DG) and the AC DG. Simulation studies demonstrate that the proposed approach achieves effective harmonic compensation.
The generation process of diffusion models is often complex and slow due to the numerous iterative steps and the high dimensionality of token sequences involved. To address the substantial inference latency and computational overhead associated with diffusion models, several token-reduction methods have emerged. However, existing token reduction methods lack effective evaluation criteria and practical architectural design, hindering the screening of important tokens. This paper introduces a cross-modal token reduction (CTR) method that preserves and optimizes important tokens to accelerate diffusion model inference and reduce computational burden. The CTR method employs token-level cross-modal contrastive learning (TCCL) to align image tokens with the global text condition in a shared semantic space. A semantically aware importance metric (SAIM) then quantifies the contribution of tokens during generation. The paper further presents a spatial-semantic token reduction (SSTR) method that combines spatial and semantic information to prune redundant tokens early and merge less important tokens later, thereby significantly reducing computational load while maintaining generation quality. Importantly, CTR requires no retraining or fine-tuning of the diffusion model backbone network. Only lightweight projection layers are trained offline and then used as plug-andplay modules during inference. Experiments on the COCO30 K dataset demonstrate that CTR achieves a 1.89-fold speedup on Stable Diffusion v1.5 and a 1.60-fold speedup on Stable Diffusion v2.1, with the Fréchet inception distance (FID) decreasing by 1.08 and 2.14.
Matrix-variable triconvex optimization is a significant generalization of vector-variable triconvex or biconvex optimization and has been found to have popular applications. To reduce computation time and storage requirements, this paper presents a matrix-form iterative method for quickly solving matrix-variable constrained triconvex optimization problems. The proposed method is based on a matrix-form alternating projection iteration scheme in the form of matrix state spaces, where an efficient line search strategy is adopted by exploiting the optimality conditions of the problem for a larger step length. Compared with the existing vector-form alternating projection gradient method, the proposed method reduces storage requirements and computational cost, and thus is more computationally efficient. Each sequence generated by the proposed method is guaranteed to be globally convergent to a partial optimum under mild conditions. Finally, the proposed method is effectively applied to blind image deblurring problems. Computed results show that the proposed algorithm is superior to related iterative algorithms in terms of computation time and solution quality.
Distributed nonsmooth nonconvex optimization is prevalent in practical applications. However, the inherent nonsmoothness and nonconvexity of such problems pose significant challenges to the development of efficient optimization approaches. This article proposes a multiagent system with finite-time consensus for solving this problem. A smooth approximation technique is leveraged to handle the nonsmoothness in the problem, and a state-dependent gain function is incorporated into the proposed approach to handle the nonconvexity in constraints. The states of the multiagent system remain within their local feasible regions and reach consensus in a finite time. In addition, the states are proven to be convergent to the critical-point set of the problem under consideration. Furthermore, the states are proven to be convergent to a globally optimal solution, under the nonsmooth Polyak-& Lstrok;ojasiewicz condition or other generalized convexity conditions. The simulation results are elaborated to substantiate the effectiveness and viability of the proposed approach.
In this article, we propose a constrained optimization approach to portfolio selection by maximizing nine risk-adjusted return metrics, subject to second-order stochastic dominance (SSD) constraints. The SSD constraints ensure that the portfolio returns are no less than an amplified proportion of returns from a benchmark in the sense of SSD. Because the number of SSD constraints is extremely large, the resulting constrained optimization problems are computationally challenging. To reduce computational complexity, we develop an efficient algorithm to solve the problems iteratively by incrementally adding SSD constraints. We experimentally demonstrate the superiority of the proposed approaches to several baselines in terms of out-of-sample performance criteria based on financial data from major world stock markets.
Effectively solving high-dimensional and multimodal optimization problems remains a critical challenge in com putational intelligence due to premature convergence and the difficulty of balancing global exploration with local exploitation. Inspired by the cryptobiosis survival strategy of tardigrades, this paper proposes the Tardigrade Optimization Algorithm (TOA), a novel bio-inspired metaheuristic designed to overcome population stagnation and diversity loss. TOA incorporates three key mechanisms: (1) a sigmoid-based fitness transformation enabling hierarchical population partitioning for enhanced diversity; (2) a dynamic cryptobiosis state that adapts search intensity and helps escape local optima; and (3) a DNA-repair operator with adaptive L & eacute;vy-flight perturbations for refined local exploitation. Comprehensive experiments on the CEC 2017, 2020, and 2022 benchmark suites demonstrate that TOA consistently achieves superior performance and statistical rankings across various high-dimensional settings. Applications in engineering design, UAV path planning, and medical diagnosis further confirm its robustness, efficiency, and scalability. These results indicate that TOA provides a flexible and effective framework for complex real-world optimization.
Time-series images with rich spatiotemporal features contain comprehensive and accurate context information for image segmentation. Due to the variability of time-series images, a random offset phenomenon may occur in targets, interfering with the continuity of temporal features. Although windowed attention mechanisms are adopted to capture the complete image information, they are prone to triggering the edge-jagged phenomenon. To address the above issues, this article presents a Swin transformer with spatiotemporal feature correction (SwinTSFC) for the semantic segmentation of time-series images. A convolutional long-short-term memory (ConvLSTM) module with dynamic correction is proposed to adjust the target deviation of temporal data by capturing the offset relationship among sequences. It learns image semantic association and maintains object alignment among dynamic data. A global-to-local learning strategy is adopted to extract spatial features. Swin transformer blocks are adopted to capture the long-range dependencies of images by strengthening interaction capabilities among windows and to improve the overall recognition ability of SwinTSFC. Self-calibrated convolution (SCConv) adaptively extracts fine-grained information to optimize edge continuity features and overcome the phenomenon of edge-jagged. The superiority of the SwinTSFC to state-of-the-art algorithms is demonstrated via experimentation. The code is available at: https://github.com/fjc1575/Marine-Aquaculture/tree/main/SwinTSFC