Generative AI is significantly transforming creative industries. However, design practitioners typically lack systematic frameworks for effectively integrating these tools. This deficiency results in fragmented workflows, reduced efficiency, and underutilized human-AI collaborative potential. This study addresses these gaps by proposing and validating the Synergistic AI-Driven Methodology for Dynamic Poster Design (SADM). The framework was developed through analysis of design theory, human-computer interaction models, and current AI capabilities. SADM structures the workflow into five defined stages: concept exploration, asset generation, dynamic orchestration, generative iteration, and intelligent deployment. Each stage specifies collaborative roles for designers and AI systems. Comparative case studies evaluated SADM against traditional processes across three dimensions: efficiency, creative diversity, and output quality. Findings indicate SADM reduces project timelines by approximately 70% and increases creative solution diversity by over 350%. Critically, the methodology shifts the designer's role from technical executor to strategic guide, enhancing creative strategic value. SADM provides a systematic operational framework addressing theoretical limitations in AI-era creative workflows. It offers practical guidelines for dynamic poster design and establishes a transferable human-machine collaboration paradigm for broader design applications.
Abstract Multipartite structures are common in diverse networked systems, where edges form endogenously only between nodes of different partite sets. Disruptive attacks often degrade large-scale connectivity, which is critical for maintaining the functional order of these networks. However, despite existing work on node-removal vulnerabilities, little is known about multipartite robustness when attackers lack precise network information and thus exhibit imperfect target discrimination. Here we develop a theoretical framework to investigate robustness under such realistic attack behaviors. We define and compute key probabilities governing connectivity to the giant component, constructing a percolation theory for multipartite networks with imperfect target discrimination and arbitrary degree distributions across partite sets of the networks. We derive equations that reveal the intrinsic mechanisms of imperfect target discrimination effects and analytically quantify the resulting robustness for the first time. Numerical simulations on synthetic and real-world multipartite networks demonstrate the effectiveness and sharpness of our theoretical results.
Cross-domain recommendation (CDR) addresses the data sparsity and cold-start problems in the target domain by leveraging knowledge from data-rich source domains. However, existing CDR methods often rely on domain-specific features or identifiers that lack transferability across different domains, limiting their ability to capture inter-domain semantic patterns. To overcome this, we propose SemaCDR, a semantics-driven framework for cross-domain sequential recommendation that leverages large language models (LLMs) to construct a unified semantic space. SemaCDR creates multiview item features by integrating LLM-generated domain-agnostic semantics with domain-specific content, aligned by contrastive regularization. SemaCDR systematically creates LLM-generated domain-specific and domain-agnostic semantics, and employs adaptive fusion to generate unified preference representations. Furthermore, it aligns cross-domain behavior sequences with an adaptive fusion mechanism to synthesize interaction sequences from source, target, and mixed domains. Extensive experiments on real-world datasets show that SemaCDR consistently outperforms state-of-the-art baselines, demonstrating its effectiveness in capturing coherent intra-domain patterns while facilitating knowledge transfer across domains.
Most federated recommender systems represent each user with a single embedding learned from local interaction data, implicitly assuming that user preferences are fixed and precisely identifiable. In federated settings, however, each client observes only a limited and fragmentary view of user behavior, rendering such point estimates inherently brittle. To address this mismatch, we model user preferences as distributions rather than points, allowing multiple compatible preference representations to coexist. Rather than collapsing evidence into a single embedding, our approach preserves uncertainty and diversity in user representations, providing a richer basis for preference modeling. We instantiate this idea with a diffusion-based generative framework that produces diverse user embeddings and derives recommendation scores by aggregating predictions across them. This distributional formulation yields more stable ranking behavior and improved robustness under ambiguous feedback. Extensive experiments on federated recommendation benchmark datasets demonstrate consistent and significant improvements over baselines. Our code is available.
Diffusion-based methods have shown strong potential for stochastic human motion prediction, but existing ap proaches typically formulate the task as a conditional diffusion process initialized from Gaussian noise. This introduces a substantial distribution mismatch between the Gaussian prior and the conditional future-motion manifold, increasing the difficulty of reverse inference and often requiring many denoising steps. Although standard Brownian bridge diffusion provides a more condition-aware alternative, it fixes the endpoint to a deterministic target, resulting in a zero-variance terminal endpoint that may restrict its ability to model the stochastic and multimodal nature of future human motion. To address these issues, we propose HumanBBD, a generalized Brownian bridge diffusion framework for stochastic human motion prediction. HumanBBD explic itly constructs a bridge process between the observed-motion domain and the future-motion domain, thereby reducing initialization mismatch and providing stronger condition-aware generation. Moreover, by relaxing the zero-variance terminal endpoint into a distributional endpoint, HumanBBD improves the flexibility and expres siveness of bridge diffusion for modeling multimodal futures. Built upon this formulation, we further develop a unified spatio-temporal generation backbone that jointly captures long-range temporal dynamics and structured spatial dependencies among joints. Extensive experiments on HumanEva-I, Human3.6M, and AMASS demonstrate that HumanBBD achieves strong prediction performance while requiring only 5-10 reverse steps, substantially improving sampling efficiency compared with conventional diffusion-based methods. Comprehensive empiri cal analyses further validate the effectiveness, efficiency, and robustness of the proposed framework. Overall, HumanBBD provides an effective and efficient alternative to standard conditional diffusion for stochastic human motion prediction.
Cycles are ubiquitous in various networks such as social, biological, and technological systems, where they play a significant functional and dynamical role. This paper proposes a node similarity measure based on minimal simple cycles, referred to as cycle similarity. Specifically, the metric quantifies the similarity between two nodes by considering the minimal cycles that connect them through their neighboring nodes, with an upper bound imposed on the cycle size to ensure computational feasibility. We then systematically examine the effectiveness and applicability of this similarity measure through two fundamental tasks: link prediction and community detection. To address the scarcity of cycles in link prediction, an edge-addition correction strategy is introduced, whereby the existence of a candidate edge is hypothetically assumed before computing node similarity. Experimental results demonstrate that this correction leads to improved performance on datasets including karate, INT, PPI, and Grid. In hierarchical community detection using cycle similarity, we find that the significance of cyclic structures (reflected by Z-scores), the presence of pendant nodes with degree one, and the existence of cut vertices are the primary factors influencing the algorithm's performance.
Federated learning has emerged as a promising paradigm to alleviate privacy risks in recommender systems. However, most federated recommendation methods mainly focus on user-item interactions and cannot fully exploit preferences from user-bundle and bundle-item graphs in bundle recommendation. Extending federated learning to bundle recommendation introduces two challenges: (1) parameter aggregation to handle non-IID data while adapting to bundle recommendation, and (2) bundle representation enhancement by integrating multi-view user preferences under privacy constraints. To this end, we propose Dual-Channel Personalized Federated Bundle Recommendation (DPFBR), which comprises two channels: first is the group personalization channel that groups users to generate personalized models, and second is the bundle enhancement channel that integrates multi-view preferences to guide local updates. Experiments show DPFBR achieves strong recommender performance. The code is available at https://github.com/SHER-afk/DPFBR.
Lead discovery is a critical step in drug development, yet current computational methods often face a trade-off between exploring novel chemical space and leveraging target-specific knowledge priors. Score-oriented algorithms can support broad exploration but lack medicinal-chemistry guidance, whereas knowledge-driven design benefits from target-specific knowledge while often staying close to existing chemotypes, limiting their utility in the pivotal real-world scenario of designing leads for targets that lack bioactive ligands. Here, we present a lead-discovery framework, AutoLeadDesign, that couples large language model (LLMs) reasoning with fragment-based chemical exploration without requiring target-specific ligand templates. Through a mutually reinforcing feedback loop between the LLM and chemical fragments, AutoLeadDesign enables target-aware and interpretable exploration of chemically diverse scaffolds while preserving critical binding interactions. Computational analysis reveals that AutoLeadDesign mitigates the tendency towards local optimization often observed in LLM-based molecular design. Analysis of the design trajectories indicates that the design process shares similar mechanisms to medicinal-chemistry strategies reported in expert-led studies (specifically fragment linking, merging, and growing), supporting the interpretability of the binding modes. Furthermore, we synthesized and experimentally validated leads targeting two therapeutically relevant targets: SARS-CoV-2 PLpro and the oncogenic KRAS G12D mutant. One designed inhibitor, PLP011, showed potent antiviral activity by suppressing viral replication and restoring the viability of infected cells. Another candidate, KP032, achieved a half-maximal inhibitory concentration (IC 50 ) of 50.24 nM against the challenging target KRAS G12D. Finally, we released the fragment libraries targeting the therapeutic proteins investigated in this study. These libraries are intended to serve as fragment-based priors, providing medicinal chemists and future machine learning frameworks with essential structural insights for targeted drug design. By combining LLM-based biochemical reasoning with fragment-grounded exploration, AutoLeadDesign provides a practical strategy for more adaptive and interpretable lead discovery.
User-centric recommendation has become essential for delivering personalized services, as it enables systems to adapt to users' evolving behaviors while respecting their long-term preferences and privacy constraints. Although federated learning offers a promising alternative to centralized training, existing approaches largely overlook user behavior dynamics, leading to temporal forgetting and weakened collaborative personalization. In this work, we propose FCUCR, a federated continual recommendation framework designed to support long-term personalization in a privacy-preserving manner. To address temporal forgetting, we introduce a time-aware self-distillation strategy that implicitly retains historical preferences during local model updates. To tackle collaborative personalization under heterogeneous user data, we design an inter-user prototype transfer mechanism that enriches each client's representation using knowledge from similar users while preserving individual decision logic. Extensive experiments on four public benchmarks demonstrate the superior effectiveness of our approach, along with strong compatibility and practical applicability. Code is available.
While Vision-Language Models (VLMs) have significantly advanced remote sensing interpretation, enabling them to perform complex, step-by-step reasoning remains highly challenging. Recent efforts to introduce Chain-of-Thought (CoT) reasoning to this domain have shown promise, yet ensuring the visual faithfulness of these intermediate steps remains a critical bottleneck. To address this, we introduce GeoSolver, a novel framework that transitions remote sensing reasoning toward verifiable, process-supervised reinforcement learning. We first construct Geo-PRM-2M, a large-scale, token-level process supervision dataset synthesized via entropy-guided Monte Carlo Tree Search (MCTS) and targeted visual hallucination injection. Building upon this dataset, we train GeoPRM, a token-level process reward model (PRM) that provides granular faithfulness feedback. To effectively leverage these verification signals, we propose Process-Aware Tree-GRPO, a reinforcement learning algorithm that integrates tree-structured exploration with a faithfulness-weighted reward mechanism to precisely assign credit to intermediate steps. Extensive experiments demonstrate that our resulting model, GeoSolver-9B, achieves state-of-the-art performance across diverse remote sensing benchmarks. Crucially, GeoPRM unlocks robust Test-Time Scaling (TTS). Serving as a universal geospatial verifier, it seamlessly scales the performance of GeoSolver-9B and directly enhances general-purpose VLMs, highlighting its remarkable cross-model generalization.
The core-periphery structure is a typical representative of the mesoscale structure in networks, which dominates the basic and potential function of the networked systems. In reality, networks inevitably face attacks, among which localized attacks have received much increasing attention since they provide efficient models for typical disruptive scenarios. However, the existing research primarily focuses on a single localized attack, with limited studies addressing multiple localized attacks, particularly regarding the impact on core-periphery structures. Here, we propose a framework for multiple localized attacks and introduce a new index to assess the vulnerability of core-periphery structure within our framework. Enhancing algorithm based on the proposed framework and index is exquisitely designed. Experimental results indicate that the algorithm can effectively measure and enhance the robustness of core-periphery structure within a limited defense budget. Furthermore, we discover the structural characteristics of the network that influence its core-periphery vulnerability. The interesting differences in defense strategies when optimizing various networks are also discussed.
Building polygon extraction is a critical task in remote sensing analysis and a fundamental component of modern urban management. Conventional segmentation-based methods often suffer from geometric distortions during the conversion from masks to polygons. End-to-end polygon prediction approaches (e.g., PolyWorld) alleviate this issue by directly predicting building polygons; however, existing PolyWorld-like methods remain limited in accurate corner vertex detection and polygon reasoning due to insufficient representation learning, particularly for geometry. In this work, we propose PolyGeom, an end-to-end framework equipped with a geometry-aware graph transformer for accurate and robust building polygon extraction. PolyGeom employs the Segment Anything Model (SAM) as its backbone to leverage large-scale pretrained features, thereby capturing both local and global semantics. Moreover, we propose a geometry-aware graph transformer that explicitly models geometry of building polygons, facilitating more reliable polygon reasoning. Extensive experiments on three challenging benchmarks, CrowdAI, WHU, and BONAI datasets, demonstrate that PolyGeom consistently outperforms existing methods in terms of building detection accuracy, topology correctness, and geometry alignment. Ablation studies further validate the effectiveness of the two key proposed designs in building polygon extraction.
Autoregressive models are structurally misaligned with the inherently parallel nature of geospatial understanding, forcing a rigid sequential narrative onto scenes and fundamentally hindering the generation of structured and coherent outputs. We challenge this paradigm by reframing geospatial generation as a parallel refinement process, enabling a holistic, coarse-to-fine synthesis that resolves all semantic elements simultaneously. To operationalize this, we introduce GeoDiT, the first diffusion-based vision-language model tailored for the geospatial domain. Extensive experiments demonstrate that GeoDiT establishes a new state-of-the-art on benchmarks requiring structured, object-centric outputs. It achieves significant gains in image captioning, visual grounding, and multi-object detection, precisely the tasks where autoregressive models falter. Our work validates that aligning the generative process with the data's intrinsic structure is key to unlocking superior performance in complex geospatial analysis.
Customized sports training routines take into account individual physiology, fatigue, and recovery to maximize performance. Proximal Policy Optimization (PPO)-based reinforcement learning is used to adjust training intensity, duration, and rest in a simulated endurance-training environment for runners, using real-time wearable and performance data. The environment models athlete status utilizing heart rate variability, VO₂ max, fatigue ratings, and injury-risk indicators. PPO is trained to maximize performance gains, recovery quality, and safety over repeated sessions. Simulated policy improves performance (18.6%), injury-risk deviation (−22.4%), recovery compliance (91.3%), training load variability control (±7.2%), reward-signal evolution convergence (+41.7%), session completion rate (94.6%), personalized adaptation score (87.5%), and fatigue index stability (94.3%). Results show that a PPO-based RL setup, specifically defined by state design, reward shaping, and multi-episode training, can provide adaptive and data-driven tailored sports training.
Sparse Mobile Crowd Sensing is a practical paradigm for data collection, where mobile users report partial measurements and the platform infers the missing values by exploiting spatiotemporal correlations. Most existing methods assume uniform spatial scales and aggregate observations into fixed time slots. In practice, however, heterogeneous devices and asynchronous reporting lead to observations that vary in spatial coverage and arrive at irregular timestamps. Spatially, coarse-scale measurements often cannot be perfectly decomposed into fine-scale grids due to offset or partial overlaps; tempo-rally, continuous irregular observations may misalign with the desired target reporting moments. These spatial and temporal misalignments jointly make data inference substantially more challenging. To tackle this problem, we propose ST-NAI, a unified framework for spatiotemporal non-aligned data inference. ST-NAI models continuous temporal dynamics with Time-Mamba by fusing physical time intervals, and resolves geometric mismatches via a scale-adjusted resampling strategy. We further enhance spatiotemporal integration with ST-Mamba and Cross-Mamba to capture global context with linear complexity. Experiments on six real-world datasets demonstrate the effectiveness of the proposed framework in handling data inference challenges.
Vision-language pretraining models have made significant progress in bridging remote sensing imagery with natural language. However, existing approaches often fail to effectively integrate multi-granular visual and textual information, relying primarily on global image-text alignment. This limitation hinders the model's ability to accurately capture fine-grained details in images, thus restricting its performance in complex, fine-grained tasks. To address this, we propose GeoAlignCLIP, a unified framework that achieves fine-grained alignment in remote sensing tasks by learning multi-granular semantic alignments and incorporating intra-modal consistency, enabling more precise visual-semantic alignment between image regions and text concepts. Additionally, we construct RSFG-100k, a fine-granular remote sensing dataset containing scene descriptions, region-level annotations, and challenging hard-negative samples, providing hierarchical supervision for model training. Extensive experiments conducted on multiple public remote-sensing benchmarks demonstrate that GeoAlignCLIP consistently outperforms existing RS-specific methods across diverse tasks, exhibiting more robust and accurate fine-grained vision-language alignment.
As a classic problem in computer vision, object detection has become one of the essential challenges that re searchers continue to explore. The emergence of You Only Look Once (YOLO) has transformed object detection from two-stage to single-stage detection, enhancing real-time performance. By transforming the object detec tion task into a regression problem, the detection speed and efficiency have also been significantly improved. This article elaborates on the development history of YOLO object detection algorithm in the past decade, with a focus on the technological evolution, evaluation indicators, dataset selection, and variant improvements from 2016-2025. We have systematically reviewed the technological innovations and major contributions from YOLOv1 to YOLOv13, including the anchor box mechanism, multi-scale prediction, attention module, lightweight design, and anchor-free architecture. Meanwhile, the frequency of use of evaluation metrics for object detec tion, including Frames Per Second (FPS), Giga Floating-Point Operations Per Second (GFLOPs), Precision (P), Recall (R), Receiver Operating Characteristic (ROC), Intersection over Union (IoU), F1-score, PR curve, Average Precision (AP), and Mean Average Precision (mAP), was analyzed using statistical literature methods. YOLO algorithm was analyzed for its proportion of utilization in object detection, image classification, and semantic segmentation on various datasets through commonly used datasets, PASCAL VOC, MS COCO, and ImageNet. Finally, the article summarizes the technological innovations and future development trends of the YOLO series, providing a reference for researchers.
Symbolic Regression (SR) is a core challenge in both physics and artificial intelligence, aiming to identify mathematical equations from experimental data. Recent impressive advancements in Large Language Model (LLM)-based SR methods have addressed the limitations of traditional approaches in flexibly incorporating prior knowledge to enhance accuracy, but they still face challenges such as high costs and scalability with numerous variables. To overcome these issues, we introduce SymBOL. This general-purpose symbolic learning framework uses Bayesian Optimization (BO) to guide the generation of high-quality mathematical expressions from the LLM while accommodating complex tasks with a large number of interdependent variables. Experiments on benchmark datasets demonstrate that SymBOL significantly outperforms baseline methods in both accuracy and efficiency. Notably, compared to advanced LLM-based approaches, SymBOL achieves a 24.85% improvement in average accuracy while reducing computational costs by 28.73%. This advantage extends to high-dimensional SR tasks, where SymBOL substantially lowers the average error. Furthermore, when applied to real-world systems in materials science and epidemiology, SymBOL accurately recovers governing equations and provides interpretable pathways for equation discovery. These findings underscore the potential of SymBOL for advancing scientific discovery.
Structure-based drug design (SBDD) can be effectively realized through an iterative refinement via the Design-Make-Test-Analyze (DMTA) cycle, which is a common workflow used by human experts. However, most LLMs function as one-shot generators that lack feedback mechanisms, leaving the DMTA loop disconnected. In this work, we propose K-BTS, a Knowledge-Driven Bi-level Thompson Sampling framework that formalizes iterative SBDD as a Dynamic Hierarchical Multi-Armed Bandit problem. K-BTS closes the DMTA loop by decoupling decisions into two levels: an upper-level policy that prioritizes high-potential molecular lineages and a lower-level mechanism that retrieves explicit chemical rules to guide LLM generation. By integrating a dual-level Bayesian update, the framework transforms sparse docking scores into reusable experience. On the CrossDocked2020 benchmark, K-BTS achieves a state-of-the-art Top-1 average docking score. The results from diverse dimensions show that K-BTS ensures search determinism through a smooth, monotonic convergence that synchronizes structural drift with affinity improvement.
Sparse sensing is a critical task in many real-world systems, heavily relying on effective data imputation. However, most existing methods overlook the severe sparsity of historical data available for training, which poses a fundamental challenge to achieving accurate imputation. Intrinsic priors-based models (i.e., MF, ODEs) are robust to sparsity but have limited performance with abundant historical data, whereas deep learning models excel with rich data but fail dramatically when data are sparse. Furthermore, this data sparsity introduces significant solution uncertainty and creates a need for more computationally efficient dependency modeling. To address these, we propose SH-Imputer, a novel adaptive framework. It first employs our proposed VDMF and AVDCDE modules to model the uncertainty of intrinsic spatiotemporal properties. Subsequently, a lightweight ST-Mamba module efficiently learns complex spatiotemporal dependencies. The entire process is governed by an adaptive mechanism that balances prior-based robustness with data-driven expressiveness, ensuring superior performance under varying historical data sparsity. We also provide theoretical justifications for our core design choices. Extensive experiments validate that SH-Imputer significantly outperforms state-of-the-art methods. Our code is available at https://github.com/JLUDhhh/SH-Imputer.