In the monitoring of multi-condition industrial systems, the lifelong adaptation of vision Pre-trained Models (PTMs) is fundamentally trapped in a capacity-cost dilemma. When confronting endless visual tasks, fixed architectures inevitably hit plasticity bottlenecks, while periodic parameter expansions lead to unsustainable memory explosions. We introduce Trans-MoE, a capacity-boundary constrained self-expanding framework. Rather than blind periodic scaling, it quantitatively evaluates the transferability of knowledge through the Representational Transferability and routes novel concepts to targeted member experts. It architecturally prevents catastrophic forgetting while defeating the capacity-cost dilemma. Extensive experiments with our Trans-MoE on standard benchmarks (e.g., ImageNet-R, 81.23%, CIFAR-100, 89.27%) surpass baseline SEMA by 6.70% and 2.29%, which indicates the superiority of ours on CIL tasks. Meanwhile, the effectiveness of the proposed method is also verified on industrial datasets. Remarkably, it requires only a compact 1M-parameter architecture to surpass unconstrained 4M-parameter baselines, breaking the capacity-cost dilemma.
Accurate and timely diagnosis of crop pests and diseases is vital for safeguarding agricultural productivity and quality. Artificial intelligence has emerged as a powerful tool in this domain, significantly enhancing decision-making in crop health management. However, existing approaches primarily rely on single-modality data for diagnosing specific crops and lack the ability to provide explainable diagnostic reasoning, thereby limiting their scalability and generalizability to diverse crop species in practical applications. To overcome these limitations, this study proposes a large multimodal model named CropGPT to enable diagnosis across all crop types and provide interactive diagnostic explanations. CropGPT is an end-to-end framework that integrates a visual encoder and a large language model. The visual encoder employs the proposed DynamicFocus module to extract multi-level image features encompassing global, local, and object-level information. The large language model incorporates a chain-of-thought design, enabling step-by-step interactive diagnosis of crop diseases along with explanatory reasoning. To enable effective fine-tuning of our model and achieve strong performance across various crops, a dataset named CropInstruct is built based on an automated and cost-efficient paradigm, significantly alleviating the scarcity of high-quality multimodal crop disease data. In addition, we introduce a test-time knowledge augmentation strategy that enhances zero-shot diagnostic performance without requiring retraining, further improving the model's generalizability to a wide range of crops. Experimental results show that CropGPT achieves 0.931 accuracy in diagnosis (>= 35.6% improvement), 71.2 BLEU-4 in image description (>= 44.4%), and 85.3 BLEU-4 in reasoning (>= 47.3%) on 79 crop pest and disease categories, outperforming stateof-the-art multimodal models such as GPT-4o and classical deep learning models under unimodal settings. In zero-shot evaluation, it reaches 0.795 accuracy on 10 unseen crops, surpassing Qwen-VL-Max by 7.3%. These results highlight CropGPT's high precision, interpretability, and generalizability across crop species.
Autoregressive (AR) models based on next-scale prediction are rapidly emerging as a powerful tool for image generation, but they face a critical weakness: information inconsistencies between patches across timesteps introduced by progressive resolution scaling. These inconsistencies scatter guidance signals, causing them to drift away from conditioning information and leaving behind ambiguous, unfaithful features. We tackle this challenge with Information-Grounding Guidance (IGG), a novel mechanism that anchors guidance to semantically important regions through attention. By adaptively reinforcing informative patches during sampling, IGG ensures that guidance and content remain tightly aligned. Across both class-conditioned and text-to-image generation tasks, IGG delivers sharper, more coherent, and semantically grounded images, setting a new benchmark for AR-based methods.
Vision-Language-Action (VLA) models have made significant strides in embodied intelligence by integrating the powerful representations of pre-trained Vision-Language Models (VLMs). However, the massive parameter scale of VLAs imposes a heavy computational burden, and these models exhibit extreme sensitivity to parameter pruning. Current paradigms often treat the resulting performance degradation as inevitable, relying on fine-tuning or low-rank corrections to recover efficacy. We challenge this convention by questioning whether the removed parameters are truly redundant if VLA pruning necessitates performance recovery to be effective, or if this paradigm masks the indiscriminate pruning of critical parameters. We revisit parameter redundancy through the lens of VLM-to-VLA adaptation, first quantifying the spatial distribution of parameter divergence during adaptation to reveal structured patterns across different modules. Subsequently, we introduce controlled pruning as a diagnostic probe: by comparing the direct impact of removing different parameter subsets on VLA performance without any fine-tuning, we establish a causal link between adaptation-induced divergence signals and functional contributions. Based on the discovered modular heterogeneities, we design a multi-module joint pruning scheme. Evaluations on the LIBERO benchmark demonstrate that our approach reduces the parameters of OpenVLA and π_0.5 by 12%–30% while maintaining approximately 90% of the original performance without any post-pruning recovery. In contrast, existing parameter pruning criteria result in total performance collapse when evaluated under the same recovery-free constraints. Our study reveals the parameter evolution mechanism in VLA adaptation and provides a new path for deploying efficient, robust robotic policies in resource-constrained environments.
Vision-language-action models have advanced robotic manipulation but remain constrained by reliance on the large, teleoperation-collected datasets dominated by the static, tabletop scenes. We propose a simulation-first framework to verify VLA architectures before real-world deployment and introduce MobileManiBench, a large-scale benchmark for mobile-based robotic manipulation. Built on NVIDIA Isaac Sim and powered by reinforcement learning, our pipeline autonomously generates diverse manipulation trajectories with rich annotations (language instructions, multi-view RGB-depth-segmentation images, synchronized object/robot states and actions). MobileManiBench features 2 mobile platforms (parallel-gripper and dexterous-hand robots), 2 synchronized cameras (head and right wrist), 630 objects in 20 categories, 5 skills (open, close, pull, push, pick) with over 100 tasks performed in 100 realistic scenes, yielding 300K trajectories. This design enables controlled, scalable studies of robot embodiments, sensing modalities, and policy architectures, accelerating research on data efficiency and generalization. We benchmark representative VLA models and report insights into perception, reasoning, and control in complex simulated environments.
Vision-language-action (VLA) models have advanced the field of embodied manipulation by harnessing broad world knowledge and strong generalization. However, current VLA models still face several key challenges, including limited reasoning capability, lack of status monitoring, and difficulty in self-correction. In this paper, we introduce \textbf{Sentinel-VLA}, a metacognitive VLA model equipped with an active ``sentinel'' module to monitor real-time execution status. Only when necessary, such as during initial planning or upon detecting an error, the model triggers a dynamic reasoning or formulate error recovery solutions. This on-demand reasoning mechanism ensures robust decision-making while minimizing computational overhead. Notably, all training data (spanning 44 tasks and over 2.6 million transitions) is automatically generated and annotated through our designed pipeline. We also propose the Self-Evolving Continual Learning (SECL) algorithm, which allows Sentinel-VLA to identify its capability boundaries and automatically collect data for expansion, paired with Orthogonal Continual Adapter (OC-Adapter) to constrain parameter updates to an orthogonal space, thereby preventing catastrophic forgetting. Real-world experiments demonstrate that Sentinel-VLA boosts the task success rate by over 30\% compared to the SOTA model, PI0. We will open-source all the code, weights, and data generation pipeline.
Automatic movie trailer generation must select shots from a full-length film and synchronize them with background music. Existing methods either relegate music alignment to post-processing or enforce rigid one-to-one shot-music mappings, overlooking that professional editing rhythm is elastic: rapid cuts accompany high-energy passages while sustained shots span quieter bars. We introduce BEAT, a framework that addresses this gap with two core components: MuVA, a compact music-visual alignment encoder trained with Sinkhorn-regularized two-stage learning, and Bar-DP, an energy-adaptive dynamic programming algorithm that produces elastic many-to-one alignments following musical dynamics. These components are integrated into a five-phase agentic pipeline that grounds the core alignment in learned cross-modal features while coordinating higher-level creative decisions through structured text signals. To support comprehensive evaluation, we also introduce TrailerArena, a benchmark with 20+ metrics across four complementary dimensions. On TrailerArena, BEAT achieves state-of-the-art performance across shot selection, ordering, and perceptual quality, while producing fully composed trailers end-to-end.
Active 3D Gaussian reconstruction fundamentally relies on selecting informative next-best views under limited sensing budgets. Existing active 3DGS methods primarily plan viewpoints according to geometric information gain, treating object-induced hidden regions in the same manner as general unexplored space. Under tight frame budgets, such geometry-driven strategies may prioritize global scene coverage while leaving partially observed objects incompletely reconstructed. To address this limitation, we propose OccamView, an object-conditioned view-selection framework for frame-budgeted active 3D Gaussian reconstruction. Rather than predicting unseen object geometry or performing shape completion, OccamView maintains an online object memory from open-vocabulary detections grounded in measured RGB-D observations and represents unresolved local occupancy around detected objects as conservative hidden-region proxies. Candidate viewpoints are then evaluated using an occlusion-aware proxy-coverage score. Furthermore, we introduce a Geo-Floor mechanism that restricts object-conditioned re-ranking to geometrically competitive candidates, allowing object-conditioned cues to guide complementary observations while preserving the geometry-driven exploration behavior of the underlying planner. Experiments on Replica and Matterport3D under a unified frame-budgeted protocol show that OccamView consistently reduces Completion and improves Completion Ratio across five frame budgets, with particularly pronounced gains under limited frame budgets. These results demonstrate that lightweight object-conditioned cues effectively complement geometry-driven active view planning.
Efficient management and precise monitoring are essential for the sustainable control of crop diseases and pests. Traditional unimodal methods exhibit reduced reliability due to data gaps and environmental fluctuations. Multimodal artificial intelligence (AI) offers a promising alternative by integrating complementary data sources and enhancing robustness and adaptability. However, a comprehensive synthesis connecting multimodal AI with multi-scale disease and pest management is still lacking. Based on 950 publications from the past decade reflecting a 31.7% annual growth rate over the past five years, this review examines the evolution of AI-driven research and compares unimodal and multimodal approaches by summarizing major data modalities, fusion strategies, and modeling techniques. Deep learning emerges as the most widely used class of AI methods, and quantitative evidence indicates that multimodal systems achieve approximately 3–48.9% higher diagnostic accuracy than unimodal models. Evidence from 27 studies demonstrates the effectiveness of multimodal fusion across imaging, spectral, environmental, and sensor-based datasets. Building upon these findings, we propose a novel three-level management framework comprising point-level diagnosis, area-scale monitoring, and spatiotemporal forecasting, clarifying how multimodal AI strengthens each task. We further highlight the role of Plant Electronic Medical Records (PEMRs) and outline a conceptual virtual plant clinic to support continuous, data-driven crop health services. Finally, this review identifies key directions including advanced fusion strategies, lightweight and interpretable models, digital twin integration, and scalable decision-support systems, which are essential for intelligent and sustainable crop disease and pest management.
Large Vision-Language Models (LVLMs) have demonstrated capabilities in multimodal understanding, yet their vulnerability to adversarial attacks raises significant concerns. To achieve practical attacking, this paper aims at efficient and transferable untargeted attacks under limited perturbation sizes. Considering this objective, white‑box attacks require full‑model gradients and task‑specific labels, making costs scale with tasks, while black‑box attacks rely on proxy models, typically requiring large perturbation sizes and elaborate transfer strategies. Given the centrality and widespread reuse of the vision encoder in LVLMs, we adopt a gray‑box setting that targets the vision encoder alone for efficient but effective attacking. We theoretically establish the feasibility of vision‑encoder‑only attacks, laying the foundation for our gray‑box setting. Based on this analysis, we propose perturbing patch tokens rather than the class token, informed by both theoretical and empirical insights. We generate adversarial examples by minimizing the cosine similarity between clean and perturbed visual features, without accessing the subsequent models, tasks, or labels. This significantly reduces computational overhead while eliminating the task and label dependence. VEAttack has achieved a performance degradation of 94.5% on image caption task and 75.7% on visual question answering task. We also reveal some key observations to provide insights into LVLM attack/defense: 1) hidden layer variations of LLM, 2) token attention differential, 3) Möbius band in transfer attack, 4) low sensitivity to attack steps.
MLLMs require high-resolution visual inputs for fine-grained tasks like document understanding and dense scene perception. However, current global resolution scaling paradigms indiscriminately flood the quadratic self-attention mechanism with visually redundant tokens, severely bottlenecking inference throughput while ignoring spatial sparsity and query intent. To overcome this, we propose Q-Zoom, a query-aware adaptive high-resolution perception framework that operates in an efficient coarse-to-fine manner. First, a lightweight Dynamic Gating Network safely bypasses high-resolution processing when coarse global features suffice. Second, for queries demanding fine-grained perception, a Self-Distilled Region Proposal Network (SD-RPN) precisely localizes the task-relevant Region-of-Interest (RoI) directly from intermediate feature spaces. To optimize these modules efficiently, the gating network uses a consistency-aware generation strategy to derive deterministic routing labels, while the SD-RPN employs a fully self-supervised distillation paradigm. A continuous spatio-temporal alignment scheme and targeted fine-tuning then seamlessly fuse the dense local RoI with the coarse global layout. Extensive experiments demonstrate that Q-Zoom establishes a dominant Pareto frontier. Using Qwen2.5-VL-7B as a primary testbed, Q-Zoom accelerates inference by 2.52 times on Document OCR benchmarks and 4.39 times in High-Resolution scenarios while matching the baseline's peak accuracy. Furthermore, when configured for maximum perceptual fidelity, Q-Zoom surpasses the baseline's peak performance by 1.1
Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or object differs from training. Adapting to each new situation typically requires collecting more data and fine-tuning. We present StellaVLA, a framework that instead adapts at test time by conditioning on a single retrieved demonstration. The key idea is to move beyond imitating what an expert did and instead convey why: an automated offline pipeline converts each raw trajectory into a structured demonstration, e.g., a task plan, sub-goal descriptions, and verbalized 3D motion, at zero human-annotation cost. Provided as in-context guidance, this structured demonstration lets the policy reason about the task rather than mimic a pixel trajectory, which also makes it transferable across embodiments (real-robot, human-hand, or XR demonstrations). A parallel dual-training design internalizes this reasoning during training through a joint action-and-language objective, while inference uses the action expert alone, preserving real-time, high-frequency control with no added latency. On the VLA-Arena leaderboard(Aug 1, 2026), StellaVLA ranks first with an overall score of 0.63, versus 0.44 and 0.22 for the strong prior models (π_0.5 and LingBot-VLA), and it further leads on LIBERO with 98.8
Generating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may cause conflict between the style and content, harming the integration. Differently, in this work we build a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration. Moreover, we extend the stylized motion generation from one modality, i.e. the style motion, to multiple modalities including texts and images through contrastive learning, leading to flexible style control on the motion generation. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, while also enabling multimodal signals control. The code of our method will be made publicly available.
Linguistic cues such as "I believe" and "probably" offer an intuitive interface for communicating confidence, yet a generalisable, principled calibration framework for linguistic confidence expressions remains underexplored. In particular, co-occurring linguistic cues, contextual variation, and subjective audience interpretation pose unique challenges. We therefore model linguistic confidence as a distribution over plausible perceived probability values that a statement is correct, capturing interpretation variability that scalar representations discard. Within this distributional framework, we introduce faithfulness as a complementary evaluation dimension and present Faithfulness Divergence (FD), an information-theoretic metric quantifying the surprise induced in audience beliefs upon truth revelation. Building on these foundations, we present Retrieval-Augmented Linguistic Calibration (RALC), a lightweight post-hoc pipeline that propagates calibrated confidence signals back into natural language via retrieval-augmented rewriting. Across three QA benchmarks and five LLM families, RALC improves in-domain faithfulness and calibration up to 66
Confidence calibration assumes a unique ground-truth label per input, yet this assumption fails wherever annotators genuinely disagree. Post-hoc calibrators fitted on majority-voted labels, the standard single-label targets used in practice, can appear well-calibrated under conventional evaluation yet remain substantially miscalibrated against the underlying annotator distribution. We show that this failure is structural: under simplifying assumptions, Temperature Scaling is biased toward temperatures that underestimate annotator uncertainty, with true-label miscalibration increasing monotonically with annotation entropy. To address this, we develop a family of ambiguity-aware post-hoc calibrators that optimise proper scoring rules against the full label distribution and require no model retraining. Our methods span progressively weaker annotation requirements: Dirichlet-Soft leverages the full annotator distribution and achieves the best overall calibration quality across settings; Monte Carlo Temperature Scaling with a single annotation per example (MCTS S=1) matches full-distribution calibration across all benchmarks, demonstrating that pre-aggregated label distributions are unnecessary; and Label-Smooth Temperature Scaling (LS-TS) operates with voted labels alone by constructing data-driven pseudo-soft targets from the model's own confidence. Experiments on four benchmarks with real multi-annotator distributions (CIFAR-10H, ChaosNLI) and clinically-informed synthetic annotations (ISIC 2019, DermaMNIST) show that Dirichlet-Soft reduces true-label ECE by 55-87
Long-context language modeling is commonly framed as a scalability challenge of token-level attention, yet local-to-global information structuring remains largely implicit in existing approaches. Drawing on cognitive theories of discourse comprehension, we propose HiCI (Hierarchical Construction–Integration), a hierarchical attention module that constructs segment-level representations, integrates them into a shared global context, and broadcasts both to condition segment-level attention. We validate HiCI through parameter-efficient adaptation of LLaMA-2 with only <5.5
Deep neural networks frequently exhibit overconfidence, undermining reliability in safety-critical applications. Existing adaptive methods rely on indirectly learned proxies of sample difficulty. We establish the logit margin as a direct and principled hardness indicator. We prove that margin tightly bounds the feasible temperature range for any target confidence. Empirically, margin strongly correlates with decision boundary proximity and reveals systematic calibration patterns across difficulty levels. We further identify a fundamental flaw in NLL-based optimization: minimizing NLL can paradoxically worsen calibration. To address this, we introduce Charbonnier-Smoothed SoftECE, a smooth objective that provably upper-bounds the smooth calibration error (smCE). Building on these insights, we propose SMART (Sample Margin-Aware Recalibration of Temperature), a lightweight method that learns a sample-wise margin-to-temperature mapping guided by our calibration-centric objective. Experiments demonstrate state-of-the-art calibration across CNNs and ViTs on standard, long-tailed, and distribution-shifted benchmarks, with a minimal inference-time data consumption. Code: https://anonymous.4open.science/r/SMART-8B11.
Multi-label learning presents unique challenges in predicting and ranking multiple labels, especially in large-scale scenarios. Although Twin Support Vector Machine (TSVM) has demonstrated strong performance in binary classification tasks, their adaptation to the multi-label setting remains underexplored. Existing extensions typically assign a single hyperplane to each label and rely on manually defined thresholds for decision-making, which often leads to suboptimal results. In this work, we propose a new framework named Probabilistic Multi-Label TSVM (PMLTSVM). It reintroduces twin hyperplanes for each label and integrates a probabilistic ranking strategy in place of conventional geometric ranking. This probabilistic perspective not only enhances robustness but also provides a more principled ranking of label relevance. To the best of our knowledge, we are the first to combine probabilistic modeling with TSVM under the binary relevance paradigm. Furthermore, we develop a safe screening rule, capable of discarding non-influential training instances before optimization, thereby substantially reducing both training time and hyperparameter tuning efforts. The proposed method is supported by theoretical guarantees and validated through experiments, consistently outperforming state-of-the-art methods by about 3 %-10 % in key metrics such as Hamming loss and Ranking loss, while accelerating training up to 49.3 times on large datasets.
Diffusion models (DMs) have profoundly transformed the field of generative modeling, delivering exceptional image generation quality. Nevertheless, residuals in the generated images, arising from modeling errors that accumulate along the practical sampling trajectories of pre-trained DMs, remain an inevitable challenge. To mitigate this, we propose a novel residual learning framework built upon a parameterized correction function, designed to improve modeling performance through residual correction. Most notably, our framework exhibits remarkable transferable residual correction capabilities, enabling a correction function optimized for a specific pre-trained DM on a given dataset to enhance the performance of other DMs trained on the same dataset. To further improve our framework, we introduce an advanced training approach that designs an ODE sampler bank to refine residual simulation and incorporates a score consistency maintenance technique to enhance model convergence. Building on this approach, the optimized correction function achieves substantial improvements in residual correction. Extensive experiments performed on four widely used datasets and multiple pre-trained DMs validate the effectiveness and superiority of our residual learning framework.