The rapid development of large language model (LLM) agents has facilitated the advancement of multiagent systems, where cognitive knowledge sharing is crucial for the execution of complex tasks. However, achieving the synchronization of the cognitive knowledge base (KB) among agents under restricted wireless resources remains a challenge, especially in dynamic real-time environments. Therefore, we propose a hierarchical LLM-agent system that consists of a high-level cluster brain (CB) and multiple lower level LLM agents (LLAs). The cognitive KB of each LLA is represented in the form of a knowledge graph (KG). To improve the efficiency of transmitting cognitive KB updates from LLAs to CB, a KG compression framework named MED-EmPress is proposed, which adaptively compresses the semantic features of cognitive KB by applying dimensionality reduction and binary quantization, and then a joint optimization problem of semantic compression and resource allocation (SCRA) is formulated to maximize semantic fidelity of the cognitive KB being transmitted. To solve this problem, a hierarchical SCRA algorithm is designed to decouple the SCRA problem into two subproblems, which involve dynamically allocating wireless resources and rationally choosing the semantic compression (SC) ratio of the cognitive KB. The goal is to maximize the system's semantic fidelity. The evaluation results demonstrate that the MED-EmPress framework reduces the size of the cognitive updates (CUs) that need to be transmitted by 96%, with only a loss of 3.6% in the entity alignment task. Furthermore, the proposed adaptive compression and transmission scheme improves semantic fidelity by 94% compared to existing methods when wireless resources are severely limited.
Deep joint source-channel coding (Deep JSCC) has emerged as a key technology for semantic communication. However, existing methods are unable to guarantee the fidelity of Regions of Interest (ROI) under severe bandwidth constraints and adverse channel conditions. To address this issue, we propose a text-guided ROI-aware JSCC framework, leveraging textual cues to direct the encoder’s focus on the ROI and incorporate semantic priors for improved robustness. In this framework, a bidirectional multi-scale alignment strategy is introduced to ensure cross-modal consistency and preserve semantic alignment between visual features and text across various feature hierarchies. Experimental results demonstrate that our method achieves 15–30% improvement over baselines under various channel conditions and compression ratios, with the largest gains under low signal-to-noise ratios and bandwidth-constrained conditions.
Referring Image Segmentation (RIS) aims to generate specified target masks in the image using natural language. While existing methods have made progress in modeling the relationship between words and pixels, they often overlook sentence-level semantic information. This limits the model's ability to fully comprehend the deeper meaning of language, affecting target localization and segmentation. To address this problem, we propose a Context-aware Mutual Attention Network (CMANet), which integrates both word-level and sentence-level semantic information to guide visual features in generating precise object masks. Specifically, during the feature encoding stage, we design a Shallow Mutual Attention (SMA) module to reduce the discrepancy between visual and linguistic representations, enhancing pixel-word alignment. In the global representation stage, we introduce a Context-aware Mutual Attention (CMA) module that utilizes sentence-level target semantics to guide the contextual representation of multi-modal features. Experiments conducted on several commonly used RIS datasets, including the natural image referring segmentation dataset, Robust RIS dataset, and referring remote sensing image dataset, show that CMANet outperforms current state-of-the-art methods on all these datasets, demonstrating superior segmentation accuracy.
Sufficient embodied scene understanding serves as the foundation for embodied agents to perceive, interpret, and solve scene-related questions in a scene. Such understanding is often constrained by limited perception, which can be summarized into two aspects: 1) the lack of perceptual abilities in the agent and 2) the scene itself is incomplete. Existing large models-based methods attempt to leverage implicit knowledge to overcome these limitations, but this lacks interpretability and controllability. Inspired by human explainable associative thinking, we propose SceneReasoner framework, which imposes explicit functional associative rules on LLMs to guide the process of the scene understanding. This framework mines deeper functional relationships between objects, enabling the agent to gain sufficient scene understanding from limited perception in a controllable manner. Specifically, SceneReasoner employs an associative knowledge base to provide such rules from two aspects: 1) functional complementarity of objects in a scene. For instance, in a computer workspace, if the agent perceives a monitor and other unclear objects, it can first analyze the function of the monitor in this area (content display), and then infer the presence of other related objects (e.g., a mouse and keyboard for content input); and 2) commonality of objects in a scene. For example, in a picture-hanging task, when the hammer in the scene is missing, the agent needs to first identify the hammer's attributes (hard and applying force), and then associate a suitable substitute in the scene (e.g., a hard wrench). Due to such explicit functional association, the agent can rapidly form a sufficient scene understanding and effectively solve scene-related questions. Experimental results demonstrate that such explicit association augmented with functional reasoning can significantly enhance agents' scene understanding under limited perception. It improves perceptual quality by 9.75% and scene reasoning ability by 21.42% compared with other methods.
Current video understanding models follow a passive perception paradigm, relying on fixed sampling strategies that can introduce substantial temporal redundancy and fine-grained detail loss. To overcome this, we transition from passive perception to active evidence acquisition over a compact and reusable video environment. We propose HIVE (Human-Inspired Video Explorer), a training-free agentic framework driven by a coarse-to-fine spatiotemporal strategy. HIVE first constructs a compact temporal environment by pruning visually redundant segments and augmenting the retained segments with semantic captions, forming a query-ready bimodal memory bank. On top of this abstracted environment, an LLM-based agent executes an iterative Planning-Action-Reasoning-Reflection loop to adaptively determine “when and where to look.” Guided by the user query and a global event synopsis, the agent invokes perceptual tools for temporal localization, caption retrieval, active spatial cropping, and localized video VQA, enabling it to select informative temporal moments and inspect question-relevant spatial regions only when necessary. Without task-specific fine-tuning, HIVE achieves 60.8% accuracy on the long-video subset of Video-MME, 67.4%/65.2% on the EgoSchema subset/full benchmark, and 75.8% on NExT-QA. These results demonstrate that active evidence acquisition provides an efficient alternative to uniformly encoding dense video inputs, improving long-video understanding and fine-grained visual reasoning under limited computational budgets.
Scene graph generation denotes the process of parsing a visual input into a graph representation for downstream reasoning tasks. Most existing research models input scenes as flat scene graphs, capturing solely on horizontal relationships between two objects, while disregarding vertical dependencies. Such flat scene graphs merely describe direct relationships between objects, and cannot represent higher-level semantics. The lack of hierarchical relationships also limits the performance of scene graphs in downstream reasoning tasks. In this work, we propose hierarchical scene graph generation, a new problem task that requires the model to generate scene graphs with clear hierarchy. To achieve this goal, a hierarchical scene graph generation with coarse-to-fine reasoning framework is proposed. It contains three stages: first, summarize the main caption of the scene; second, describe the main visual elements related to the theme; and finally, add secondary elements with background. To facilitate training, a high-quality hierarchical scene graph dataset that comprises 40 K image-scene description pairs is constructed. Utilize rich knowledge and powerful multi-turn conversation capabilities of multi-modal large language models, the proposed framework can not only achieve better semantic understanding and relational modeling, but also seamlessly integrating its relational modeling capabilities to enhance visual-language tasks. Extensive experimental results demonstrate that our method constructs an effective scene representation, outperforming the state-of-the-art by 3.35 % in recall while generating significantly fewer average triplets 8.7 vs. 9.2. The results on downstream tasks also indicate that hierarchical scene graph significantly contributes to visual-language interactive reasoning.
In the information age, individuals and organizations have accumulated massive volumes of digital data. As society enters the intelligence era, the primary demand has shifted toward effectively utilizing these data through intelligent models. However, prevailing large language models (LLMs) typically require data to be uploaded to centralized cloud servers for processing, making it difficult to ensure data privacy. Consequently, users increasingly seek small, locally deployable models that can operate directly on private data. To realize this expectation, this paper explores intelligent systems built upon small models rather than monolithic large models. On this basis, we first revisit the essence of intelligence as the capability to utilize knowledge to solve problems and accomplish tasks, and further propose the first principle of intelligence, which models intelligence as a closed-loop, goal-oriented process. Within this process, an agent, under external constraints, leverages knowledge to evaluate discrepancies between its internal state and task objectives, formulates strategies to reduce these discrepancies, and iteratively executes actions to converge toward goal completion. From this perspective, intelligence is viewed as a collaborative system of interdependent cognitive functions. Grounded in this theoretical principle, we propose the six-capability network, a knowledge-driven cognitive architecture that decomposes intelligence into six fundamental capabilities: observation, attention, understanding, discrimination, memory, and execution. These capabilities constitute the core cognitive model of the proposed framework and are each realized by lightweight, deployable small models operating over structured knowledge representations. Finally, to demonstrate the executability of the network, we consider a networked intelligence system. In this system, agents exchange semantic symbols that encode knowledge rather than raw data, enabling each agent to operate within its functional scope while acquiring missing knowledge from others when needed. This approach offers a new intelligence paradigm, enabling deployment in private, local, and resource-constrained environments.
Multi-layer dictionary learning (MDL) has demonstrated significantly improved performance for image classification. However, most of the existing MDL methods just overall shared dictionary learning architecture, which weakens the discrimination ability of the dictionaries. For this, we proposed a powerful framework called the Multi-layer Graph Constraint Dictionary Pair Learning (MGDPL). Our MGDPL integrates multi-layer dictionary pair learning, structure graph constraint, and discrimination sparse representations into a unified framework. First, the multi-layer structured dictionary learning mechanism is applied to dictionary pairs to enhance the discrimination performance by rebuilding the reconstruction error of the previous layer via the latter layer. Second, it subjects the structure graph constraint on the sub-sparse representations to ensure the discrimination capability of the near neighbor graph. Third, the multi-layer discriminant graph regularized constraint term can ensure high intra-class tightness and inter-class dispersion of dictionary atoms in reconstruction space. Extensive experiments show that MGDPL can achieve excellent performance over other state-of-the-arts.
This study aims to improve human pose estimation under occlusion by hierarchically fusing vision and language in a parse graph framework. Language offers rich priors, such as spatial relations, but existing visual-language fusion using global features often weakens responses in occluded regions, causing alignment and localization errors. To address this, we propose Parse Graph-based Visual-Language interaction (PGVL) with a core novel Guided Module (GM), where low-level nodes preserve local features and high-level nodes preserve global features. PGVL performs hierarchical top-down decomposition and bottom-up composition via recursive cross-attention, guided by GM. GM enables high-semantic nodes to guide feature updates of cross-attention-processed low-semantic nodes, ensuring correct cross-modal fusion. We also design network based on PGVL, which achieves 68.2 MAP on CrowdPose (+0.7 over HRNet-W32), 82.1 MAP on AP-10K (+4.3 over CLAMP) and 79.3 MAP on Animal-Pose (+5.0 over CLAMP). PGVL and our network is validated on major pose estimation datasets. The code link is at https://github.com/lushbng/PGVL.
Large pre-trained vision-language models (VLMs) like CLIP have shown great potential for solving the unsupervised domain adaptation (UDA) problem. Existing prompt learning for UDA based on the unsupervised-trained VLMs requires distribution alignment between source and target domains in the common space for both vision and language branches. However, it is difficult for rough cross-domain alignment to maintain the discriminative semantic structure of both domains. Besides, the coarse features with non-informative noises due to ignoring the pseudo-label noises may cause failures to concentrate on precise semantics alignment. In this work, we propose a Prompt-Based Invertible Mapping Alignment (PIMA) method to incorporate discriminative domain knowledge into prompt learning, which is featured with refined cross-domain alignment in two separate space with a well-kept structure. Specifically, we design an invertible neural network-based homeomorphism mapping, and then achieve distribution alignment through such invertible mapping for connecting source and target visual feature space, which can preserve the data semantic structure. For better semantic alignment in vision-language space, we develop cross-modal implicit contrastive learning module to regularize non-informative features, which aims to find the low-rankness of implicit representation space. We conducted extensive experiments on three benchmark datasets to prove the advantages of our proposed PIMA over state-of-the-art methods.
This study focuses on Embodied Complex-Question Answering task, which means the embodied robot need to understand human questions with intricate structures and abstract semantics. The core of this task lies in making appropriate plans based on the perception of the visual environment. Existing methods often generate plans in a once-for-all manner, i.e., one-step planning. Such approach rely on large models, without sufficient understanding of the environment. Considering multi-step planning, the framework for formulating plans in a sequential manner is proposed in this paper. To ensure the ability of our framework to tackle complex questions, we create a structured semantic space, where hierarchical visual perception and chain expression of the question essence can achieve iterative interaction. This space makes sequential task planning possible. Within the framework, we first parse human natural language based on a visual hierarchical scene graph, which can clarify the intention of the question. Then, we incorporate external rules to make a plan for current step, weakening the reliance on large models. Every plan is generated based on feedback from visual perception, with multiple rounds of interaction until an answer is obtained. This approach enables continuous feedback and adjustment, allowing the robot to optimize its action strategy. To test our framework, we contribute a new dataset with more complex questions. Experimental results demonstrate that our approach performs excellently and stably on complex tasks. And also, the feasibility of our approach in real-world scenarios has been established, indicating its practical applicability.
Transformer has demonstrated remarkable performance in various computer vision tasks. However, its potential is not fully explored in skeleton-based action recognition. On one hand, existing methods primarily utilize fixed function or pre-learned matrix to encode position information, while overlooking the sample-specific position information. On the other hand, these approaches focus on single-scale spatial relationships, while neglecting the discriminative fine-grained and coarse-grained spatial features. To address these issues, we propose a Multi-Scale Adaptive Skeleton Transformer (MSAST), including Adaptive Skeleton Position Encoding Module (ASPEM), Multi-Scale Embedding Module (MSEM), and Adaptive Relative Location Module (ARLM). ASPEM decouples spatial-temporal information in the position encoding procedure, which acquires inherent dependencies of skeleton sequences. ASPEM is also designed to be dependent on input tokens, which can learn sample-specific position information. The MSEM employs multi-scale pooling to generate multi-scale tokens that contain multi-grained features. Then, the spatial transformer captures multi-scale relations to address the subtle differences between various actions. Another contribution of this paper is that ARLM is presented to mine suitable location information for better recognition performance. Extensive experiments conducted on three benchmark datasets demonstrate that the proposed model achieves Top-1 accuracy of 94.9%/97.5% on NTU-60 C-Sub/C-View, 88.7%/91.6% on NTU-120 X-Sub/X-Set and 97.4% on NW-UCLA, respectively.
Referring image segmentation aims to segment the target by a given language expression. Recently, the bottom-up fusion network utilizes language features to highlight the most relevant regions during the visual encoder stage. However, it is not comprehensive that establish only the relationship between pixels and words. To alleviate this problem, we propose a mixed-scale cross-modal fusion method that widens the interaction between vision and language. Specially, at each stage, pyramid pooling is used to augment visual perception and improve the interaction between visual and linguistic features, thereby highlighting relevant regions in the visual data. Additionally, we employ a simple multi-scale feature fusion module to effectively combine multi-scale aligned features. Experiments conducted on Standard RIS benchmarks demonstrate that the proposed method achieves favorable performance against state-of-the- art approaches. Moreover, we conducted experiments on different visual backbones respectively, and the proposed method yielded better and significantly improved performance results.
Visual relationships are different, and their types play an important role in visual scene understanding. However, most of the existing works ignore the different types of visual relationships and adopt the unified approach to learn all visual relationships. It not only limits the flexibility of the model, but also causes fuzzy relationship representation. To address this problem, we deeply study four visual relationship types. Among them, the geometric reflects the spatial interaction, the semantic reflects an action, and the possessive and misc indicates an intrinsic correlation. And then, we propose a novel method- Types Determine Methods (TDM) - which designs different learning strategies according to the relationship types to infer visual relationships. Experiments demonstrate that our approach achieves superior or competitive performance over previous methods, validating its effectiveness.
Visual expertise-the ability to discriminate highly similar exemplars quickly and accurately within a category-supports skilled performance across real-world domains and is supported by distributed neural systems. We focus on non-face expertise to test cross-domain convergence in acquired real-world visual skills, treating faces separately because socially embedded, sensitive-period-constrained processing could blur this inference. It remains unresolved whether non-face expertise across heterogeneous domains converges on a shared, domain-general whole-brain architecture, or instead recruits domain-contingent neural configurations that vary with task and stimulus demands. We conducted a coordinate‑based meta‑analysis of 22 task‑fMRI studies spanning 11 real‑world non‑face expertise domains (579 participants, 210 peak‑activation foci). Primary analysis revealed a robust, right‑lateralized parieto‑temporo‑occipital circuit centered on the middle occipital gyrus, middle temporal gyrus, angular gyrus and adjoining inferior parietal lobule. We propose that this circuit constitutes a domain‑general neural core that integrates fine-grained visual features with semantic associations while supporting attention‑guided recognition of visually similar objects in expert performance. Subgroup and meta‑regression analyses uncovered a complementary adaptive component, the engagement of which varied systematically with representational and contextual factors. Pictorial stimuli and expert-novice contrasts reliably strengthened recruitment of the right-hemisphere core, whereas symbolic stimuli engagement toward left temporal regions while selectively re-engaging right-core nodes. In addition, male-skewed samples showed attenuated left-hemisphere activation. Together, these findings delineate a stable right-hemisphere neural scaffold underlying non-face visual expertise, flexibly supplemented by left-hemisphere systems as a function of stimulus format, task demands, and demographic context, providing a whole-brain reference framework for future studies. ### Competing Interest Statement The authors have declared no competing interest. the National Key R&D Program of China, Grant No.2022YFF1202400
Existing models for recognizing human-object interaction (HOI) in videos mainly rely on visual information for reasoning and generally treat recognition tasks as traditional multi-classification problems, where labels are represented by numbers. This supervised learning method discards semantic information in the labels and ignores advanced semantic relationships between actual categories. In fact, natural language contains a wealth of linguistic knowledge that humans have distilled about human-object interaction, and the category text contains a large amount of semantic relationships between texts. Therefore, this paper introduces human-object interaction category text features as labels and proposes a natural language supervised learning model for human-object interaction by using natural language to supervise visual feature learning to enhance visual feature expression capability. The model applies contrastive learning paradigm to human-object interaction recognition, using an image-text paired pre-training model to obtain individual image features and interaction category text features, and then using a spatial-temporal mixed module to obtain high semantic combination-based human-object interaction spatial-temporal features. Finally, the obtained visual interaction features and category text features are compared for similarity to infer the correct video human-object interaction category. The model aims to explore the semantic information in human-object interaction category label text and use a large number of image-text paired samples trained by a multi-modal pre-training model to obtain visual and textual correspondence to enhance the ability of video human-object interaction recognition. Experimental results on two human-object interaction datasets demonstrate that our method achieves the state-of-the-art performance, e.g., 93.6% and 93.1% F1 Score for Sub-activity and Affordance on CAD-120 dataset.
Vision Transformers (ViTs) have recently demonstrated significant potential in computer vision, but their high computational costs remain a challenge. To address this limitation, various methods have been proposed to compress ViTs. Most approaches utilize spatial-domain information and adapt techniques from convolutional neural networks (CNNs) pruning to reduce channels or tokens. However, differences between ViTs and CNNs in the frequency domain make these methods vulnerable to noise in the spatial domain, potentially resulting in erroneous channel or token removal and substantial performance drops. Recent studies suggest that high-frequency signals carry limited information for ViTs, and that the self-attention mechanism functions similarly to a low-pass filter. Inspired by these insights, this paper proposes a joint compression method that leverages properties of ViTs in the frequency domain. Specifically, a metric called Low-Frequency Sensitivity (LFS) is used to accurately identify and compress redundant channels, while a token-merging approach, assisted by Low-Frequency Energy (LFE), is introduced to reduce tokens. Through joint channel and token compression, the proposed method reduces the FLOPs of ViTs by over 50
Action coordination in human structure is indispensable for the spatial constraints of 2D joints to recover 3D pose. Usually, action coordination is represented as a long-range dependence among body parts. However, there are two main challenges in modeling long-range dependencies. First, joints should not only be constrained by other individual joints but also be modulated by the body parts. Second, existing methods make networks deeper to learn dependencies between non-linked parts. They introduce uncorrelated noise and increase the model size. In this paper, we utilize a pyramid structure to better learn potential long-range dependencies. It can capture the correlation across joints and groups, which complements the context of the human sub-structure. In an effective cross-scale way, it captures the pyramid-structured long-range dependence. Specifically, we propose a novel Pyramid Graph Attention (PGA) module to capture long-range cross-scale dependencies. It concatenates information from various scales into a compact sequence, and then computes the correlation between scales in parallel. Combining PGA with graph convolution modules, we develop a Pyramid Graph Transformer (PGFormer) for 3D human pose estimation, which is a lightweight multi-scale transformer architecture. It encapsulates human sub-structures into self-attention by pooling. Extensive experiments show that our approach achieves lower error and smaller model size than state-of-the-art methods on Human3.6 M and MPI-INF-3DHP datasets.
Recent methods using diffusion models have made significant progress in human image generation with various control signals such as pose priors. In portrait generation, both the accuracy of human pose and the overall visual quality are crucial for realistic synthesis. Most existing methods focus on controlling the accuracy of generated poses, but ignore the quality assurance of the entire image. In order to ensure the global image quality and pose accuracy, we propose Knowledge-Based Global Guidance and Dynamic pose Masking for human image Generation (KB-DMGen). The Knowledge Base (KB) is designed not only to enhance pose accuracy but also to leverage image feature information to maintain overall image quality. Dynamic Masking (DM) dynamically adjusts the importance of pose-related regions. Experiments demonstrate the effectiveness of our model, achieving new state-of-the-art results in terms of AP and CAP on the HumanArt dataset. The code will be made publicly available.
Parse graphs have been widely used in Human Pose Estimation (HPE) to model the hierarchical structure and context relations of the human body. However, such methods often suffer from parameter redundancy. More importantly, they rely on predefined network structures, which limits their use in other methods. To address these issues, we propose a new context relation and hierarchical structure modeling module, RMPG (Refinement Module based on Parse Graph). RMPG adaptively refines feature maps through recursive top-down decomposition of feature maps and bottom-up composition of sub-node feature maps with context information. Through recursive hierarchical composition, RMPG fuses local details and global semantics into more structured feature representations, thereby improving HPE accuracy. RMPG can be flexibly embedded as a plug-in into various mainstream HPE networks. Moreover, by supervising sub-node features map, RMPG learns the context relations and hierarchical structure between different body parts with fewer parameters. Extensive experiments show that RMPG improves performance across different architectures while effectively modeling hierarchical and context relations of the human body with fewer parameters. The code will be found at https://github.com/lushbng/RMPG.