
Current proactive defense mechanisms, though effective, predominantly concentrate on impeding deepfake models rather than regulating them. In light of the pervasive demand for deepfake creation for legitimate purposes, we introduce a proactive deepfake control framework based on a “whitelist” mechanism. This framework conditionally governs deepfake output through the utilization of a key-protected watermark, guaranteeing robustness throughout the entire process. The scheme proposes an encrypted embedder and an encrypted extractor. The former ensures that only the data owner can embed watermarks on their own facial images, while the latter safeguards against watermark leakage while facilitating watermark extraction and matching. Subsequently, the deepfake model is retrained to exclusively accept facial images that contain the specified watermark. Furthermore, a three-stage training strategy is proposed to bolster robustness and ensure watermark traceability. Experimental findings underscore the efficacy and rationality of our approach in regulating deepfake generation. Additionally, the results demonstrate the impressive performance of our scheme in terms of image quality and robustness.
Content authentication is a significant problem of the research of image security. This article exploits visual and structural feature maps to construct a new image hashing algorithm for content authentication. The visual feature map is extracted by the luminance contrast model. Since the visual feature map can reflect image regions of human visual attention, the features of the tampered areas are captured in the hash. In addition, the structural feature map is extracted by the dual-tree complex wavelet transform. Next, the block Krylov singular value decomposition is applied to the visual feature map and the structural feature map for constructing feature vectors. Finally, feature vector distances are calculated and encrypted to produce hash. Extensive experiments are performed on the public databases to demonstrate effectiveness. The results demonstrate that the proposed algorithm can correctly detect 89.83% similar images and 93.93% forged images under the optimal threshold. The comparative results show that the proposed algorithm outperforms some baseline algorithms in content authentication.
Recent advancements in immersive communication technologies have facilitated the exploration of novel modalities of social interaction, utilizing platforms that offer a spectrum of representations from simplified avatars to photorealistic volumetric and point-cloud reconstructions. This study focused on how these different representations and degrees of immersion impact communication processes and collaborative task performance. The primary objective was to examine potential variances in performance and interaction quality across different systems during a basic communication task. To achieve this, we employed Meta Horizon Workrooms as a three degrees of freedom (3DoF) system, VR2Gather as a six degrees of freedom (6DoF) system, and the MS Teams meeting application as a comparative benchmark. A user study was conducted whereby participants engaged in a charade game, alternating between the roles of “mimic” and “guesser.” Throughout the experimental session, data were collected to assess comfort, presence, task load, and interaction quality, with performance quantified by the number of words guessed per minute. The findings indicated no significant differences in workload, presence, or simulator sickness across platforms; however, significant differences were observed in audio–visual quality, communication effort, and enjoyment. Performance metrics revealed that participants exhibited the highest task performance within MS Teams, followed by VR2Gather, and subsequently Horizon Workrooms. These findings underscore the influence of embodiment fidelity and system familiarity on communication quality within immersive environments. We discuss system-level constraints and propose considerations for the design of future social extended reality (XR) platforms that effectively balance realism with usability and performance.
3D Gaussian Splatting has become a main technique for fast 3D scene reconstruction and editing, leveraging an efficient and flexible explicit representation for high-fidelity real-time rendering. However, the quality of the point clouds used to initialize Gaussians remains a key factor that limits fine-grained geometry reconstruction. To address this limitation, we propose an Iterative Spatial Decomposition (ISD) framework that bridges dense geometric priors from Multi-View Stereo (MVS) with Gaussian Splatting. ISD mitigates the mismatch between dense MVS point clouds and Gaussian sparsity by iteratively partitioning the scene into voxels of adaptive granularity and performing density-aware point assignment. Building on ISD, we introduce Hierarchical Geometric Prior Sampling (HGPS) to substantially reduce redundancy in MVS point clouds while preserving critical details, thereby providing a more robust geometric foundation for reconstruction. We further develop Hierarchical Geometry-aware Initialization (HGI), which uses a voxel-radius-based adaptive parameter initialization and replaces the iterative KNN-based procedure with batch computation, enabling efficient and robust Gaussian initialization. Additionally, we propose a Hierarchical Geometry-aware Densification (HGD) method. By dynamically identifying over-reconstructed or under-reconstructed regions through voxel constraints, HGD enhances detail reconstruction quality while controlling storage overhead. Extensive experiments on Mip-NeRF360, Tanks & Temples, and Deep Blending demonstrate significant improvements in rendering quality, achieving state-of-the-art LPIPS performance. These results indicate that our approach effectively alleviates deficiencies in the geometric priors of initial point clouds and recovers richer geometric details.
Rubbing served as a crucial cultural carrier, transferring characters from steles onto paper using ink, thereby facilitating greater dissemination and longer-term preservation. However, prolonged corrosion, inclement weather, and other factors cause various rubbing deteriorations. Therefore, an automated rubbing image restoration method is highly desirable to reduce the complexity of manual restoration procedures. In particular, when we refer to restoring a rubbing image, we mean recreating the original text on the rubbing image to reflect its original appearance. Existing methods, such as image denoising, can address minor deterioration but cannot handle severe deterioration. Image inpainting methods attempt to handle large areas of deterioration but generally struggle to produce contextually consistent strokes due to a lack of control in the restoration process. Additionally, currently there is no reliable method for rubbing character recognition. In this paper, we propose a novel two-stage rubbing image restoration method that performs well on different levels of deterioration. In the first stage, we integrate visual and contextual information for text recognition in rubbing images, including those that are heavily deteriorated. The second stage includes a multi-prior conditional latent diffusion model for rubbing image restoration, utilizing the character identities recognized in the first stage and undeteriorated characters as priors to guide the restoration process. Extensive experiments have demonstrated the effectiveness of our method compared to existing approaches.
Video Quality Assessment (VQA) technology is of significant importance for improving video transmission, storage, and processing. Although Convolutional Neural Networks (CNNs)-based and Transformer-based methods have achieved significant progress, they still suffer from some drawbacks. The previous methods treated video data as independent samples, thereby neglecting the close or distant relationships between different quality levels and consequently constraining the model’s discriminative ability; meanwhile, most existing fusion strategies utilize fixed architectures that lack the ability to adapt to the data for optimal integration, resulting in insufficient utilization of spatio-temporal information and limited expression capabilities of the fused features. To address the above issues, this article proposes a VQA method with In-Batch Contrastive Learning and Two-Phase Feature Fusion (IBCL-VQA). Firstly, the spatial features are extracted through two branches, which not only preserve the global semantics but also focus on the local regions. The temporal characteristics are obtained through a pre-trained video recognition model. Secondly, we propose an in-batch contrastive learning mechanism which, through the principles of maximizing intra-class similarity and minimizing inter-class similarity, combined with a dynamically adjusted penalty strategy, models the correlation between video quality levels. Thirdly, a two-phase feature fusion strategy, consisting of the Gated Spatio-temporal Attention Unit (GSTU) and the Adaptive Fusion Cell (AFC), is further proposed. The former achieves spatio-temporal feature fusion through dynamic weight allocation, and the latter adaptively integrates the features from the two branches based on data characteristics. Finally, a regression module outputs the quality score. Experimental results on five real-world VQA datasets demonstrate the superior performance of the IBCL-VQA. Furthermore, the strong generalizability is verified through cross-database testing. The code and pre-trained weights will be publicly available at: https://github.com/BoHu90/IBCL-VQA .
Few-shot medical image segmentation (FSMIS) has become one of the potential solutions for limited annotated medical image analysis. However, the realistic multiple domains of medical data demands the FSMIS models generalizing across domains. Thus, the cross-domain few-shot medical image segmentation (CD-FSMIS) is introduced, and we propose the Adversarial Prototypical Perturbation (APP) model which employs the adversarial learning strategy in the prototype learning process for gaining the domain robust prototypes. Specifically, the method consists of two components: adversarial signal formulation (ASF) and interactive prototypical attack (IPA). The ASF module collects perturbations from the gained gradients from both of the intra-class variation measurement loss and the inter-class variation measurement loss, and the IPA module aims to impose gained perturbations on the prototypical representation construction process with the two stages of interactive attacking manner. Additionally, a local-imbalance aware whitening loss is designed to resist the shift-sensitive local components for further enforcing to learn the domain robust prototypical representation in the IPA module. Extensive experiments are conducted on three cross-domain medical imaging datasets, and the results demonstrate that our model outperforms the state-of-the-art few-shot medical image segmentation methods. The code is available at https://github.com/YazhouZhu19/APP .
Traditional weakly supervised video anomaly detection (WSVAD) tasks typically rely on coarse-grained frame-level labels for training. Although this approach reduces annotation costs, it results in weak semantic understanding and spatial localization capabilities due to the absence of fine-grained annotations, hindering precise pixel-level anomaly detection and localization. Thanks to the success of vision-language models (VLMs), e.g., CLIP, recent approaches leveraging large VLMs focus on exploiting their strong semantic understanding capabilities, but they typically feed only keyframes or short video segments into the models, without supplying sufficient prior contextual information (e.g., contextual frames around anomalies, zoomed-in anomaly regions, and detailed anomaly descriptions), which restricts the models’ capability for fine-grained anomaly understanding and precise localization. More recently, a few methods leveraging VLMs, attempt to achieve training-free spatial anomaly localization by fusing patch-level visual features with simple textual features. However, these methods employ simplistic textual descriptions, lacking deep semantic comprehension of anomalies, leading to coarse localization results with significant irrelevant background noise. To address these issues, we propose STPrompt \(++\) , a novel weakly supervised spatio-temporal video anomaly detection and localization method based on VLMs. In our work, we systematically leverage preliminary coarse localization regions derived from anomaly scores as spatial priors, together with contextual frames around keyframes, zoomed-in views of suspected anomalous regions, and refined textual descriptions of anomalies. This comprehensive prompting mechanism guides the VLMs toward deep semantic comprehension of video anomalies, enabling accurate pixel-level spatial localization. The proposed STPrompt \(++\) requires no additional training and significantly enhances the precision of anomaly understanding and localization through a carefully designed multi-round and multi-modal prompting mechanism. Extensive experiments on two widely used WSVAD benchmarks, UCF-Crime and UBnormal, show that our method achieves state-of-the-art spatial localization performance and competitive temporal anomaly detection results. Notably, on UCF-Crime dataset, our approach improves spatial localization accuracy (in TIoU) by 5.61% over the current best method (from 23.90% to 29.51%), underscoring its superior capabilities in precise anomaly localization and semantic understanding.
Deep neural networks are vulnerable and susceptible to adversarial attacks. Audio adversarial examples impose acoustically imperceptible perturbations to clean audio examples, fooling classification models into producing incorrect results. Transferability is a critical property of audio adversarial examples, making black-box attacks applicable in practice and attracting increasing interest. Despite recent studies achieving transferability across models within the same domain, they consistently fail to achieve transferability across different domains. Given that time-domain and frequency-domain models are the two predominant approaches in audio classification, we observe that adversarial examples generated for one domain demonstrate significantly constrained transferability to the other. To address this limitation, we first consider an Inter-Domain Ensemble (IE) strategy, which fuses outputs from both domains to get an ensemble loss, optimizing adversarial examples to converge toward a common adversarial space among both domains. However, we further observe that simply averaging outputs from both domains causes adversarial examples to be more transferable to one domain, while reducing transferability to the other compared to single-domain attacks. Therefore, we propose a novel Adaptive Inter-Domain Ensemble (AIE) attack, which dynamically optimizes the contributions of both domains through adaptive weighting, improving the overall cross-domain transferability of audio adversarial examples. Extensive evaluations on diverse datasets consistently demonstrate that AIE outperforms existing methods, establishing its effectiveness in enhancing adversarial transferability across domains. Our code is available at https://github.com/unclelongheu/Audio_Adversarial_Example .
Multimodal large language models (MLLMs) have achieved significant advancements in multimodal understanding, reasoning, and interaction. However, they still suffer from hallucination, where the generated text often deviates from the factual content of the input image. To mitigate this issue, prior studies have primarily employed direct preference optimization (DPO) for human preference alignment. However, these approaches treat all textual words equally, neglecting the varying significance of individual words in grounding text generation to image content. This limitation hinders fine-grained semantic alignment and consequently constrains their effectiveness in hallucination suppression. To address this limitation, we propose a vision-guided lexical DPO method, called VGL-DPO. Specifically, we quantify the significance of words in positive preference data based on their relevance to the visual input and dynamically assign different weights to different words during training. This facilitates more precise optimization by emphasizing critical words that contribute to factual grounding. Additionally, we leverage the importance differences between high-significance words in positive and negative preference data to adaptively adjust the weight of the negative preference loss. This dynamic reweighting mechanism further refines the model’s ability to suppress hallucinated content while reinforcing factual accuracy. Extensive experiments across various models demonstrate that our method outperforms existing state-of-the-art methods in reducing hallucination and enhancing factual accuracy.
In real-world scenarios, multimodal sentiment analysis faces significant challenges, particularly in cross-scenario generalization. Existing works fail to effectively deal with the variability in evaluation frameworks and modality combinations, which results in poor transfer performance across different application contexts. In this article, the cognition-driven adaptive semantic decoding framework (CASDF) is proposed to realize an evaluation system and modality-independent multimodal sentiment analysis. Specifically, the adaptive modality association module is proposed to construct the adaptive modality mapping space, which allows us to dynamically adapt to arbitrary modality combinations. This indeed breaks through the limitation of the modality number and effectively deals with the modality gap. Furthermore, similar to the human hierarchical cognition (“perception-concept-decision”), the evaluation system progressive alignment module is presented to establish the unified evaluation system. This consists of the perception, concept, and decision analysis, which contributes to the adaptive cross-task analysis from the discrete sentiment space to the continuous sentiment space. The above joint analysis of the evaluation system and modality number indeed leads to the more flexible and generable multimodal sentiment semantic decoding paradigm. The experiments demonstrate that our sentiment semantic analysis network can achieve state-of-the-art performance.
To enhance shares visual quality and security, meaningful secret image sharing relies on pre-input cover images to endow shadow images with interpretable semantics. However, the recently proposed schemes often yield shadows with mediocre visual quality and compromised security, such as vulnerability to statistical analysis or information leakage. Generative SIS (GSIS) introduces image generation or other operations, either to generate high-quality shadows or to eliminate the need for pre-input covers. Our prior \((2,2)\) -GSIS generated meaningful shares without covers but incurred non-critical leakage and did not support lossless reconstruction. Grayscale image colorization, being a widely adopted image processing operation, offers a promising route for GSIS by enriching semantics through chrominance synthesis. We introduce a colorization-driven GSIS. Chrominance components are shared via a \((k,n)\) -threshold SIS. Near-neutral chrominance from color templates provides structural priors that guide the synthesis of share pixels. The generated chrominance supersedes the template values and directly participates in colorization. This dynamic constraint departs from the linear modification paradigm of cover-based schemes, yielding shares that are visually natural and semantically preserved, without information leakage, and enabling lossless recovery from any \( k \) of \( n \) shares. Theoretical analysis and experiments validate the effectiveness and advantages of the framework.
Cross-modal federated learning is constrained by bandwidth and on-device compute. We present Mobiflip: a minimalist strategy that freezes a lightweight backbone and communicates only a channel-wise \(1\times 1\) scaling adapter appended to the image branch. Guided by the Information Bottleneck, we prove that under common distributional and linear-encoder surrogates, per-channel scaling attains the linear optimum; coupled with the directional geometry of (Mobile)CLIP, the adapter is, in first-order approximation, an optimal preconditioner of the cosine-similarity space—preserving discriminative directions while compressing redundancy and suppressing inter-client drift. We adopt MobileCLIP as a mobile-friendly backbone to jointly minimize compute and communication. On CIFAR-10/100 and medical imaging, a single aggregation already yields stable Bacc; each round transmits only about 0.7% of backbone parameters with \(>\!\!92\%\) reduction in communication. Compared with recent federated multimodal/large-model methods, Mobiflip maintains—or even improves—accuracy under ultra-low communication.
Traditional fracture diagnosis relies heavily on the experience of clinicians and the interpretation of medical imaging. In complex cases, the inefficiency of manual interpretation often leads to misdiagnosis or missed detection, underscoring the need for automated segmentation techniques. A major challenge in calcaneal fracture image segmentation lies in the blurred and irregular boundaries of fractures, coupled with the scarcity of high-quality annotated data. To address these issues, this study independently constructs the first dataset specifically dedicated to Calcaneal Fracture segmentation, termed CalFrac. This dataset, collected from Ruijin Hospital in Shanghai, comprises CT scans of calcaneal fractures from 139 patients, along with corresponding pixel-level annotated ground truth segmentation masks. In addition, we propose the Calcaneal Fracture segmentation-Edge detection Network (CFE-Net), a multi-task CNN-Transformer hybrid architecture that employs a dual-branch structure to jointly perform fracture segmentation and edge detection. The main segmentation network adopts an encoder–decoder design to localize the fracture region, while the edge detection branch extracts boundary information and refines the segmentation via cross-branch feature interaction. Experiments on the CalFrac dataset compare CFE-Net with eight state-of-the-art methods. CFE-Net achieves superior performance across all evaluation metrics, demonstrating its advantages in both region integrity and boundary delineation. We have released the dataset and code at https://github.com/esdszdx0/CalFrac-Dataset .
Recent progress in Multimodal Large Language Models (MLLMs) has enabled GUI agents that interpret interface screenshots and execute natural-language instructions. However, their increasing role in human–computer interaction raises security concerns. Existing work shows that MLLMs are highly sensitive to interface perturbations, yet most attacks assume unrealistic capabilities such as direct access to model-ready inputs. In real deployments, attackers cannot control the screenshot pipeline or the non-differentiable preprocessing operations preceding inference, rendering many prior attacks ineffective. We introduce RAA , a r ealistic a dversarial a ttack framework that models the complete screenshot-to-prediction pipeline. To address gradient blockage from non-differentiable preprocessing (e.g., resizing, quantization), we develop a dual-branch differentiable preprocessing module that restores gradient flow while remaining faithful to the actual inference path. To improve robustness against query phrasing changes, we further incorporate a task-level semantic variation mechanism that jointly optimizes over paraphrased task variants. Experiments on Qwen2.5-VL-3B/7B-Instruct and GUI-Owl-7B show that RAA consistently outperforms prior attacks, improving attack success rates by an average of 30% under localized perturbations and a 10.7% under global perturbations, while maintaining strong effectiveness against standard defenses such as resizing, compression, and noise. Ablation studies highlight the roles of perturbation region and semantic diversity. Our framework establishes a practical threat model for GUI agents and offers a foundation for future robustness and security research.
Recent research in text-guided video editing aims to extend image-based editing models to video domains. A significant challenge in this transition is ensuring temporal consistency across frames. However, existing methods often exhibit limited editing accuracy when processing prompts associated with motion, such as “ floating ” or “ moving .” Our analysis indicates that this limitation arises from inaccurate attention maps corresponding to motion-related prompts. To address this, we introduce the Motion-to-Attention (M2A) module, explicitly integrating motion information for enhanced video editing precision. Specifically, we first convert optical flow extracted from the video into a comprehensive motion map. Optionally, users can specify directional information to refine motion map extraction further. The proposed M2A module incorporates two complementary techniques: “ Attention–Motion Swap ,” which directly substitutes the imprecise attention map of motion prompts with the extracted motion map, and “ Attention–Motion Fusion ,” which adaptively enhances attention maps based on the correlation with the motion map using a carefully selected Fusion metric. Experimental validation demonstrates that incorporating our M2A module into existing text-to-video editing frameworks significantly improves both quantitative performance metrics (CLIP-Acc, Masked PSNR, BRISQUE) and qualitative visual quality. Extensive experiments and comparative studies confirm the superior editability and robustness of our method over current state-of-the-art approaches. Comprehensive results are publicly available at https://currycurry915.github.io/Motion-to-Attention/ .
Proxy hashing methods have attracted increasing attention in cross-modal retrieval, because they are able to learn the mapping of different modalities into a common low-dimensional hash space by leveraging global proxies. However, existing approaches typically suffer from two limitations: (1) They solely utilize prior labels to capture the global semantic information, and hence lack the ability to explore necessary fine-grained semantic information to bridge the modality gap effectively. (2) They seldom consider the guidance information of intrinsic semantic similarity on proxy-centered space, and thus fail to leverage the similarity relations among instances sufficiently. To mitigate these limitations, we propose RAPH, a novel Relation-Aware Proxy Hashing framework that learns the semantic relations between different modalities and different semantic levels to enhance the discriminative capability of hash codes. Specifically, we first propose a Local Semantic Interaction (LSI) module based on masked language modeling to achieve the interaction of multi-modal fine-grained semantic features. Second, a relation-aware hashing learning scheme is designed to simultaneously explore the intrinsic semantic relationships and the global semantic information based on proxies. This is achieved by minimizing the reconstruction error between the multi-modal affinity matrices derived from learned features and the cross-modal similarity matrix of the hash codes. The proposed framework is able to learn more discriminative hash codes and achieves superior performance to many baselines on three public datasets.
In human society, companion animals are no longer simply ”family members” or ”emotionally dependent individuals.” They are entering, alongside us, a web of relationships reshaped by diverse computational technologies. In this context, this paper innovatively proposes a three-stage analytical framework, encompassing monitoring, interpretation, and empowerment. This framework, through the philosophical lens of ”significant otherness,” reveals a clear evolutionary path in how technology mediates the human-animal bond, moving from a one-way observation to a two-way, co-constituted interaction. The specific contributions of this work are as follows: (1) the establishment of a core-task-centered approach for classifying relevant technologies; (2) a systematic exploration of the role of AI-driven multimodal technologies within this framework; and (3) the integration of the philosophical concept of ”significant otherness” with cognitive engineering principles to enrich the theoretical underpinnings of the evolving human-animal relationship. Additionally, this paper also addresses key ethical, methodological, and technical challenges and opportunities inherent in the future application of AI within the field of Animal-Computer Interaction (ACI), offering new insights and perspectives for both academic research and practical application.
Weakly supervised video anomaly detection (WVAD) aims to locate events or behaviors that deviate from normal patterns in untrimmed videos using video-level labels. Recent studies typically utilize supplementary modalities to assist anomaly detection. However, these methods suffer from two main issues: (1) The limitations of long-duration anomaly event temporal modeling. The model struggles to consistently maintain key information, resulting in the forgetting phenomenon, which affects the tracking of the event’s overall dynamic evolution and complicates anomaly event analysis and understanding. (2) The multi-modal fusion strategy is insufficient, particularly when there is temporal inconsistency between visual and audio information, causing the model to overlook key information, directly affecting the accurate detection and recognition of anomalous events. To address these issues, we propose a visual-guided long-term temporal context learning network (LTCLNet). The network consists of three key components: a cross-modal interaction module, a multi-modal fusion module, and a visual-guided parameter optimization strategy. First, to address the forgetting issue in long-duration anomaly detection, we designed a cross-modal interaction module. The key part of this module is the establishment of a cross-matrix mechanism. This mechanism achieves bidirectional temporal guidance across modalities. It allows the temporal modeling of each modality to dynamically integrate information from the other modality. This enables the model to continuously track the dynamic evolution of the event. The tracking is facilitated through shared temporal information between the visual and audio modalities. Secondly, to fully exploit the complementary characteristics between different modalities, we introduced a novel temporal reversal integration method in the multi-modal fusion module. This method reverses the feature sequences of each modality to enhance the model’s perception of temporal dynamic changes. By fusing the modality features before and after reversal, the shared temporal structure between modalities is strengthened, improving the model’s ability to capture anomalous information. Additionally, our proposed visual-guided parameter optimization strategy trains a parallel visual modality network as a semantic anchor, ensuring that the model stays aligned with a semantically stable and structurally clear visual flow during the learning process, thus ensuring stability and semantic coherence in the training. Extensive experiments on datasets such as XD-Violence demonstrate that our method significantly outperforms existing approaches, particularly achieving notable improvements in the accuracy and stability of long-term anomaly detection. Our code is publicly available at https://github.com/ibliever/LTCLNet .
This study examines PUPS, a representative Bitcoin ecosystem project, to elucidate the success mechanisms of Web3 meme projects. We test three hypotheses: (H1) community sentiment and social media virality constitute the fundamental drivers of meme asset valuation; (H2) core participants accumulate positions at low prices and distribute at peak valuations; (H3) meme diffusion is predominantly driven by internal imitation, significantly outweighing external marketing effects. Applying event study methodology, social network analysis, and the Bass diffusion model to social media and on-chain data, our findings support all hypotheses, revealing a ”propagation–sentiment–trading” pathway. We identify a distinctive ”community fingerprint” comprising 348 original holders and 5,036 6-core addresses, characterizing them as both community stabilizers and hype catalysts. This pattern illustrates the paradox of ”economic recentralization” within technically decentralized systems. Paradoxically, the founder's public assertion that ”everything will eventually go to zero” evolved into a cultural ritual that reinforced community consensus. This study concludes by proposing a ”meme financialization” framework, offering novel perspectives for understanding ”Attention as Capital”, ”Consensus as Value”, and ”Narrative as Asset” in Web3 ecosystems.