Sustainable Digital Infrastructure plays an increasingly important role in strengthening regional economic resilience and promoting inclusive growth. However, its dynamic effects on nascent entrepreneurship remain underexplored. Treating China’s smart city pilot policy as a quasi-natural experiment, this study uses data on newly registered enterprises in 288 Chinese cities from 2005 to 2019 and a multi-period difference-in-differences (DID) model to examine this relationship. The results show that Sustainable Digital Infrastructure significantly promotes new venture creation. Owing to environmental dynamism—manifested in creative destruction, strategic wait-and-see behavior, and learning-curve effects—its entrepreneurial impact is initially volatile before becoming significantly and persistently positive. The industry-level analysis reveals heterogeneous effects across sectors: the largest estimated effect occurs in wholesale and retail trade, while the estimated effect in scientific research and technical services strengthens over time. Mechanism analyses provide suggestive evidence for three channels: technological innovation, human capital development, and digital finance. The effects are also more pronounced in eastern China and smaller cities. This study provides empirical evidence and practical guidance for emerging economies seeking to promote new venture creation through digital infrastructure. It conceptualizes “sustainability” in the SDI context as the capacity of digital infrastructure to support enduring and inclusive economic development, rather than as a property of the infrastructure’s environmental performance.
With the advancement of energy conservation and emission reduction initiatives at ports, electric container trucks are being increasingly adopted in port collection and distribution logistics. The charging strategy for electric container trucks directly affects their transportation routing. This paper considers three different charging strategies for electric container trucks and develops the optimization models for their transportation routes. An improved genetic algorithm is designed to solve the models. Based on real-world data from Logistics Company Q, an empirical analysis is conducted. The comparison analysis results show that the lowest transportation cost occurs under the rapid battery-swapping mode, followed by the partial charging strategy, while the full charging strategy incurs the highest transportation cost. The sensitivity analysis results show that, under the full charging mode, total distribution costs increase significantly as the delivery time window penalty coefficient rises. Under the partial charging mode, when the charging amount is between 60% and 80% of the battery capacity, the total distribution cost decreases with the increase of the charging amount. However, when the charging amount is between 80% and 100% of the battery capacity, the total distribution cost also exhibits increasing trend. Under the battery-swapping mode, as the rated battery capacity gradually increases from its original level, the total transportation cost gradually decreases. However, it introduces operational challenges, especially time conflicts, which may have a negative impact on the service quality of the transportation company.
Video Anomaly Detection (VAD) aims to automatically identify anomalous events in videos that significantly deviate from normal behavioral patterns. Self-supervised learning motivates models to learn effective features from unlabeled data by designing proxy tasks. However, existing approaches often rely on coarse-grained modeling, focusing mainly on global sequence order or holistic scene structures, which may limit their ability to capture subtle motion changes or localized anomalies. Therefore, this paper proposes a self-supervised learning framework combined with fine-grained spatio-temporal proxy tasks to extract key features more accurately. For the temporal branch, we design a time interval prediction task: given a fixed middle frame and randomly sampled frames from both sides, the model predicts their temporal intervals relative to the center frame, thereby modeling the dynamic patterns of behavior. To enhance temporal modeling capabilities, we introduce a multi-head self-attention mechanism to capture inter-frame dependencies in the input sequence. The spatial branch employs a noise classification task inspired by diffusion models, where varying levels of noise are added to image patches, and the model predicts the corresponding noise levels. This encourages learning of local appearance features and patch-level sensitivity to perturbations. Our method is trained in an end-to-end manner and does not rely on pre-trained models. Experiments on three benchmark datasets demonstrate stable performance: the method achieves AUC scores of 98.6% on UCSD Ped2, 91.7% on CUHK Avenue, and 83.7% on ShanghaiTech. These results suggest that the proposed approach can generalize well across different scenes, perspectives, and types of anomalous behavior.
Fully capitalizing the complementary advantages of multiviews while minimizing their limitations remains a key challenge in point cloud semantic segmentation. Traditional multi-view approaches face two critical issues: (1) overlooking the adverse effects of inter-modal data distribution discrepancies; (2) ineffective fusion mechanisms fail to maximize the potential of each view. To address these challenges, we propose PVContrast, a novel point-voxel fusion network, employing a contrastive learning strategy to optimize inter-modal synergy and feature representations. For the first issue, we introduce a Cross-Modal Quadruplet Alignment Loss based on a Semantic-Driven Feature Decomposition strategy, which soft-groups features by labels into anchor, positive, and negative samples at the semantic group level. Additionally, a Bi-Directional Quadruplet Loss ensures intermodal semantic consistency and alignment through bidirectional modeling. To our knowledge, this is the first application of this technique to multimodal indoor segmentation. For the second issue, we propose a Two-Stage Fusion strategy. First, a Cross-View Multiscale Feature Propagation (CMFP) mechanism leverages cross-self-attention for effective feature interaction. Second, an Adaptive Fusion Refinement (AFR) module fuses and enhances structural representations. Our method achieves 72.5% mIoU on S3DIS and 75.8% on ScanNetv2, outperforming most state-ofthe-art algorithms and demonstrating superior performance in multimodal point cloud segmentation.
Anomaly detection and localization techniques have seen widespread adoption in industrial manufacturing due to their efficiency and strong generalizability. However, although unsupervised methods do not require anomalous samples during the training stage, they still face many challenges in actual deployment: on the one hand, nominal samples in training data may contain minor defects or noise; such contamination can mislead the learning process, causing the model to miss defects and thereby increasing false-positive or false-negative rates. On the other hand, the scarcity of anomalous samples often leads to model overfitting, while the complex and variable morphology of anomalies renders precise identification and localization difficult. To address these issues, we propose a Prototype-corrected Multi-Scale Feature Reconstruction Network (PMSR). This network leverages a purified multi-scale prototype normality library as guidance to accurately reconstruct normal regions. PMSR encompasses two core components: a multi-scale prototype correction mechanism that refines training data via the prototype library, and a composite masking strategy that enhances the model's robustness to both local structural and global pattern perturbations. In addition, a dedicated feature-matching module maximizes the alignment probability between features to further boost model precision. Experimental results on three representative industrial datasets-MVTec AD, VisA, and MPDD-demonstrate that PMSR achieves state-of-the-art detection performance, confirming its effectiveness and applicability for industrial defect inspection.
Video anomaly detection (VAD) identifies events that deviate from normal patterns, but faces challenges due to scarce anomalies, high annotation costs, and overfitting. To address these issues, we propose a VAD framework that combines spatio-temporal pseudo-anomaly generation with a dual-branch contrastive discrimination strategy. In the temporal dimension, motion rhythms are explicitly disrupted by randomly sampling frames at varying intervals, enabling the model to capture temporal inconsistencies. A multi-head self-attention mechanism is further employed to model inter-frame dependencies. In the spatial dimension, inspired by diffusion-based progressive perturbation, graded noise is injected into local regions to generate appearance pseudo-anomalies of varying intensity. This design allows the model to learn both subtle and severe appearance deviations. Compared with conventional Gaussian noise-based perturbation strategies, this progressive design provides more robust and expressive supervision for modeling complex appearance anomalies. Furthermore, we design a dual-branch contrastive discrimination loss for temporal and spatial modeling, ensuring effective separation of normal and abnormal patterns in the feature space. Our method achieves consistent and superior performance on three representative datasets: UCSD Ped2, Avenue, and ShanghaiTech, with Area Under the Curve(AUC) scores of 98.5%, 90.5%, and 79.3%, respectively. The results demonstrate that the proposed method exhibits strong robustness and generalization across various scenes, viewpoints, and anomaly types.
This study examines pricing and green innovation decisions in an e-retail platform supply chain where a national brand (NB) competes directly with a platform-owned private brand (PB). Using a Stackelberg game framework, we analyze four scenarios: No innovation, NB-only innovation, PB-only innovation, and bilateral innovation. The key findings are as follows. Network externalities consistently benefit all supply chain members and enhance green innovation levels across scenarios, but their impact on retail prices is scenario-dependent. The NB manufacturer achieves its highest profit under NB-only innovation, whereas the e-retail platform prefers bilateral innovation, which constitutes the unique Nash equilibrium. This misalignment creates a green innovation prisoner's dilemma: Although bilateral innovation maximizes the total supply chain profit, the manufacturer earns less than under NB-only innovation. Moreover, standalone green innovation by either party yields higher green levels than bilateral innovation. These findings suggest that coordinating bilateral green innovation may require profit-sharing mechanisms, and that innovation strategies should be tailored to the market conditions.
Exploiting the geometric directional anisotropy of point clouds to better model local structures is an essential challenge for 3D point cloud scenario understanding. While existing isotropic feature aggregation or static multiaxial decomposition strategies have either failed to take this property into account, resulting in diluted anisotropy representation, or have only considered the modeling of relationships among points. To address this problem, we propose GMSFormer, which improves perception on local structures by explicitly associating point cloud structures with axial patterns. The core innovations of our work reside in the Geometry-aware Multi-Structured (GMS) Attention and Geometry Prompt modules. The former additionally decomposes a cube window into three orthogonal planes (XY, XZ, and YZ, each representing a geometric pattern of the missing axial direction) in the point cloud native space, subsequently utilizing self-attention to explicitly model the geometric properties in different axial directions innovatively. The latter leverages geometric properties and implicit features to quantify planes importance, rendering the model consistently emphasizing the effective planes to enhance axial modes of the cube. Moreover, we also design the Dilated Gated (DG) Refiner to detail multi-source features from various axes and the cube through multi-scale dilation convolutions and a learnable gating mechanism. Finally, we integrate the Kolmogorov-Arnold Network (KAN) into GMSFormer blocks for the first time, realizing differential processing of multi-source features with learnable activation functions, which explores the application of KAN in the point cloud domain. Our approach achieves 74.0% mIoU on S3DIS and 76.7% mIoU on ScanNetv2, significantly outperforming most state-of-the-art algorithms.
Video anomaly detection (VAD) plays a critical role in various engineering applications such as safety monitoring and traffic management. Current deep learning-based methodologies for VAD predominantly focus on frame reconstruction and prediction strategies. However, insufficient modeling of long-range spatio-temporal regularities and latent motion transitions in video sequences limits the effectiveness of both paradigms. Inspired by video coding and decoding, we formulate unsupervised video anomaly detection as a restoration task in which intermediate frames are recovered from the start and end frames of an object-centric clip. This formulation requires the model to infer multiple missing frames from sparse temporal observations, thereby encouraging stronger modeling of object-centric spatio-temporal context, including motion continuity, local appearance evolution, and temporal dependency patterns. To support this task, we introduce a Contextual Memory Bank (CMB) with a two-stage memory-coordination learning strategy to store latent motion-context representations from normal sequences and retrieve relevant spatio-temporal priors when only boundary frames are available. We further employ memory query decomposition for local context matching, which improves the consistency between the recalled context and the current input. Extensive experiments on benchmark datasets demonstrate the effectiveness of the proposed framework for video anomaly detection.
The conflicting imperatives of collective supply chain resilience and data sovereignty create a critical 'data silo paradox'. Whilst Federated Learning (FL) offers a technical solution by enabling collaborative modelling without raw data exchange, its organisational mechanisms remain under-theorised. Drawing on Dynamic Capabilities Theory and Transaction Cost Economics, this research conceptualises Federated Learning Maturity (FLM) and examines its impact on Supply Chain Resilience (SCR). Based on a structural equation model derived from 1,199 questionnaires, the research confirms that FLM acts as a strategic antecedent to resilience. Mediation analysis reveals three distinct pathways: External Organisation Trust, Information Sharing, and Operational Flexibility. Crucially, the results resolve the 'trust paradox', demonstrating that technical maturity fosters relational trust, which serves as the most potent mediator. The findings suggest that cryptographic governance functions as a credible commitment, reducing transaction costs and enabling a new paradigm of 'non-disclosing collaboration'. This research provides empirical evidence that resilience in the algorithm economy is driven by the quality of 'learned information' rather than raw data transparency.
Anomaly detection has deeply penetrated various aspects of industrial manufacturing due to its high efficiency and wide applicability. However, in practical application scenarios, it still faces multiple challenges. On the one hand, most current methods compare the samples to be tested with pre-collected normal samples. This approach, which relies on fixed reference samples, significantly increases the difficulty of aligning images. On the other hand, when logical anomalies are highly similar to the background distribution of normal samples, detection models often struggle to accurately distinguish between normal and anomalous regions, leading to missed detections or misjudgments. To address these issues, we propose the Intrinsic Prototype Guided Reconstruction Network (IPG-FRN). IPG-FRN has two core components: First, it dynamically mines the most representative intrinsic prototypes (IPs) directly from the image to be tested through a center-edge strategy and an intrinsic prototype extractor, eliminating the reliance on the normality of an external training set and avoiding the alignment difficulties caused by appearance and position differences in reference samples in traditional methods. Second, the global IP memory library updated iteratively during the training phase screens and replaces the extracted intrinsic prototypes through an IP projection strategy, ensuring that the retained intrinsic prototypes are highly stable in terms of normality and semantic information, thereby providing reliable and powerful guidance for subsequent normal feature reconstruction. Experimental results show that IPG-FRN achieves advanced detection performance on MVTec AD, VisA, and MPDD, fully demonstrating its effectiveness and applicability in industrial defect detection tasks.
Multi-attribute features such as appearance, optical flow, and pose have been widely applied in the field of Video Anomaly Detection. However, existing studies often overlook the semantic consistency between these features, which limits the potential relationships among different attributes and hinders the effective utilization of this information in anomaly detection. To address this issue, we propose a video anomaly detection framework based on Semantic Consistency and Multi-Attribute Feature Complementarity (SC-MAFC). Specifically, we design a three-branch encoder-decoder network to separately encode and predict appearance, optical flow, and pose features. By comparing the consistency differences between appearance and other attribute features at normal and anomalous moments, we can use these differences as the basis for anomaly detection. To better capture these differences, we introduce a Spatial-Channel Feature Complementarity Module (SCFCM), allowing different attribute features to complement each other and helping the model more accurately understand the semantic consistency of normal events, thus improving its ability to recognize the semantic consistency in normal events. Additionally, to further enhance detection performance, we introduce a memory module to degrade the reconstruction quality of features at anomalous moments, by which large errors are generated in future frame predictions, making anomalies more pronounced and easier to detect. It was evaluated on three benchmark datasets: UCSD Ped2, Avenue, and ShanghaiTech. The method proved effective, achieving AUC scores of 99.4% on UCSD Ped2, 93.9% on Avenue, and 84.6% on ShanghaiTech, demonstrating its robustness in various scenarios. The code is publicly accessible at the following link: https://github.com/jzt-dongli/SC-MAFC.
Introduction: International shipping faces the dual imperative of decarbonization and the maintenance of safety and service reliability. However, evidence concerning the conditions under which artificial intelligence (AI) generates verifiable carbon benefits remains fragmented. Methods: This review examines 479 English-language articles and reviews retrieved from the Web of Science Core Collection and Scopus and published from 2016 to 23 August 2026 through bibliometric analysis, science mapping, auxiliary document-level thematic coding, and full-text synthesis of six representative reviews and one perspective. A post hoc domain-validation sensitivity analysis used two independently specified deterministic rule sets to test whether broad search terms altered the main conclusions. Results: Publication output accelerated markedly after 2022, with 277 papers (57.83%) published during 2022–2025 and a further 137 records already indexed in the partial year 2026. The two screening rules agreed on 96.87% of records (Cohen’s kappa = 0.753). A conservative sensitivity subset of 437 records, obtained through a strict rule-based title-abstract screen and removal of one retracted and one withdrawn record, reproduced the principal temporal, source-journal, and leading-keyword patterns. Machine learning remained the most frequent keyword, while recent studies increasingly addressed deep learning, ship energy efficiency, port operations, federated learning, and energy management. Discussion: Based on these findings, the review advances an evidence-informed AI-to-Carbon Value Chain (AICV) conceptual synthesis comprising data observability, model credibility, decision executability, system coordination, and carbon verification. This synthesis is interpretive rather than a validated causal framework. Future research should prioritize carbon-ready benchmarks, calibrated physics-informed and causal models, human-in-the-loop field evaluation, network-level coordination, and auditable well-to-wake assessment. Review registration and appraisal: This review was not registered, and no formal protocol was prepared. Because no effect-size synthesis was undertaken, formal study-level risk-of-bias, reporting-bias, and certainty assessments were not applied. Funding: The review received no external funding.
To address the limitations of traditional two-stream networks, such as inadequate spatiotemporal information fusion, limited feature diversity, and insufficient accuracy, we propose an improved two-stream network for human action recognition based on multi-scale attention Transformer and 3D convolutional (C3D) fusion. In the temporal stream, the traditional 2D convolutional is replaced with a C3D network to effectively capture temporal dynamics and spatial features. In the spatial stream, a multi-scale convolutional Transformer encoder is introduced to extract features. Leveraging the multi-scale attention mechanism, the model captures and enhances features at various scales, which are then adaptively fused using a weighted strategy to improve feature representation. Furthermore, through extensive experiments on feature fusion methods, the optimal fusion strategy for the two-stream network is identified. Experimental results on benchmark datasets such as UCF101 and HMDB51 demonstrate that the proposed model achieves superior performance in action recognition tasks.
Point cloud multiplane mechanisms effectively capture rich geometric information from different projections, but this approach lacks explicit consistency constraints across planes, and the loose coupling mechanism of concatenation and summation leads to inconsistent cross-plane features. To address this challenge, we propose a contrast plane transformer, an innovative multiview network that utilizes contrastive learning to achieve cross-view feature coupling. Specifically, we first decompose the point cloud window into three orthogonal planes (XY, XZ, and YZ) and use the attention and graph-based local enhancement module to extract features from each plane. Subsequently, to realize cross-plane feature coupling, we propose a patch-level plane alignment mechanism. Specifically, we partition the planes into a certain number of patches, designate one patch from one plane (XY) as an anchor, and designate the corresponding patches from the other two planes as positive samples while utilizing other randomly selected patches as negative samples to construct an efficient and robust comparison mechanism. We not only achieve cross-plane consistency but also significantly reduce the computational complexity of point-level comparison. Finally, considering the imbalance of geometric information provided by different planes, we propose a dynamic plane fusion mechanism that can automatically adjust the weights of different planes according to the local geometric features of the input point cloud, thus realizing adaptive information fusion. Our method achieves 73% and 75.5% mIoU on the S3DIS and ScanNetv2 datasets, respectively, significantly outperforming most algorithms, which fully demonstrates the effectiveness of our algorithm.
The multimodal fusion of point cloud data is a challenging task in 3D computer vision, particularly in fine object segmentation. In this study, we propose an innovative unified multimodal LiDAR segmentation network, named Align and Blend (A2Blend for short), which skillfully integrates three representative forms of point clouds: point view, voxel view, and range view. The core innovation of A2Blend lies in addressing two primary tasks: Align and Blend. For the “Align” task, we have carefully designed a learnable cross-modal association module, whose core is a unique “Cross-Modal Triplet Alignment Loss” mechanism. To the best of our knowledge, this is the first application of such a mechanism in the field of LiDAR segmentation. This mechanism applies principles from contrastive learning, promoting the tight clustering of similar semantic groups both within and across modalities, while increasing the separation between dissimilar groups by expanding the distance between similar and dissimilar sample clusters in the feature space. This significantly enhances the discriminative power and representational capacity of the feature embeddings. For the “Blend” task, we propose a fusion strategy that integrates intragroup self-attention within modalities and intergroup self-attention across modalities. This approach combines key concepts from standard self-attention and cross-self-attention mechanisms to achieve more comprehensive multimodal fusion. The segmentation performance on two large-scale outdoor datasets and one indoor dataset surpasses that of most state-of-the-art algorithms.
As global e-commerce expands, efficient cross-border logistics services have become essential. To support the evaluation of logistics service providers (LSPs), we propose HD-CBDTOPSIS (Technique for Order Preference by Similarity to Ideal Solution with heterogeneous data and cloud Bhattacharyya distance), a hybrid multi-criteria group decision-making (MCGDM) model designed to handle complex, uncertain data. Our criteria system integrates traditional supplier evaluation with cross-border e-commerce characteristics, using heterogeneous data types—including exact numbers, intervals, digital datasets, multi-granularity linguistic terms, and linguistic expressions. These are unified using normal cloud models (NCMs), ensuring uncertainty is consistently represented. A novel algorithm, improved multi-step backward cloud transformation with sampling replacement (IMBCT-SR), is developed for converting dataset-type indicators into cloud models. We also introduce a new similarity measure, the Cloud Bhattacharyya Distance (CBD), which shows superior discrimination ability compared to traditional distances. Using the coefficient of variation (CV) based on CBD, we objectively determine criteria weights. A cloud-based TOPSIS approach is then applied to rank alternative LSPs, with all variables modeled using NCMs to ensure consistent uncertainty representation. An application case and comparative experiments demonstrate that HD-CBDTOPSIS is an effective, flexible, and robust tool for evaluating cross-border LSPs under uncertain and multi-dimensional conditions.
Given the importance of time varying term structure relationships, the connectedness of various time charter contract with the ship price and demolition market varies under different market conditions. This paper investigates dynamic connectedness among different shipping segment of dry bulk sector utilizing the time and frequency domain TVP-VAR connectedness method covering global crisis and geopolitical shocks. Our empirical results indicate strong asymmetric interconnectedness across all segments particularly spillover effect is relatively higher in Panamax vessel. We find evidence of strong long-term spillover impact followed by short-term except Capsize vessel. When analyzing market influence, the average spot freight market leads the system from short to long-term followed by time charter contracts (12-month) across all shipping markets. Our analysis shows that the newbuilding market equally sensitive to external shocks in the short and long-term period while the secondhand market is relatively sensitive to other shocks in the long-term period. In particular, the secondhand ship price of Panamax and Handymax receives strong spillover from the short to long-term period. Moreover, we observe that the demolition market play critical role as positive transmitter of different sizes specially in the Panamax and Handymax segment. Our results enhance understanding of decision-making in cash flow, ship and asset management by showing asymmetric proactive and risk-tolerant behavior in the short and long-term periods.
Video anomaly detection (VAD) aims to identify rare events in videos that deviate from normal behaviors. Due to the scarcity of anomalous samples, most methods in this field rely on datasets composed solely of normal behaviors to learn standard patterns and identify significant deviations as anomalies. However, these approaches have several limitations. Primarily, the exclusion of anomalous samples during training can lead to the neglect of highly discriminative spatio-temporal characteristics intrinsic to normal behaviors. Moreover, the variability of spatial scales at which anomalies occur is often overlooked by current methods. To address these challenges, an innovative method that learns multi-grained spatio-temporal features has been introduced. This approach begins by generating appearance and motion pseudo-anomalies through simple geometric transformations. A dual-branch network model for spatio-temporal decoupling, comprising separate spatial and temporal branches, is then proposed. Each branch tackles two proxy tasks: (1) assessing the presence of appearance or motion anomalies and (2) pinpointing the locations of these anomalies within specific spatial and temporal contexts. The model is designed to mine multi-grained features from both macroscopic and microscopic perspectives. Experimental results on the Avenue, ShanghaiTech, and UCSD Ped2 datasets demonstrate the efficacy of our method, achieving frame-level AUC scores of 91.5%, 80.2%, and 98.2% on each dataset, respectively.
We aim to identify individual actions and group activities from videos. Existing solutions primarily model spatial and temporal relationships based on the positional information of individual participants, but there is a clear limitation: due to the relative motion of the camera and the discrete nature of pixels, the pixel coordinate sequences of characters in videos often fail to accurately reflect their true location information in the real physical world, and no research team has yet proposed an effective solution to this problem. To overcome this challenge, we propose a method based on structure from motion (SFM). We utilize SFM, combined with deep learning which is called sub-bundle adjustment for secondary correction of three-dimensional pose keypoint information as part of the input of the model, and employ transformer and Swin Transformer architectures to richly represent local and global spatial features; then, we designed a multi-channel feature fusion module. This module integrates local color features (RGB), global RGB features, and pose features through an adaptive weighting and fusion approach, thereby providing a more comprehensive feature representation for subsequent group behavior recognition. Finally, followed by long short-term memory to capture temporal feature dependencies at the video level. In our empirical study, we explored multiple approaches to fuse these feature representations and confirmed the complementary advantages among them. On the Volleyball dataset, our method achieves 93.9% under multi-class accuracy (MAC) indicator and 94.3% under mean per class accuracy (MPAC). On the Collective dataset, our method achieves 95.8% under MAC metric and 95.9% under MPAC. Experimental results demonstrate the importance of motion recovery and deep learning-based secondary reconstruction result correction. More importantly, our proposed method has achieved significant superior performance on recognized benchmarks for group activity recognition.