
Background The use of artificial intelligence (AI) in business operations is increasing rapidly. However, existing research provides limited systematic insight into how AI integration reshapes business models. This is especially unclear when big data capabilities within digital transformation support AI. This lack of understanding hampers both theoretical development and practical decision-making for strategic AI deployment. Methods Drawing on dynamic capability theory and the resource-based view, this study employs structural equation modelling using a pooled cross-sectional dataset of 847 multinational corporations across 23 industries. The executive survey was collected in 2024 and complemented with retrospective survey information and archival firm-level indicators covering 2019–2024. The study adopts a sequential explanatory mixed-methods design, combining quantitative analysis with 312 executive interviews. Results The findings identify four pathways through which AI-driven business model innovation is positively associated with competitive advantage: algorithmic value proposition enhancement (β = 0.487, p < 0.001), intelligent ecosystem orchestration (β = 0.423, p < 0.001), predictive resource optimization (β = 0.392, p < 0.001), and autonomous competitive positioning (β = 0.356, p < 0.001). Together, these pathways explain 52.3% of the variance in competitive advantage (R2 = 0.523) and account for 75.8% of the total association between AI integration and competitive advantage. Pathway effectiveness varies across business model contexts: algorithmic value proposition enhancement shows a stronger association in B2C contexts (β = 0.612) than in B2B contexts (β = 0.389), while intelligent ecosystem orchestration shows a stronger association in platform-based models (β = 0.534) than in traditional models (β = 0.318). Conclusions The study advances dynamic capability theory by offering a fine-grained empirical analysis of AI-enabled capability deployment and its association with competitive advantage. It also offers practitioners an evidence-based framework for prioritizing AI and big data investments according to their competitive context.
The coprime sampling structures offer significant advantages in achieving large apertures and high degree of freedom (DOF), but sensor failures can severely impact their space time adaptive processing (STAP) performance. To address this issue, a data completion STAP method based on a multi-norm constrains adaptive alternating direction method of multipliers (AADMM) is proposed. The method first utilizes difference technology to construct virtual space–time snapshots and establishes a multi-norm constrained rank optimization model based on low-rank properties specifically for sensor failure scenarios, aiming to accurately reconstruct missing data and estimate the clutter covariance matrix (CCM). Subsequently, the alternating direction method of multipliers (ADMM) framework is employed to decompose the complex constrained problem into multiple sub-problems for alternating iterative solutions. To overcome the bottleneck of slow convergence in traditional iterative algorithms, an optimal adaptive iteration step-size selection strategy is introduced, which significantly enhances computational efficiency. Validation was conducted through 500 Monte Carlo trials based on simulated side-looking airborne phased array radar data with a clutter-to-noise ratio (CNR) of 30 dB and a signal-to-noise ratio (SNR) of 10 dB. Numerical results demonstrate that in sensor failure scenarios, while the DOF of the traditional STAP method drops to 4 and that of the conventional C-STAP method drops to 9, the proposed AADMM-C-STAP method consistently maintains a virtual DOF of 19, significantly enhancing the system’s clutter suppression capability and target detection robustness.
Predicting ABL1/BCR–ABL1 mutation resistance to tyrosine kinase inhibitors (TKIs) remains challenging because available mutation–drug evidence is limited, heterogeneous, and distributed across public bioactivity resources. This study presents an optimized calibrated stacked-learning framework for computational screening of ABL1/BCR–ABL1 mutation–TKI resistance. Public bioactivity records from BindingDB and ChEMBL were harmonized into a conservative binary mutation–drug pair-level dataset, with activity-derived variables excluded from model predictors to prevent target leakage. The proposed model combined automatic base-learner selection, sparse logistic stacking, probability calibration, and recall-prioritized threshold optimization. In the random stratified holdout experiment, the model achieved strong discrimination, with AUROC = 0.970 and PR-AUC = 0.922. Repeated stratified cross-validation confirmed stable performance, with mean AUROC = 0.932 ± 0.054 and mean PR-AUC = 0.869 ± 0.095. In the source-aware BindingDB-to-ChEMBL robustness experiment, performance decreased to AUROC = 0.795 and PR-AUC = 0.581, indicating sensitivity to database shift and assay heterogeneity. Benchmark analysis showed that simpler probabilistic and regularized models generalized better under cross-source evaluation. Overall, the framework provides a leakage-controlled and probability-calibrated screening approach for prioritizing resistant-candidate mutation–TKI pairs, but its outputs should be interpreted as hypothesis-generating predictions rather than clinical treatment recommendations.
Industrial gas turbines operate under dynamic conditions in which startup, steady-state, and transient operation exhibit distinct combustion characteristics. Conventional machine learning-based Predictive Emissions Monitoring Systems (PEMS) generally employ a single model trained using aggregated operational data, implicitly assuming a stationary relationship between turbine operating variables and emissions. This assumption often degrades prediction accuracy during operating transitions due to non-stationary combustion behavior. To address this limitation, this study proposes a regime-aware heterogeneous stacked ensemble framework for multi-pollutant emissions prediction. Historical industrial gas turbine data are first segmented into startup, steady-state, and transient operating regimes, after which independent predictive models are developed for each regime. The proposed framework combines XGBoost, Decision Tree (DT), and LightGBM as complementary base learners, while ElasticNet serves as a regularized meta-learner to integrate their predictions. The proposed model predicts CO, CO2, SO2, and NOx emissions and is evaluated using independent validation datasets. During startup, validation MAPE ranges from 1.38% to 5.88%, with R2 values between 0.39 and 0.95. Under steady-state operation, validation MAPE ranges from 1.23% to 8.07%, with R2 values between 0.88 and 0.97, where the largest prediction error is associated with CO while the majority of pollutant–regime models satisfy the preferred industrial target of 5% MAPE. During transient operation, validation MAPE ranges from 1.20% to 5.57%, with R2 values ranging from 0.73 to 0.98. These results demonstrate that predictive performance depends on both the operating regime and pollutant characteristics, while consistently providing strong explanatory capability across diverse operating conditions. Model interpretability is investigated through complementary sensitivity analyses using Mean Decrease in Impurity (MDI), Permutation Importance, and SHapley Additive exPlanations (SHAP), which consistently identify regime-dependent dominant process variables. Furthermore, measured inference latencies of less than 20 ms per prediction demonstrate the computational efficiency of the proposed framework for near-real-time industrial deployment. Overall, the proposed regime-aware stacked ensemble provides an interpretable, computationally efficient, and practically deployable solution for industrial multi-pollutant Predictive Emissions Monitoring Systems.
Robust single-time-point (STP) dosimetry critically depends on selecting an imaging time that yields reliable estimates of the time-integrated activity coefficient (TIAC). This study developed and evaluated a virtual-patient nonlinear mixed-effects modelling (NLMEM) framework to optimise STP imaging schedules for renal TIAC estimation, using [¹⁷⁷Lu]Lu-PSMA-617 therapy as a proof of concept. A previously published NLMEM (sum-of-exponentials structure and parameter distributions) describing renal [¹⁷⁷Lu]Lu-PSMA-617 biokinetics in 63 patients served as the generative model [1]. Based on fixed- and random-effects parameters, 500 virtual patients (VPs) were sampled, and reference TIAC values were computed analytically (rTIAC). For each candidate imaging time, renal activity measurements were simulated by applying proportional noise (7.9
Heart disease, which includes conditions such as heart failure, coronary artery disease, and ventricular fibrillation, represents a major health challenge worldwide. The increasing mortality rate from cardiovascular disease (CVD) underscores the need for accurate and effective diagnostic methods. Although past studies have investigated machine learning (ML) and deep learning (DL) algorithms, issues like data imbalance, feature selection, and model optimization remain. This research introduces a comprehensive heart disease prediction system utilizing the Cleveland and UCI datasets. The methodology involves gathering and analyzing patient data, followed by data preparation using the Interquartile Range (IQR) and Synthetic Minority Oversampling Technique (SMOTE) to handle missing values. The Deep Graph Correlation Network (DGCN) is employed to extract distinct features from the pre-processed data. The Density-based Spatial Clustering of Applications with Noise (DBSCAN) method, enhanced with fuzzy logic, is used to segment the dataset. For predicting disease, the Adaptive Walrus Optimization Algorithm (WaOA) improves the Evolutionary Attention-based Deep Long Short-Term Memory (EA-DLSTM) model. This approach achieves accuracy rates of 99.62% for the UCI dataset and 99.71% for the Cleveland dataset, with precision scores of 99.75% and 99.68%, respectively. SHAP (SHapley Additive exPlanations) values ensure the model’s transparency and interpretability by emphasizing the contributions of key features to the predictions. This integrated strategy not only achieves high accuracy but also offers an interpretable and dependable system for the early detection of heart disease, effectively bridging the gap between automated predictions and clinical validation.
Background Cancer is a major health issue caused by various chemical imbalances and genetic problems, ranking as the second leading cause of death worldwide. Approximately one in six individuals succumb to this disease. Lung cancer is among the most prevalent and lethal types, significantly affecting human health due to its high mortality rate. Objectives Early detection is vital in combating this life-threatening illness, as it greatly enhances survival rates. Computed tomography (CT) scans are key diagnostic tools for lung cancer, but their high cost can limit access and lead to varied interpretations by different observers. Moreover, analyzing these images demands significant time and expertise. Recent advancements in deep learning (DL) have accelerated lung cancer detection. Methods This paper presents a flexible and scalable feature-fusion and classification pipeline to enhance computer vision tasks by integrating features from multiple sources and utilizing deep sequential modeling. The framework combines high-level representations from the Shifted Window (Swin) transformer and handcrafted descriptors from the Histogram of Oriented Gradients (HOG), enabling a thorough fusion of semantic, structural, and texture-based information. A support vector machine (SVM) classifies by learning sequential relationships among the fused features. The proposed hybrid Swin-HOG-SVM model allows pathologists to effectively and affordably evaluate more patients, leading to improved healthcare outcomes. This optimized model assists in the early detection of lung cancer, reducing the necessary effort, time, and costs. Results Experimental results demonstrate that the proposed Swin–HOG–SVM pipeline outperforms standalone deep and handcrafted approaches across multiple evaluation metrics, achieving improved accuracy, precision, recall, and F1-score, achieving an accuracy of 99.39 %, specificity of 97.78 %, precision of 99.42 %, recall of 97.78 %, and an F1-score of 98.56 %. Conclusions This innovative classifier shows significant promise in accurately and efficiently detecting lung tumors, providing a valuable tool for oncology specialists looking to streamline the diagnostic process.
In this paper, a new hybrid conjugate gradient (CG) parameter is proposed for solving convex constrained monotone equations, with particular emphasis on compressed sensing image recovery. The proposed method introduces a novel combination parameter μk, defined through a modified Dai–Liao conjugacy condition. Unlike existing hybridizations, this formulation ensures the sufficient descent property while maintaining numerical stability, thereby guaranteeing global convergence. To evaluate the effectiveness of the method, we compare it against recent algorithms including SGCS, CHCG, and DTCG1 on both numerical test problems and image restoration tasks. The numerical experiments demonstrate that the proposed MHCG method achieves faster convergence in terms of iteration counts and computational time, while also provides higher solution accuracy. In compressed sensing image restoration, MHCG delivers superior perceptual quality, with sharper edge recovery and reduced artifacts compared to competing methods. Overall, the results confirm that the proposed hybridization offers a robust and efficient alternative for large-scale monotone equations and signal processing applications.
Underwater Wireless Sensor Networks (UWSNs) are vital for applications including environmental observation, disaster prevention, and deep-sea exploration, yet intermittent node defects severely degrade data fidelity and network reliability. This paper introduces a Distributed Extreme Learning Machine (D-ELM) framework for self-fault diagnosis in UWSNs, designed to efficiently detect intermittent sensor failures through localized learning and collaborative decision exchange among neighboring nodes. Leveraging the inherent advantages of extreme learning machines, fast training speed and strong generalization, D-ELM enables each sensor node to independently perform real time diagnostics while sharing minimal diagnostic information to maintain low communication overhead in bandwidth limited acoustic channels. The distributed architecture not only reduces computational complexity but also enhances fault isolation precision and robustness against transient defects. Experimental evaluations conducted on multi-hop underwater sensing datasets with simulated environmental disturbances demonstrate that D-ELM achieves an average diagnostic accuracy of 83.21% on unseen data, outperforming state-of-the-art classifiers such as Support Vector Machine (SVM), Random Forest (RF), Decision Tree (DT), Multilayer Perceptron (MLP), and Extra-Trees in terms of precision, F1-score, and training time. The proposed scheme provides a scalable, energy-efficient, and autonomous solution for long-term underwater monitoring, significantly improving network reliability, diagnostic accuracy, and operational lifespan in dynamic underwater environments.
This paper proposes a cooperative communication scheme based on the trust-partner mechanism to enhance the information security of cell-edge users. A linear weighting model is developed to integrate multiple social relationship metrics, including interest similarity, user association strength, social mutual assistance and user social-distance distribution, thereby deriving a comprehensive user trust value for identifying optimal relay partners. An optimization problem is formulated to maximize the secrecy rate of cell-edge users, and the approximate lower bound of the maximum secrecy rate is derived using the Difference of Convex (DC) programming. The closed-form solution expressions of the transmission power are derived by using Karush–Kuhn–Tucker (KKT) conditions. The simulation results show that compared with other schemes, the proposed scheme significantly improves the probability of success in relay selection and achieves a secrecy rate for cell-edge users that closely approaches that of the exhaustive method.
This paper presents a feature-engineered ensemble method for fault detection in medium-voltage insulator ultrasound signals. We develop a feature extraction framework that captures time-domain, frequency-domain, wavelet-domain, and envelope characteristics of ultrasonic emissions from porcelain insulators under various conditions: normal operation, artificial contamination, and simulated lightning damage. A robust signal augmentation methodology enhances classification robustness while preserving class-discriminative properties. We evaluate standard ensemble approaches (random forest, extreme gradient boosting, and light gradient boosting) and introduce two novel algorithms: EvoBagging, which employs evolutionary operators to optimize bootstrap samples, and proximal policy optimization bagging (PPOBagging), which leverages reinforcement learning through proximal policy optimization to construct ensembles via sequential decision-making. Experimental results demonstrate that our feature engineering approach significantly outperforms raw signal classification (91.67% vs. 55.00% accuracy). The proposed PPOBagging algorithm achieves the highest performance with 94.17% accuracy and a macro-F1 score of 0.9412, surpassing conventional methods. Comparative analysis against convolutional kernel transforms (standard, mini, and multi random convolutional kernels) reveals that while these methods offer competitive accuracy (up to 88.33%), they fall short of our feature-engineered ensemble approach in both performance and interpretability. Comparison against deep learning methods also reaffirms the usefulness of the proposed feature engineering technique. This work contributes practical tools for non-destructive monitoring of power distribution systems, enabling proactive maintenance strategies to prevent failures.
This study proposes an optimized method for immersive and interactive dramatic narratives, grounded in integrated design principles that center on multimodal perception, dynamic decision-making, and collaborative generation. The proposed approach establishes an adaptive decision-making mechanism driven by both emotion and behavior, thereby moving beyond the conventional optimization paradigm based on isolated modules. Specifically, the method leverages Adaptive Reinforcement Learning (ARL) and Multimodal Generative Adversarial Networks (MM-GAN). A multimodal adaptive neural network is first employed to model sequences of user actions, speech, and visual behaviors, while an integrated emotion analysis module predicts real-time emotional states, enabling accurate capture of multidimensional interaction demands. Subsequently, an ARL model incorporating Double Deep Q-Network (Double DQN) and Prioritized Experience Replay (PER) is adopted to maximize the long-term reward of user experience, thus facilitating dynamic branching decisions in the plot. Third, MM-GAN is utilized to generate personalized narrative content across text, image, and speech modalities, with a multi-task joint optimization framework ensuring collaborative training across all components. Experimental results on a self-built VR-IID dataset demonstrate that the proposed method outperforms mainstream baseline models in terms of plot coherence (BERTScore-F1 = 0.897), behavioral response accuracy (92.8%), and immersion score (4.6 points). The system achieves an average response time of 132 ms, satisfying the real-time interaction requirements of VR environments, and reaches training convergence within 58 rounds. Furthermore, this study introduces a narrative adaptability calculation model to quantify the matching degree between user states and plot content, offering a computable theoretical foundation for the intelligent optimization of interactive narratives.
The integration of Large Language Models (LLMs) into software engineering workflows has demonstrated significant potential for automating code refactoring across programming languages. However, cross-language refactoring—particularly between high-level orchestration languages (Python) and shell scripting environments (PowerShell)—presents unique security and semantic challenges. This paper presents a novel framework for LLM-refactored Python-based PowerShell command generation that addresses critical security vulnerabilities inherent in automated code translation.We evaluate the framework using a curated dataset of 8886 labeled examples (yielding 7,719 unique, deduplicated patterns via AST-based hashing and semantic filtering) mapped across ten MITRE ATT&CK categories. Using 5-fold stratified cross-validation, our analysis reveals that 21.95% of PowerShell commands contain high-risk security patterns, with Invoke-Expression (8.3%), DownloadString (5.8%), and execution policy bypasses (2.5%) being most prevalent. We demonstrate that Retrieval-Augmented Generation (RAG) techniques reduce vulnerability introduction rates by 57.9% relative to standard GPT-4o (absolute reduction: 22.3 pp; p<0.001, McNemar test) and outperform security-aware prompting (81.5% vs. 54.3% security compliance). Comprehensive comparisons against rule-based sanitizers, prior frameworks, and larger models (CodeLlama-34B, DeepSeek-Coder-V2-Lite-16B) show that our 7B-parameter framework achieves superior security outcomes with significantly reduced computational requirements.We address prompt injection risks through defense-in-depth strategies, including semantic intent classification, Unicode normalization, randomized delimiter spotlighting, and multi-turn stateful risk tracking. Furthermore, we subject the framework to a rigorous adversarial evaluation utilizing Atomic Red Team, PyRIT, and Garak benchmarks, achieving a 100% Detection Rate (DR) and reducing the Attack Success Rate (ASR) to 0% across 76 complex adversarial scenarios. We provide complete dataset and code artifacts for reproducibility (https://github.com/tamerelserwy-research/Secure-LargeLanguageModel-PS-Refactoring), following the Reproducibility Maturity Model (RMM) guidelines.
Tourism and point-of-interest (POI) recommendation is inherently intent-dependent and context-sensitive, yet many existing methods primarily rely on user–item interactions or short-term behavioral sequences, making it difficult to capture the rich semantics of natural-language travel demands. To address this limitation, we propose GUIDE, a Generative User Intent Dual-Encoder framework. We first distinguish a user’s current intent from long-term preference, and construct target-free semantic user contexts from pre-target histories and contextual metadata to avoid label leakage. GUIDE then uses an open-source large language model strictly as a semantic encoder to represent user-side intent and item-side POI content in a shared embedding space. On top of this backbone, a bidirectional contrastive alignment module improves fine-grained user–item matching, and a variational latent preference module models the uncertainty and multi-faceted nature of travel interests through latent preference sampling. This design combines efficient approximate nearest-neighbor retrieval with diverse and robust re-ranking. We evaluate GUIDE primarily on a genuine tourism/POI check-in dataset, Foursquare-TKY, and further examine its cross-domain behavior on Yelp and MovieLens, comparing it against popularity-based, collaborative, graph-based, sequential, self-supervised, POI-specific, and LLM-enhanced baselines. Under a statistically validated protocol with five random seeds, GUIDE consistently improves top-K ranking on the POI dataset, with significant gains over the strongest baseline (paired Wilcoxon signed-rank test, p<0.01), while auxiliary results on Yelp and MovieLens suggest that the semantic dual-encoder design also transfers to broader recommendation settings. Ablation, parameter-matched, efficiency, fairness, and case analyses further verify the contribution of each component, and a dedicated analysis maps each gain onto the prior finding it confirms, extends, or qualifies. Our conclusions are limited to the evaluated public datasets and are not claimed to generalize to all users or destinations. Code and preprocessing scripts are released for reproducibility.
The rapid expansion of IoT systems fuels the reuse of vulnerable code in firmware binaries, exposing IoT devices to network attacks. Current dynamic analysis, which relies on emulation and fuzzing, suffers from high latency, hindering large-scale vulnerability detection that is especially critical today. Static analysis leverages program structure and statistical features for vulnerability detection. However, existing Static approaches fail to accurately reconstruct vulnerability logic due to limited feature sets, and IoT programs are typically closed-source binary programs lacking code features, leading to high false positive and false negative rates. We propose TIR-VD, an innovative method that addresses the limitations of existing approaches, significantly improving binary vulnerability detection in IoT system. First, we convert binary sample into images using the Bin2Img algorithm and extract texture features with a lightweight deep learning model. Next, we construct a weighted input-related control flow graph (WIR-CFG) and extract semantic features by a semantic-aware deep learning framework. For vulnerability detection, we match the target binary against vulnerability samples by integrating both texture and semantic features. The experimental results indicate that TIR-VD surpasses existing open-source tools in vulnerability detection accuracy and efficiency. Additionally, we identified two 0-day vulnerabilities in binary and reported them to the vendor. TIR-VD significantly enhances vulnerability detection accuracy and efficiency in IoT firmware binaries by integrating texture and semantic features, enabling scalable and accurate analysis of closed-source IoT binaries.
Metaheuristic optimization has evolved from classical handcrafted strategies to increasingly adaptive, hybrid, and intelligence-driven systems. However, existing classification schemes remain largely static and fragmented, limiting their ability to capture the progressive transformation of optimization paradigms. This paper proposes a novel five-generation evolutionary framework (1G–5G) that systematically organizes metaheuristic development from Classical Metaheuristics (1G), Inspiration-Based Metaheuristics (2G), Hybrid Metaheuristics (3G), and Self-Adaptive Metaheuristics (4G), to Intelligent Optimization Systems (IOS, 5G). Unlike conventional taxonomies that group algorithms primarily by inspiration source or structural similarity, the proposed framework introduces an evolutionary perspective based on increasing levels of adaptivity, autonomy, and intelligence integration. The framework further conceptualizes IOS as a system-level optimization paradigm characterized by machine learning integration, autonomous decision-making, surrogate modeling, and dynamic environmental responsiveness. By synthesizing the historical progression of optimization methodologies into a unified generational model, this work provides both a structured analytical taxonomy and a forward-looking conceptual foundation for next-generation intelligent optimization research.
Obtaining optimal logical rules and maintaining the explainability of the logical rules are two crucial issues in developing a logic mining model. To address these challenges, a novel symbolic logical rule namely J-type Random 2,3 Satisfiability was proposed to represent the attribute in the datasets. The proposed logical rule consists of second and third order clauses where the order of the clause is randomly generated. A well-built logical rule will be integrated with a Discrete Hopfield Neural Network and applied to real-life datasets. To achieve the selection of important attributes, Topological Data Analysis was utilized to capture the structural characteristics of the dataset. After the attribute selection, the permutation operator will be applied during the training phase to increase the solution space of the logic mining. In this context, the expanded search space will ultimately yield the optimal induced logic. The proposed logic mining model was evaluated using numerous real-life datasets from various fields of study. Experimental results demonstrate that the proposed model outperforms all state-of-the-art logic mining models, achieving an average accuracy of 0.8375 and superior performance in precision, F1-score, and Matthews’ Correlation Coefficient.
Modern healthcare networks increasingly rely on interconnected medical devices, electronic health records, and real-time monitoring systems, making them highly vulnerable to sophisticated cyberattacks that can directly compromise patient safety and clinical operations. Ensuring secure and reliable network behavior in such environments requires intrusion detection systems that not only achieve high detection accuracy but also minimize false alarms to reduce alarm fatigue while enabling early threat identification. The primary challenge in healthcare cyberattack detection lies in simultaneously achieving high precision, near-perfect attack recall, low false positive rates, and timely detection, under highly imbalanced and dynamic network traffic conditions. Conventional machine learning and deep learning based intrusion detection methods often lack probabilistic risk calibration and are therefore limited in their suitability for safety–critical healthcare monitoring environments. To address these limitations, a risk-aware hybrid framework is introduced that integrates GraphSAGE-based unsupervised node embeddings for modeling relational network behavior, Gaussian Mixture Models (GMMs) for probabilistic anomaly scoring, and conformal calibration to enforce statistically guaranteed false alarm control. Network flows are represented through graph-induced behavioral embeddings that capture protocol service state interactions, while anomaly likelihoods are transformed into interpretable risk scores with controlled alarm rates. Extensive experiments conducted on healthcare-relevant traffic derived from the UNSW-NB15 dataset show promising detection performance, achieving 96 % overall accuracy, 99.97 % attack detection recall, 0.9864 ROC-AUC, and a false positive rate of only 5.05 % under conformal risk constraints. The framework further enables early attack detection, behavioral drift monitoring, and interpretable risk attribution across network entities. To the best of current knowledge, this work presents an early healthcare-oriented risk-calibrated cyberattack detection framework combining GraphSAGE, GMM-based anomaly scoring, and conformal calibration. Evaluation on UNSW-NB15 supports its promise, while validation on real medical IoT and hospital network data remains necessary for clinical translation.
In field ecological monitoring, search and rescue, and disaster prevention missions, carpet search strategies are typically employed. The importance and quantity of tasks often change dynamically, requiring multi-robot systems to possess the capability to flexibly respond to dynamic variations. However, traditional task allocation methods are predominantly designed for static small-scale environments and struggle to meet the demands of large-scale dynamic task scenarios, nor are they optimized for carpet search strategies. To address this issue, this paper formulates carpet search as the sequential execution of multiple tasks within elongated regions and proposes a Two-stage Adaptive Sampling Rolling Task Reallocation (TASRTRA). In TASRTRA, to resolve the problem of high-importance tasks being overlooked due to random probability sampling, an adaptive probability sampling mechanism is constructed that dynamically adjusts the sampling probability distribution to increase the sampling probability of high-importance tasks, thereby ensuring a balance between task value maximization and computational efficiency. Considering the time-varying characteristics of task importance during execution, the marginal cost calculation method is improved by introducing a dynamic importance variation factor, enabling real-time perception of task importance evolution and dynamic reconstruction of task allocation schemes in multi-robot systems. Furthermore, to mitigate the elevated collision risk caused by intersecting task paths, an intersecting path exchange strategy is introduced to optimize the task allocation scheme, reducing collision probability while enhancing overall system performance. Experimental results demonstrate that TASRTRA exhibits superior performance under dynamic and various task-scale scenarios. Compared with traditional methods, it achieves an average reward improvement of 7.62% and an average runtime reduction of 15.15%. The method adapts effectively to various dynamic task conditions, demonstrates strong robustness, and significantly optimizes task allocation efficiency.
Recognizing human activities and detecting anomalous behaviors in crowded scenes from unmanned aerial vehicle (UAV) imagery remains a challenging problem due to severe occlusions, scale compression, viewpoint distortion, and weak appearance cues, which limit the effectiveness of conventional ground-based and 2D appearance-driven methods. To address these challenges, this paper proposes CrowdVoxel-Net, a geometry-aware multimodal framework for UAV-based activity recognition and anomaly detection, in which a novel aerial silhouette extraction network, SA-Net, is introduced to enable reliable human foreground isolation under complex drone viewpoints. Building upon the extracted silhouettes, pseudo-3D human structure is reconstructed via monocular depth estimation and encoded using voxelized orthogonal depth projections to capture spatial and volumetric cues in a view-consistent manner. Long-range and occlusion-resilient spatiotemporal motion dynamics are modeled using CoTracker, while robust and domain-invariant semantic representations are extracted using a Vision Transformer–enabled DINO self-supervised backbone. In addition, local structural cues are strengthened using a compact keypoint–driven descriptor and a fast cross-frame matching module, improving geometric alignment under challenging aerial viewpoints. These complementary geometric, semantic, spatial, and temporal representations are fused and modeled using a Transformer-based classifier for joint activity recognition and anomaly detection. The proposed framework is evaluated on four challenging benchmarks, including UAV-Human, Drone Action, ShanghaiTech, and CADG, achieving classification accuracies of 53.3% on the full 155-class UAV-Human dataset, 85.2% on its 15-class subset, 91.2% on Drone Action, 95.2% on ShanghaiTech, and 85.7% on CADG, consistently outperforming existing state-of-the-art methods and demonstrating strong robustness and generalization under diverse UAV surveillance scenarios.