
Smart water conservancy is a crucial measure for promoting the efficient utilization of water resources and the modernization of water governance,and has become a core direction for the transformation and upgrading of the water conservancy sector.During the current construction process,it still faces problems such as heterogeneous multi-source data,insufficient coupling between mechanisms and data models,and limited level of business intelligence,which makes it difficult to fully support the demand for scientific decision-making in complex water conservancy scenarios.Large lan-guage model(LLM),relying on its advantages in knowledge representation,semantic reasoning and natural interaction,shows broad potential in empowering smart water conservancy.From the three core dimensions of technical system,busi-ness paradigm and system integration,this paper systematically reviews the research progress of LLM in the water conser-vancy field,sorts out the core technical systems of LLM empowering smart water conservancy,including knowledge en-hancement,parameter optimization,interaction optimization,multi-modal fusion and agent development,and analyzes the dataset construction and evaluation methods of LLM in the water conservancy field.It deeply examines the transforma-tion of business paradigms driven by LLM in core scenarios such as basin flood control,water resource allocation,and construction,operation and maintenance of water conservancy projects.It also summarizes three types of system integra-tion models and technical pathways for business implementation,namely knowledge enhancement-driven,professional model coupling,and agent collaboration.Finally,the development trends of the in-depth integration of LLM and smart water conservancy are summarized and prospected,aiming to provide theoretical reference and practical experience for the in-telligent and smart transformation of the water conservancy sector.
Ancient texts are characterized by their conciseness,flexible syntactic structures(such as inversions and ellipses),and significant semantic differences between classical and modern Chinese.Consequently,existing models trained on modern corpora exhibit limited performance in recognizing character relationships within ancient Chinese texts,particularly in capturing deep semantic associations.To address this challenge,this paper proposes an entity pair enhanced relationship recognition model for Chinese ancient books(EPERRM).Firstly,the GuwenBERT pre-trained model is employed to extract deep semantic representations from ancient texts.A multi-dimensional entity pair feature extraction module is then constructed to capture semantic features of entity pairs,incorporate part-of-speech tags to identify semantic roles,and compute relative positional features to enhance structural awareness.To effectively integrate these features,a gated channel weighting algorithm is designed to dynamically screen and weight semantic and part-of-speech features.Furthermore,a multi-head cross attention mechanism is introduced to enable deep interaction between global text representations and fused entity pair features,thereby facilitating accurate relationship classification.Experimental results on the Twenty-Four Histories corpus demonstrate that the EPERRM model achieves an F1Macro score of 90.28%,outperforming the GuwenBERT baseline.The model also exhibits superior performance in Zero-shot and Few-shot settings compared with large language models such as Qwen and DeepSeek.Ablation studies confirm that the integration of semantic,part-of-speech,and relative position features significantly enhances model performance,with entity pair features contributing more critically than plain text features.
The detection of fake news is critical to mitigating the spread of misinformation.A primary challenge lies in extracting supplementary information that supports the assessment of news authenticity from inherently limited textual content.Traditional approaches are constrained by their limited capacity for deep textual reasoning and the effective integration of multi-dimensional contextual data,making it difficult to leverage external information related to the news content.This paper proposes a novel text-based fake news detection method that combines large language models with fine-tuned small language models.The large language model performs multi-level analysis and semantic mining on both the news content and associated comments,generating enriched analytical features.These features,along with the original news content,are then fed into a compact,task-specific small language model,which excels in classification performance due to its spe-cialization and efficiency.By delegating complex reasoning and information extraction to the large language model and utilizing the small language model as a precise classifier,the proposed framework achieves a synergistic effect that enhances overall detection accuracy.Experimental results demonstrate that this collaborative approach achieves an accuracy of 95.18%on real-world datasets,offering a highly accurate and practically viable solution for detecting fake news in plaintext.
In visually degraded scenarios,text image restoration aims to enhance both the visual quality and content integ-rity of images,and its core tasks can be categorized into two main types:image quality restoration and image content com-pletion.Among these,image quality restoration focuses on reversing the physical degradation process under the premise of a complete character sequence.Image content completion addresses situations where characters are partially or entirely missing,requiring the generation of visually consistent and semantically reasonable content based on contextual informa-tion.The primary objective of both tasks is to improve the overall quality of the images,thereby enhancing the readability,recognition accuracy,and information integrity of text images.In light of this,this paper provides a comprehensive and systematic summary of the research progress in this field.Firstly,for the task of image quality restoration,the review focuses on three primary directions,including super-resolution,geometric distortion correction,and image enhancement,with a discussion of their representative methods.Secondly,regarding the task of image content completion,the focus is placed on character-level completion generation,with a comparative analysis of two mainstream approaches:those based on glyph structure priors and those based on semantic context reasoning.Subsequently,the commonly used datasets and eval-uation metrics in the field are systematically summarized and compared,thereby revealing the existing limitations of cur-rent data resources in terms of scale,authenticity,and task adaptability.Finally,this paper provides the prospect of future research trends in the field,with an emphasis on the potential value of text image restoration in practical applications such as cultural heritage preservation,document digitization,and intelligent text processing,thereby offering valuable references for subsequent related practical applications.
Distributed key generation technology has garnered significant attention in recent years.It enables participating parties to collaboratively generate public-private key pairs without relying on a trusted third party,utilizing threshold secret sharing.This makes it a fundamental building block for decentralized technologies like threshold signatures and blockchain.Traditional distributed key generation protocols typically assume a synchronous network,making them vulnerable to attacks in practical asynchronous environments.To address this challenge,recent research has proposed asynchronous distributed key generation schemes,which can effectively initialize threshold cryptosystems without requiring a global clock.However,these schemes are heavily dependent on intricate underlying cryptographic knowledge,along with their associated setups and constraints,thereby rendering the overall system architecture complex and difficult to construct.This paper presents a modular analysis of the key components of asynchronous distributed key generation technology,starting from its architectural perspective.The definition and basic settings of asynchronous distributed key generation technology is clearly defined,and its system architecture is described,which is divided into a random generation module and a consensus generation module according to the execution process and functions.This paper classifies and gradually describes the asynchronous secret sharing technology relied upon by the randomness generation module and the asynchronous Byzantine consensus technology relied upon by the consensus generation module,and systematically introduces the relevant underlying cryptographic knowledge.This paper introduces the application scenarios of asynchronous distributed key generation technology,summarizes and looks forward to its current important challenges and future research directions.
With the growing concern over energy consumption and heat dissipation in data centers,temperature control has become a pivotal factor in maintaining operational efficiency and extending equipment lifespan.Accurate temperature prediction plays a vital role in this process.However,traditional methods primarily rely on physical models and rule-based frameworks,which struggle to adapt to complicated and changeable operational environment.In recent years,data-driven methodologies have emerged as a promising alternative.This review systematically reviews data-driven research on data center temperature prediction,focusing on core dimensions such as hierarchical structure,time scales,and cooling systems.It refines core concept definitions,analyzes data types and feature distributions,and evaluates their influence on model performance.It synthesizes research progress in machine learning,deep learning,and hybrid modeling approaches,highlighting their contributions to enhancing predictive accuracy and robustness.The review indicates that deep learning models are particularly effective in capturing complex spatiotemporal dependencies,while hybrid models offer comple-mentary advantages by integrating the strengths of multiple paradigms.Potential research directions are discussed,includ-ing multimodal data fusion,adaptive learning,and edge-intelligent temperature prediction frameworks.
Brain tumors are highly fatal neurological diseases,for which precise diagnosis is crucial to improving prognosis.Traditional imaging techniques are often limited in detecting small lesions,characterizing tumor heterogeneity,and facilitating quantitative analysis.Deep learning technology provides crucial technical support for overcoming such limitations.This paper systematically reviews deep learning research and applications in brain tumor diagnosis.Firstly,widely used public datasets are summarized,and key evaluation metrics along with their clinical significance are outlined.Secondly,the review focuses on three core tasks.For tumor detection,it systematically compares two-stage and one-stage approaches.For tumor segmentation,it summarizes methods based on 2D slices,3D volumes,and multi-contrast sequences.For tumor classification,it analyzes models from single-network and multi-network architectural perspectives.The evolution of mainstream methods is traced,the defining characteristics of different technical routes are highlighted,and the integration strategies and application prospects of large language models in these domains are discussed.Furthermore,the current status of clinical translation for deep learning models is evaluated.Core challenges are analyzed,and potential pathways toward clinical adoption are outlined.Finally,limitations of existing methods are summarized.Future research directions are proposed from three perspectives:data,model improvement,and clinical application.
In response to the challenges of detecting social media misinformation under conditions of implicit semantic expression,complex content structures,and diverse propagation behaviors,this paper investigates a unified detection framework that integrates the reasoning capabilities of large language models with multi-source feature representations.The objective is to improve misinformation detection accuracy through the joint modeling of reasoning-based knowledge and multi-dimensional features,while enhancing the transparency of the model's decision process.A large language model-enhanced dual-gated fusion framework(LLM-EDGF)is developed,in which a structured prompting procedure is employed to guide the large language model to generate reasoning-based knowledge features related to factual consis-tency and logical coherence of the text,which are then jointly modeled with deep semantic representations extracted by BERT.In addition,user profile features and propagation behavior features are incorporated to characterize the attributes of information publishers and diffusion patterns.At the feature fusion stage,a two-stage gating mechanism is designed at both the"knowledge-semantic"and"content-social"levels to achieve dynamic weighting of multi-source features and suppression of noisy information,and a lightweight classifier is finally applied for misinformation detection.Experimen-tal results on multiple public datasets and a self-constructed Weibo dataset demonstrate that the proposed framework out-performs comparative models in overall detection performance and exhibits more stable recognition capability in complex misinformation scenarios such as partial fabrication of facts and incoherent narrative structures.Furthermore,the intro-duced reasoning-based knowledge features provide traceable information for the model's decision process,supporting the applicability of the proposed approach to social media misinformation detection tasks.
To address the issues of suboptimal modality feature extraction,inadequate feature fusion,and insufficient information interaction in multimodal sentiment analysis,this paper proposes a multimodal sentiment analysis method based on multi-head high-low frequency attention and mutual information calculation.Firstly,a multi-head high-low frequency attention mechanism is designed for cross-modal feature fusion,enabling parallel extraction and integration of local fine-grained features and global contextual representations.This effectively combines multi-scale features and enhances the model's capability to extract both local and global emotional feature information.Secondly,a multimodal feature interactive enhancement strategy is introduced,which employs cross-modal bidirectional interaction and comple-mentary learning mechanisms to achieve thorough fusion and multi-level feature complementarity within and between modalities.Additionally,a mutual information calculation task is incorporated to strengthen the model's ability to learn shared information among features,thereby improving the effectiveness of cross-modal interaction.The proposed method is experimentally validated and tested on the public datasets CMU-MOSI and CMU-MOSEI.The binary classification accuracies achieved are 90.40%/88.48%and 87.81%/84.59%,respectively,which are 1.38/1.75 and 0.69/1.20 percentage points higher than the baseline model.The F1-scores reach 90.36%/88.40%and 87.78%/84.93%,representing improvements of 1.46/1.86 and 0.39/1.06 percentage points over the baseline.The Pearson correlation coefficients are 0.865 and 0.801,showing increases of 0.004 and 0.002 compared with the baseline.The comparative results demonstrate that the proposed model outperforms many state-of-the-art methods and can effectively enhance the accuracy of multimodal sentiment analysis.
The pervasive diffusion of Deepfake techniques has posed severe threats to public security and social trust.Videos forged by cascading multiple synthesis models or post-processing modules increasingly exhibit frame-different manipulative traces,a phenomenon that remains under-explored.Existing detectors,built on the homogeneous-forgery assumption,struggle to deliver reliable and fine-grained judgments when unknown and mixed forgery strategies appear across frames.To address this limitation,this paper formalizes the analysis of such samples as the frame-wise heterogeneous Deepfake(FH-Deepfake)problem and introduces the ETMA(EfficientNet-Transformer with multi-attention aggregation)framework.ETMA casts FH-Deepfake detection as a fine-grained multi-label learning task,in which a hierarchical attention mechanism is employed to capture frame-specific forgery patterns.Concretely,(1)texture enhancement coupled with spatial attention highlights subtle tampered regions;(2)multi-attention aggregation fuses multi-level features to produce discriminative frame representations amenable to temporal modelling;(3)a multi-label frame-level forensic transformer encodes long-range inter-frame dependencies while aligning semantic categories,enabling simultan-eous identification and localization of multiple forgery types.To evaluate the approach,this paper constructs and releases the first FH-Deepfake benchmark.Extensive experiments demonstrate the superiority of ETMA,and subsequent face-restoration tests driven by the detection results further corroborate the practical necessity of investigating heteroge-neous deepfakes.
Machine vision-based icing transmission line image segmentation is one of the primary methods for estimating ice thickness.However,high-voltage transmission lines are typically located in remote and open areas.During winter icing events,the pixel-level distribution of ice-covered transmission lines often closely resembles their surrounding environ-ment.Images of ice-covered conductors often exhibit two extreme characteristics:low illumination and highlight satura-tion.These conditions result in overly smooth textures and sparse features in the images,posing challenges for effective segmentation.This study adapts the segment anything model(SAM)adapter for the specific task of ice-covered conductor segmentation by redesigning the prompt encoder.This study proposes a multi-scale search mechanism for salient features in the frequency domain of ice-covered images,enhancing the ability of the model to extract texture information through dense searches of sparse features.This study introduces a local zero-mean salient feature prompting module,which reduces feature similarity by zero-meaning normalization,amplifies feature disparities,and strengthens the edge-capturing capability of the model to guide contour segmentation.The proposed algorithm is validated on a custom ice-covered conductor segmentation dataset and open conductor images.Experimental results demonstrate that the proposed method accurately captures ice-covered conductor contours and correctly classifies internal pixels.The average values of technical indicators such as mIoU,mDice,F1,recall and precision in icing transmission line image segmentation reach 0.907,0.949,0.910,0.927 and 0.897 respectively,outperforming mainstream ice-covered conductor segmentation algorithms.
As one of the key devices in urban infrastructure systems,surveillance cameras play an important role in urban traffic management,information support,and public security monitoring.However,the proliferation of unauthorised sur-veillance camera installations exacerbates privacy violations and even posing threats to public security.Surveillance camera deployment exhibits configuration imbalances,with redundant installations in certain areas while others remain covered by blind spots.Methods for detecting surveillance cameras using street view imagery can provide technical support to address these issues.Nonetheless,surveillance cameras in street-view images typically have few pixels,complex background and varied shapes,which leads to low detection accuracy and high miss rate in the tasks of detection.To address this issue,this paper proposes a street-view surveillance camera detection algorithm called LH-RTDETR(LDConv and HiLo enhanced real-time detection transformer)based on RT-DERT(real-time detection transformer).Linear deformable convolution(LDConv)is introduced to improve downsampling,which adaptively changes the convolution sampling shape while main-taining a stable number of parameters,thereby enhancing the flexibility of information extraction and effectively reducing false detections.The high-low frequency attention(HiLo)mechanism is incorporated into the encoder to independently process high and low frequency components,thereby strengthening the model's understanding of both local and global contexts and reducing the missed detection.A high-resolution feature extraction layer P2 is added to capture more detailed information,thereby improving detection accuracy.Experiments conducted on the self-constructed surveillance camera detection(SCD)dataset show that,compared with the baseline model RT-DETR,with a reasonable increase in parameter count,LH-RTDETR improves the mean average precision by 2.7 percentage points for IoU at 50%and by 3.8 percentage points for IoU from 50%to 95%,as well as increasing the recall by 2.5 percentage points.Ablation experiments conducted on the same dataset demonstrate that LH-RTDETR can effectively reduce the miss rate and improve detection accuracy.
With the deepening recognition of the value of data elements across society,establishing a reasonable data pricing mechanism has become a core issue in promoting the development of the data elements market.In recent years,path query technology has achieved remarkable results in optimizing industry resource allocation efficiency and enhancing decision-making efficiency,continuously generating significant economic benefits for society.However,current research on graph data pricing primarily focuses on issues such as social network pricing and statistical query pricing,with limited studies on path query pricing.Therefore,this paper conducts research on path query pricing in graph data and proposes a person-alized path query pricing mechanism.Specifically,this pricing mechanism,based on providing consumers with optimal solutions,considers different types of consumers'preferences for path length and price,and returns the final price.It also allows consumers to query multiple paths in a single request.Additionally,this paper analyzes potential arbitrage issues in path query pricing and designs corresponding arbitrage-free mechanisms for each type of scenario.Furthermore,it develops an exact query pricing algorithm and an approximate query pricing algorithm for different query conditions.To avoid the additional computational costs induced by graph updates,this paper investigates dynamic query pricing and proposes a new solution to prevent recalculating query prices from scratch,thereby reducing computational costs.Finally,experi-ments using real and synthetic data verify that the proposed algorithms can effectively price large-scale graph data based on pricing theory.
Large language models(LLMs)have demonstrated remarkable capabilities in natural language understanding and multi-step reasoning.Transferring these reasoning advantages to graph-structured tasks has become a crucial direction in graph learning.However,directly applying LLMs to graph data faces challenges such as significant modal disparity and the difficulty of effectively encoding structural information.To address the limitations of existing methods,including insufficient graph-text alignment,the absence of explicit reasoning processes,and high fine-tuning costs,this paper proposes GraphCoT,a graph representation learning approach based on efficient chain-of-thought(CoT)fine-tuning.The core of this approach lies in a high-quality CoT distillation mechanism,where a powerful teacher model is used to generate instruction data containing intermediate reasoning paths,thereby explicitly guiding the student model to learn multi-step reasoning with minimal training data.Additionally,GraphCoT introduces a graph-text alignment module to achieve cross-modal alignment between graph representations and language space,and adopts a lightweight two-stage training strategy that prioritizes modal alignment before task-specific fine-tuning.This approach significantly reduces training costs while balancing efficiency and transferability.Experimental results on node classification and link prediction tasks across multiple datasets demonstrate that GraphCoT outperforms existing state-of-the-art methods in accuracy and generalization,validating the effectiveness of both the CoT distillation mechanism and the proposed alignment strategy in graph representation learning.
The reduction of internal unit sizes of computer makes the devices more susceptible to short-channel effects,current noise and electromagnetic interference,leading to transient bit-flip errors(also known as soft errors).Some of soft errors may result in silent data corruption,which is a serious threat to the reliability of computer systems.Therefore,it is urgent to predict and evaluate the impact of soft errors effectively.Compared with CPUs,the semantics of GPU are highly coupled with the parallel behavior between threads,and thus the propagation of soft errors is more complex.That makes it difficult for existing soft error resilience assessment methods to be applied at the GPU assembly level.To address the problem,a soft error resilience prediction model for GPU programs,EPKGP(soft error resilience prediction model with error propagation knowledge graph),is constructed based on an error propagation knowledge graph,so as to predict the soft error resilience of program fault points at SASS(streaming assembly)assembly level.Firstly,large-scale fault injection experiments are carried out to analyze and extract dynamic heuristic information from data,such as instruction types,bit-flip positions,bit-flip types and thread block numbers.Then,a SASS assembly code corpus is constructed to pre-train the Word2Vec model,mine abundant semantic knowledge and obtain static instruction semantic embedding features.Finally,this paper constructs an error propagation knowledge graph by virtue of the control flow and data dependency relationships between program instructions.The two heterogeneous features are fused as input and the graph convolutional network(GCN)is utilized to efficiently predict soft error resilience of GPU programs.Experiments show that compared with existing soft error prediction methods like G-SEPM,the accuracy of EPKGP is increased by up to 4.05 percentage points and the F1 score rises by up to 3.92 percentage points.
Event extraction is a key task in natural language processing that aims to automatically identify event triggers,event types,arguments,and argument roles from unstructured or semi-structured text and convert them into structured representations.To address the challenges faced by traditional deep learning models in extracting financial events,such as high annotation costs,low efficiency in handling long documents,and limited capability for parsing complex events,this paper proposes a large language model event extraction method based on multi-dimensional instruction set fine-tuning(MIFEE).MIFEE begins by designing a multi-dimensional instruction library that addresses various requirements of financial event extraction.It then proposes an instruction importance-driven stacking strategy,which first quantifies the independent contribution of each instruction dimension,and then stacks them in order of importance.This approach maximizes the complementary effects across dimensions while avoiding instruction redundancy,thereby constructing a multi-dimensional instruction set.Finally,the model fine-tunes a large language model based on this multi-dimensional instruction set to enhance financial event extraction performance.Experimental results show that MIFEE outperforms the compared baseline methods on the general financial event extraction datasets ChFinAnn and DuEE-Fin,achieving optimal F1 scores of 90.7%and 77.3%,respectively,demonstrating the superiority of MIFEE.In performance validation experiments conducted under long-text and complex event scenarios,MIFEE surpasses few-shot in-context learning(Few-shot ICL)methods in both F1 score and inference speed,confirming its strong capability in parsing efficiency and accuracy when handling complex semantic structures and long-range dependencies.Additionally,ablation experiments further validate the effectiveness of the instruction importance-driven stacking strategy.
To address the issues of missed and false detections caused by background interference,object occlusion,small target size,and texture blur in remote sensing object detection,this paper proposes a detection method named RAP-YOLO.Firstly,an RCSOSA module is designed,which integrates the multi-branch convolutional structure of RCS(residual con-textual spatial)with the one-shot feature aggregation mechanism of OSA(one-shot aggregation).By applying structural reparameterization,it enhances feature extraction efficiency and then unifies multi-level feature fusion,thereby improving the model's perception of remote sensing images.Secondly,a CBAM(convolutional block attention module)attention module is introduced.The channel attention module extracts semantically significant channel features by combining global average pooling and max pooling with an MLP(multi-layer perceptron)to model important channels.Then,the spatial attention module captures key positional information using pooling operations and a 7×7 convolution to model spatial relationships,enabling precise localization and enhanced feature representation of target areas.Finally,the PKIBlock(poly kernel inception block)module is developed,which integrates multi-scale Depthwise convolutions,convolutional feed-forward network(ConvFFN)for nonlinear transformation,and a CAA(context anchor attention)mechanism for context-aware feature modeling.This enhances both the expressive power of features and semantic understanding,signifi-cantly improving detection accuracy and robustness for multi-scale objects in remote sensing images.Comparative experi-ments on DIOR and NWPU VHR-10 datasets show that,compared with YOLO11,the proposed method achieves mAP scores of 85.2%and 91.2%,with recall reaching 79.0%and 84.1%,respectively,both outperforming baseline models.Overall,the proposed algorithm demonstrates superior accuracy and robustness in remote sensing object detection.
Accurate application of music style recognition technology in intelligent teaching systems significantly enhances learners'classroom engagement and learning outcomes.Given the limitations of existing single deep learning models in audio feature extraction and temporal dependency modeling,alongside the potential of large language model(LLM),this paper investigates a hybrid ensemble strategy spanning four deep learning architectures—convolutional neural network(CNN),recurrent neural network(RNN),long short-term memory(LSTM),and Transformer.The approach uses convolutional layers in the first-stage architecture to extract spectral-spatial features from audio,and leverages attention mechanisms in the second stage to capture long-range temporal dependencies among these features.To further assess performance differences across technical approaches,this paper not only compares multiple metrics of existing single-model baselines but also evaluates the zero-shot classification capability of large language models such as GPT-4.Experiments show that a CNN-RNN hybrid ensemble model achieves 99.0%classification accuracy on the GTZAN public music-genre dataset,significantly outperforming standalone CNN,RNN,LSTM,and Transformer models as well as other hybrid ensembles,and markedly exceeding the 62.3%accuracy of large language models.ROC curve analysis further validates the classification performance of the CNN-RNN hybrid model,with an AUC approaching 1.000.Confusion matrix results indicate that the model nearly eliminates misclassification of acoustically similar genres.A developed mini-program successfully integrates the best-performing model into intelligent teaching practice,enabling precise and real-time music style feedback,which is anticipated to effectively advance the development of adaptive music education.
Tracking device locations continuously at a global scale is fundamental to cybersecurity research.Existing tech-niques typically rely on cooperation with network operators to deploy dedicated hardware and software,require users to install tracking clients with approvals,or obtain side-channel information through hacking methods.To address it,this paper proposes a privacy attack method that enables a remote attacker with no privileges to leverage legacy IPv6 configure mechanisms to infer devices'hardware MAC addresses from their IP addresses,thereby achieving persistent global tracking of a large number of devices.Specifically,the method involves periodic,large-scale IPv6 network scans to collect online devices'IPv6 addresses and their leaked MAC addresses by reverse engineering,and then uses the MAC addresses as device identifiers combined with open-source IP geolocation databases to continuously track device locations.To overcome the challenge of identifying IPv6 addresses that leak device identifiers,global IPv6 scans were conducted continuously over one month with a 24-hour cycle,resulting in the collection of 1.27 billion IPv6 addresses containing device identifiers,corresponding to about 518.6 million unique devices(based on MAC address counts),among which the location changes of approximately 100 million mobile devices were successfully tracked.Experimental results demonstrate that even a single scanning node allows a non-privileged attacker to easily track hundreds of millions of devices,posing a severe threat to user privacy and network security.
Aiming at the challenges of limited multi-scale object detection performance and densely distributed small-scale lesions in grape leaf disease detection, this paper proposes an improved high-precision detection model based on YOLOv11, named VitiDetect-Net (vitis detection network). A phytomorph fusion module (PFM) is designed, which constructs multi-granularity feature representations through heterogeneous dilated convolutions and incorporates channel compression and attention reweighting mechanisms to achieve adaptive cross-level feature fusion, significantly enhancing the discriminative capability of the model for multi-scale targets. A patho-attention gate (PAG) module is introduced, which establishes a dynamic gating filter mechanism in the channel dimension to enhance the response to diseased regions, and incorporates deformable convolutions in the spatial dimension to strengthen the perception of lesion morphology, achieving synergistic optimization in both channel and spatial domains. This improves the detection performance for dense small targets while effectively reducing the number of parameters. An EdgeAlign sampler (EAS) is proposed, which employs an offset prediction network to generate adaptive sampling grids and combines channel attention to enhance key feature selection, effectively preserving the structural integrity of small target edges and reducing missed and false detections of small-scale lesions. Experiments on a self-built grape disease dataset show that the proposed model achieves a 3.6 percentage points improvement in mAP50:95 over the baseline model. Deployment tests on an NVIDIA Jetson Xavier NX edge device demonstrate an inference speed of 49.5 FPS, meeting real-time diagnostic requirements. Generalization experiments on the VisDrone dense small object dataset show a 1.4 percentage points improvement in mAP50, verifying the strong cross-domain adaptability of the model. This study achieves a good balance between detection accuracy and inference efficiency, providing a reliable solution for real-time disease diagnosis systems in smart agriculture.