
Satellite-Ground Integrated Networks (SGINs) have become novel network architecture to provide mega-access services for terrestrial users to release network resources, reduce loading pressure and obtain seamless communication functionalities. Nevertheless, due to limited loading capacities and complex constellation configuration of low earth orbit (LEO) satellites, this makes it hard to directly apply conventional resource scheduling policies in the dynamic SGINs environment. Moreover, substantial multi-modal information in the SGINs environment hinders superior service performance. Thus, we establish a multi-modal enabled task scheduling model to minimize network latency and privacy authentication overhead for processing stochastic task types, time-varying channel gains and dynamic LEO locations. Next, we design a transformer based group relative policy optimization (GRPO) algorithm to further optimize the CPU cycle frequency, the number of quantization bits and block size. Meanwhile, we present a blockchain based privacy authentication protocol to guarantee secure task transactions and demonstrate the quantization error bound via extensive mathematical analytical derivations. Finally, massive simulation results show that the proposed Transformer-GRPO algorithm has superior performance gains in term of network latency and privacy overhead compared with some state-of-the-art benchmarks.
Multi-path routing in Space-Air-Ground Networks (SAGINs) efficiently utilizes heterogeneous resources to optimize end-to-end quality of service, thereby proving critical for next-generation networks. However, dynamic network topologies, resource scarcity, and multi-objective tradeoff pose significant challenges to routing protocols, thus driving their evolution from mere path selection toward holistic resource management that strikes a balance between network performance and energy consumption under stringent operational constraints. To tackle these challenges, we investigate multi-path routing for SAGINs to optimize flow control, path selection, power control, and mission scheduling for energy-efficient and high-performance network operation. Firstly, we propose a Dual-Scale Time-Varying Graph (DSTVG) model for SAGINs that captures topology evolution at large timescales and data transmission at small timescales, achieving significantly lower construction complexity than traditional models. Using DSTVG, we formulate the studied problem as a mixed-integer non-linear program to minimize energy consumption while maximizing mission weight sum. For mission-guaranteed energy minimization scenario, we achieve polynomial-time global optimality by exploiting structural properties under mild assumptions. For the mission weight-energy trade-off scenario, we propose a low-complexity, two-layer algorithm yielding high-quality solutions. Simulation results demonstrate the proposed method achieves a performance improvement of at least $178.5\%$ compared with baseline approaches.
Rapid and precise semantic segmentation of flash floods, such as landslides and debris flows, is critical for emergency response and post-disaster assessment. In recent years, Unmanned Aerial Vehicle (UAV) swarms offer flexible large-area coverage but face two major challenges for collaborative model training: (i) highly non-Independent and Identically Distributed (non-IID) data distributions that degrade global accuracy, and (ii) prohibitive communication costs in mobile ad hoc networks. To address these challenges, we propose Federated Multi-Level Knowledge Distillation (FedMLKD), a novel framework that enables UAVs to exchange both logits and adaptively pooled intermediate features. This design preserves critical spatial structure while drastically reducing transmission overhead. A client quality–aware aggregation strategy further ensures robust global knowledge construction. Experiments on the Landslide segmentation dataset demonstrate that compared to state-of-the-art baselines, FedMLKD improves mIoU by 9.19% and the average Boundary F1-score by 85.54%, while reducing communication cost by nearly 200 times. These results highlight FedMLKD's ability to balance accuracy and efficiency, making it a practical candidate for real-time disaster mapping in bandwidth-constrained UAV swarms.
The explosive growth of Industrial Internet of Things (IIoT) has been driving adoption of mobile edge computing (MEC) technique, which mitigates the tension between computation-intensive tasks and the limited capabilities of IIoT terminal devices (TDs). To meet the growing computational demands in IIoT applications, unmanned aerial vehicle (UAV)-enabled MEC has emerged as a promising framework. However, stemming from the inherent openness of air-to-ground (A2G) links, the covertness of data offloading behavior in UAV-enabled MEC is exposed to considerable risks. To address this issue, we investigate a UAV-enabled covert MEC network against the detection from an illegal warden under the assistance of multiple jammers. To enhance the computation efficiency, we formulate a joint design problem of UAV trajectory, communication resource and computation resource to maximize the computation data amount while adhering the mobility, energy consumption, resource limitation and especially the covertness constraints. Then, to address the problem with complicated covertness metric under assistance of multiple jammers, we utilize the hypothesis testing theory to derive closed-form expressions of the covertness metric, which is for the first time to analysis the covertness performance with multiple friendly jammer. Subsequently, the monotonicity of the derived metric is analyzed and further utilized to simplify the original covertness constraint into an equivalent but analyzable form. Afterward, a novel convex approximation approach is developed to construct tight convex problem, enabling an efficient algorithm to attain a high-quality solution. Finally, simulations validate the effectiveness of proposed design.
Private Transformer inference for IoT network intrusion detection (NIDS) is limited by the communication and synchronization cost of secure attention and activation layers. We present PrivSentry, a two-party workload-aware private inference framework for IoT NIDS. The design is driven by a calibration result from the evaluated compact detector: on public IoT NIDS datasets, row-wise max-shifted attention logits are concentrated near zero with a long negative tail. PrivSentry uses this workload structure to instantiate a public-cutoff Softmax, a bounded-domain GELU surrogate, and a fixed-schedule LayerNorm routine over additive shares. These operators replace secure exponentiation, wide-range nonlinear evaluation, and data-dependent control flow with low-degree arithmetic dominated by batched secure multiplications. Across six public NIDS datasets, the secure BERT-Tiny pipeline remains close to plaintext inference, with only a 0.1 percentage-point average F1 decrease. Under the default 400 Mbps / 4 ms edge-WAN setting, PrivSentry completes a 128-token BERT-Tiny query in 10.269 s, which is $3.7\times$ faster than the measured Bolt baseline and $1.9\times$ faster than the measured BumbleBee baseline. These results support workload-calibrated secure-operator design for compact private IoT NIDS, while delimiting the scope of the system for deeper backbones and more constrained links.
Internet of Vehicles (IoV) helps detect traffic anomalies by analyzing data from vehicles and roadside devices, but raises concerns regarding data privacy. Current solutions struggle with inefficiency over encrypted data or sacrifice security for better performance, making it hard to find a balance between security and availability. To address the above issue, this paper proposes a privacy-preserving road traffic anomaly detection scheme (GRANT), which can provide efficient anomaly detection services over encrypted data. Specifically, we first design a Euclidean Distance-based Inner Product Functional Encryption (ED-IPFE) protocol for measuring the similarity between encrypted data. Next, to handle tight on-board resource constraints and real-time requirements, we introduce an Enhanced hierarchical K-Means Tree (EKM-Tree) to reduce the search space and improve detection efficiency. Finally, by integrating ED-IPFE and EKM-Tree, we present GRANT, which can preserve sensitive information while enabling accurate and efficient anomaly detection services. Security analysis reveals that GRANT provides post-quantum security under the hardness assumption of the Hint Learning With Errors (HintLWE) problem. Experimental results show that GRANT substantially improves both detection accuracy and efficiency. The detection accuracy reaches nearly 95%, and the average retrieval latency remains below 0.19 ms over 10, 000 samples with 256 dimensions.
With the ongoing increase of Mobile Internet of Things (mIoT) devices and their data, ensuring data security and user privacy has emerged as critical challenges. Encryption techniques are extensively employed to safeguard the mIoT data confidentiality during transmission. Nonetheless, the risk of adversaries forging mIoT data and compromising user privacy remains a concern. To alleviate the issue, some scholars turned to ring signatures to uphold the mIoT data integrity and user anonymity. But to our best knowledge, existing mIoT data management systems fail to simultaneously guarantee data confidentiality and integrity, ciphertext access control, and user anonymity during data transmission, thereby hindering convenient and secure mIoT data sharing. In this paper, we propose a cloud-assisted framework named DRSC_mIoT, which utilizes our designed DualRing signcryption (DRSC) to enhance the convenience and security of mIoT data sharing. To achieve this, we integrated the DualRing signature with tag-based encryption to present a novel DRSC construction tailored for secure mIoT networks. Assisted with the DRSC, our DRSC_mIoT ensures both data integrity and confidentiality, providing ciphertext access control while safeguarding the sensitive information privacy of data senders. Rigorous analysis indicates that our DRSC_mIoT satisfied the security properties of IND-iCCA, EU-iCMA, and anonymity. The comprehensive performance evaluation results demonstrate that our computation overhead is practical for mIoT. Our communication overhead is only 16.2% to 31.8% of others when the ring size is 64. The simulated evaluation on mIoT constrained devices demonstrates that DRSC's resource demands are practicality for real-world mIoT applications.
Satellite-based terrestrial specific emitter identification (SEI) supports wide-area, continuous electromagnetic monitoring for spectrum management and satellite communication security. However, this task faces two major challenges: (i) signal impairments caused by complex ground-satellite propagation channels and (ii) the difficulty of identifying diverse unknown emitters. To address these challenges, this paper proposes StarSense, a novel SEI framework that reformulates the classical SEI problem as a nonlinear dynamical-system identification problem based on phase-space reconstruction theory. Specifically, we introduce phase-space learning to adaptively model delay-coordinate representations over a bounded set of candidate delays and implement it using a structure-aware representation network (STAR-Net). We further develop an open-set identification mechanism that integrates embedding-space discrimination, extreme value theory (EVT)-based dual-threshold rejection, and unknown-sample clustering (OPEN-S3). This mechanism distinguishes multiple unknown emitters without requiring the number of unknown emitters to be known in advance. Experiments on publicly available RF datasets evaluate StarSense under the considered dominant ground-satellite propagation effects and varying SNR and openness conditions. StarSense improves the macro-averaged unknown-class assignment accuracy from 70.7% to 85.3% over S3R while using only about one-fifteenth as many parameters.
In this paper, we model the carrier shutdown in multi-carrier mobile networks as a deep reinforcement learning (DRL) problem. The proposed modelling framework performs energy-saving actions by selectively turning off carriers and intelligently reallocating their users consistently with temporal variations of users' density. In addition to maintaining service continuity, our policy further guarantees a novel user experience metric. Leveraging real and recent datasets, we train and evaluate our model over realistic network scenarios. Our results show more than 15% energy saving with the fulfillment of the user experience constraints, outperforming currently deployed solutions and researched approaches in the literature by almost 50%. More interestingly our approach exhibits generalization properties, a very promising characteristic for its adoption in real mobile networks deployment.
The integration of federated learning (FL) and multi-access edge computing (MEC) is a promising direction for next-generation wireless network. However, their joint deployment faces challenges from the limited bandwidth of edge devices (EDs), possible exposure of training information through transmitted updates, and Byzantine vulnerabilities in both EDs and edge servers (ESs). To address these challenges, this article proposes CEBTFL, a communication-efficient and Byzantinetolerant FL for MEC. Specifically, CEBTFL integrates four collaborative mechanisms: i) a sign-optimized momentum update (SOMU) to reduce communication overhead while maintaining model accuracy and data privacy; ii) a lightweight directional screening at ESs that employs cosine similarity-based trust evaluation to suppress unreliable local updates; iii) a hybrid projection consensus at the cloud server (CS) that combines real and dummy data projections for robust global aggregation; and iv) an error-compensated mechanism to stabilize convergence under gradient compression. Theoretical analysis demonstrates that CEBTFL maintains stable convergence, while dummy data can enhance Byzantine robustness. Comprehensive experiments show that CEBTFL sustains high model accuracy and communication efficiency, even with up to 60% Byzantine EDs and 40% Byzantine ESs.
Unmanned aerial vehicle (UAV) swarm provides effectively and diversely applications in low-altitude networks. However, due to the non-stationary open nature wireless environments, and the untrusted UAVs can participate in wireless communication tasks, which cause severe security threats. In this paper, we consider an integrated communication, sensing, computation and security framework, where the swarm UAV carried reconfigurable intelligent surfaces (ARIS) cooperative for communication task offloading with the existence of untrustworthy ARISes. Specifically, we propose a secure collaborative communication task offloading framework, which is federated multiagent deep deterministic policy gradient with the score-trust based aggregation refinement (FedMADDPG-STAR) strategy. In this framework, each ARIS agent learns local trajectories under limited observations and periodically uploads model parameters rather than the raw data to a global server for aggregation via FedMADDPG. On the other hand, the proposed STAR algorithm enhances robustness by comprehensively considering the consistency of training directions, training progress, and multivariate anomaly indicators, thereby enabling adaptive trust scoring and isolation of untrustworthy ARISes. In addition, to increase the training stability and the efficiency of the proposed FedMADDPG-STAR, we design the refined reward function and employ the fast optimal phase-shift configuration. The system complexity is then analyzed. Simulations comparing different robust aggregation methods under various Byzantine attacks demonstrate that our proposed STAR scheme outperforms the typical aggregation approaches. We find that the proposed FedMADDPG-STAR framework can achieve secure and efficient communication task offloading under the non-stationary open nature wireless environments and the Byzantine attacks.
Wireless Power Transfer (WPT) is widely used to replenish energy for the devices in Wireless Rechargeable Sensor Networks (WRSNs). The existing studies typically assume that chargers provide free charging service strictly follow the assigned scheduling results. However, the chargers may be owned by individuals, who expect to profit from energy trading, and may take strategic behaviors to increase their own profits. Moreover, the adversaries could infer the charging costs of chargers from the scheduling results. This exposes energy trading in WRSNs to persistent threats from both incentive attack and inference attack. To address these challenges, we propose a budget-constrained auction-based model for directional charging service in WRSNs, and formulate the Charger Deployment problem for Utility Maximization (CDUM) under the budget constraint. We adopt the auction theory and differential privacy technique to design a novel directional Charger Deployment Mechanism based on budget-feasible auction with Personalized Privacy Preserving (P $^{3}$ CDM). P $^{3}$ CDM determines the charger deployment strategy and payment for each charger based on probability distributions. The theoretical analysis demonstrates that P $^{3}$ CDM achieves computational efficiency, truthfulness, individual rationality, differential privacy, and utility maximization. The experimental results show that P $^{3}$ CDM reduces privacy leakage by an average of 71.4% compared with other privacy-preserving algorithms. Compared with other benchmarks, P $^{3}$ CDM only incurs an increase of 5.64% in payment, and 24.1% reduction in utility. The field experiments further demonstrate that P $^{3}$ CDM achieves the same utility as the non-privacy-preserving algorithm, with the cost of 4.32% increase in average payment per charger, while preserving charging cost privacy.
Many real-time mobile applications, such as video conferencing, virtual reality, and industrial internet of things (IIoT) control, impose strict deadlines on packet arrivals. Multipath transmission, supported by most modern mobile devices through both Wi-Fi and cellular interfaces, can enhance the performance of these real-time applications. However, most existing multipath schedulers fail to jointly consider deadline requirements and the monetary costs of using multiple paths, where cellular data plans are generally more expensive than Wi-Fi plans. To address this issue, we propose a Deadline-aware Multipath Packet Scheduler, DaMPS, designed to prioritize packet delivery within packets' deadlines in wireless mobile networks. DaMPS performs intelligent scheduling via optimization and adapts its decisions to varying network dynamics based on a Multi-Armed Bandit (MAB) technique. To further reduce the monetary costs associated with transmission, we introduce a Cost-aware extension for DaMPS, named CaDaMPS, to select cost-effective paths while ensuring deadline satisfaction. We implement DaMPS and CaDaMPS in MPQUIC and evaluate their effectiveness under various network conditions using Mininet. Our extensive experimental results show that DaMPS and CaDaMPS improve deadline adherence by 6.37%$\sim$63.05% and 5.32%$\sim$61.44%, respectively, while maintaining high throughput and low latency comparable to those of other schedulers. Moreover, CaDaMPS reduces the average cost per packet by 20%$\sim$59% without compromising deadline requirements.
Low-power wide area networks (LPWANs) have gained significant traction in recent years, with LoRaWAN (long-range WAN) emerging as a prominent representative. Recently, LoRaWAN has expanded into the 2.4 GHz unlicensed band to support global deployments and higher data rates. However, this shift introduces severe cross-technology interference (CTI) from coexisting Wi-Fi networks that share the same spectrum, leading to substantial communication failures. As a physical-layer solution to this problem, this paper presents CoWiL to combat CTI from Wi-Fi to LoRa (the physical layer of LoRaWAN) in the 2.4 GHz band for better coexistence between LoRaWAN and Wi-Fi networks. Existing approaches address this problem by compromising Wi-Fi performance or assuming that Wi-Fi interferes with only a small portion of LoRa signals. Unlike them, CoWiL does not affect normal Wi-Fi communications and remains effective across varying levels of CTI. This is achieved by implementing CoWiL at the LoRa receiver side to directly extract LoRa data out of the CTI from Wi-Fi. Specifically, CoWiL leverages the temporal correlation between the preamble and payload of a LoRa signal, using the demodulated preamble to construct a frequency bin mask that aids in robust payload decoding. Experimental evaluations in diverse real-world settings demonstrate that CoWiL reduces the LoRa packet error rate by up to 90% compared to state-of-the-art methods under Wi-Fi-induced CTI, significantly enhancing coexistence performance.
WiFi fingerprint-based localization technology utilizing Received Signal Strength Indicator (RSSI) has gained significant attention due to its deployment convenience and cost efficiency. While existing methods rely extensively on offline-trained models, deep learning-based localization systems remain highly vulnerable to adversarial attacks that severely degrade localization performance. To address these challenges, this paper proposes a new diffusion-inspired localization defense method named as SRLoc to reconstruct clean RSSI signals from adversarially perturbed inputs. Instead of relying on the multi-step stochastic iterations which are typical of standard diffusion models, SRLoc employs a highly efficient single-step conditional denoising network and utilizes the attack intensity directly as a conditioning vector rather than traditional time steps. Furthermore, the model replaces the conventional U-Net upsampling module with a 1D transposed convolutional architecture tailored for sparse RSSI data, and is optimized with a customized triple-loss function. This function strategically imposes physical space limits and geographical consistency constraints on the basis of standard robust Huber loss, ensuring that the reconstructed signals remain physically plausible. Experimental results demonstrate that the SRLoc effectively defends against adversarial attacks on both the CJU dataset and Tampere dataset in comparison with four state-of-the-art methods and five baseline methods, while maintaining excellent localization performance. On the CJU dataset, the SRLoc can resist 86.5%, 80.8%, and 76.9% of the attacks from FGSM, PGD, and MIM respectively, outperforming the best comparative method by 9.2%, 12.6%, and 15.8% under attack intensity of 0.2. On the Tampere dataset, the SRLoc achieves attack resistance rates of 95.7%, 92.6%, and 94.6% against FGSM, PGD, and MIM respectively, surpassing the best comparative method by 10.6%, 6.8%, and 6.8%. Furthermore, SRLoc maintains excellent defensive capability even when facing unknown and black-box attack types. These results indicate that the SRLoc method effectively mitigates localization errors induced by adversarial attacks, thereby significantly enhancing the robustness of the localization model.
Visible light positioning (VLP) is a promising indoor localization technology with strong interference immunity and easy integration with existing infrastructure. However, most existing VLP systems rely on multi-anchor deployments and Received Signal Strength (RSS), which are costly to deploy and remain sensitive to receiver pose and ambient light. To address these limitations, we propose a single-anchor LED-array localization framework based on Phase-Difference (PD) fingerprints and Graph Neural Networks (GNNs). First, to mitigate the sensitivity of RSS and the complexity of multi-anchor geometry, we introduce the first single-anchor PD LED-array localization scheme. By encoding positional information in inter-LED PDs, this design eliminates the need for multi-anchor surveying and synchronization, suppresses slow illumination drifts. Second, to handle the instability of raw measurements, we develop a robust fingerprint construction pipeline that begins with photodiode samples and performs inter-LED phase estimation, temporal unwrapping with outlier rejection, wrap-safe sine–cosine embedding, and normalized storage in a compact database. Third, to address the limitations of heuristic K-nearest-neighbor matching, we propose a query-centric GNN-based localization framework that encodes physics-aware similarity cues in node features and learns geometry-aware neighbor weights. We evaluated the proposed method in four representative indoor scenes against ten baselines. Results show that PD fingerprints consistently outperform RSS fingerprints, and the proposed GNN further improves accuracy. It achieves mean errors of 0.26 m and 0.43 m in the corridor and office–corridor scenes, respectively, and reduces the mean error by up to 56% compared with the strongest baseline. These results demonstrate a scalable and robust pathway to high-accuracy VLP with lightweight infrastructure.
Federated learning (FL) is susceptible to poisoning attacks, where malicious clients manipulate local data or models to disrupt training. The system and data heterogeneity inherent in practical FL systems exacerbates these vulnerabilities, rendering existing defense mechanisms ineffective or infeasible. Specifically, distinguishing benign local models, trained on heterogeneous client data, from poisoned ones presents a significant challenge. Moreover, semi-asynchronous FL (SAFL) paradigms, commonly employed to address system heterogeneity, further complicate this issue by preventing fair evaluation of local models originating from different global models (i.e., with varying staleness). In this work, we propose a novel defensive framework (namely Fed-Beta) for robust and accurate FL model training under system and data heterogeneity. First, we introduce a staleness-aware SAFL paradigm, where the server accepts only a fixed number of local models per round and groups them based on their staleness. Then, we implement a two-stage aggregation mechanism. Specifically, we develop a robust intra-group aggregation method using model inversion to evaluate data-domain discrepancies among clients. This method accurately identifies and excludes malicious local models from aggregation, producing a reliable representative model for each group. Moreover, we design a model-consistency-aware inter-group aggregation method, which selectively aggregates group representative models with consistent update directions to update the global model. Theoretically, we conduct rigorous convergence analysis of Fed-Beta, offering insights into how system and data heterogeneity affect the defensive performance. Empirically, extensive experiments corroborate its superiority over existing schemes.
Wireless Rechargeable Sensor Networks (WRSNs) have become a key enabler of the Internet of Things. There are directional wireless chargers capable of emitting multiple charging beams to charge the nodes in WRSNs via Wireless Power Transfer technology (WPT); yet most existing research on antenna charging direction selection assumes a single-beam model, leading to the selection of a suboptimal antenna direction set. In addition, prior research mostly focuses on two-dimensional networks and lacks effective solutions for three-dimensional (3D) scenarios. In this paper, we investigate the Charging Scheduling problem using a Unmanned Aerial Vehicle (UAV) with Dual-Conical Charging Beams in 3D-WRSNs (CSUDB-3D) for charging the nodes in the network and prove it to be NP-hard. To address this challenge, we first tackle the problem of antenna direction set selection from the infinite-size entire 3D spherical direction set by elegantly designing an algorithm via exploiting the geometric properties among the nodes. We prove that this algorithm guarantees to return an antenna direction set with the minimum size that is functionally equivalent (FuncEqv)-meaning it ensures the same optimal scheduling performance-to the original infinite 3D spherical direction set, and hence name it the Minimum FuncEqv Direction Set Algorithm (MFEDS). Then, the Lin-Kernighan Heuristic (LKH) algorithm is adopted to determine a quasi-optimal charging tour for the UAV to charge the nodes. By integrating MFEDS and LKH, we build a three-step framework termed Scheduling of a UAV Charger with Dual-Conical Charging Beams in 3D-WRSNs (UAVDCB-3D) to effectively solve the CSUDB-3D problem. Simulation results demonstrate that UAVDCB-3D outperforms the best existing benchmark in terms of both total energy loss and time span. Specifically, it reduces total energy loss by up to 51.10% and the time span by up to 46.15%.
Multi-access edge computing provides an effective computation offloading paradigm for resource-constrained vehicles by deploying resources at network edges. However, high vehicular mobility and heterogeneous task demands lead to spatio-temporal load imbalances across static roadside units (RSUs). Prior works either lack responsiveness to sudden workload surges or exhibit poor robustness under sustained fluctuations. To fill this gap, we explore an aerial-terrestrial cooperative design, leveraging UAV assistance and cross-RSU task migration to enable responsive surge absorption and system-wide robustness. Nevertheless, achieving preemptive UAV position scheduling and anticipatory cross-RSU task migration remains challenging. We argue that proactive prediction of future dynamics is crucial for guiding timely and collaborative decision-making. On this basis, we propose a Dynamics-Aware Hierarchical Multi-Agent (DHMA) deep reinforcement learning approach that jointly optimizes UAV position scheduling and vehicular computation offloading. Additionally, we develop a prediction-driven observation augmentation method to enhance the perception capabilities of agents for future dynamics based on predicted workload surges and vehicular movements. Comprehensive evaluation across four simulation scenarios, together with real-world validation, demonstrates the effectiveness, robustness, and practical viability of our approach.
Low Earth orbit (LEO) satellite networks have evolved into a pivotal communication infrastructure due to their global coverage, low propagation delay, and high deployment flexibility, where a domain-partitioned architecture is typically employed for scalability. However, the extreme topological dynamics caused by continuous orbital motion, combined with the pronounced network heterogeneity across different control domains, pose substantial challenges for efficient and reliable routing. Conventional routing methods, which often rely on relatively stable topology structures, struggle to handle frequent link disruptions, rapid path-length variations, and constantly changing neighboring satellite sets. To tackle these challenges, this paper proposes a meta-learning routing framework built upon a discovering reinforcement learning (RL) optimization mechanism for dynamic and heterogeneous LEO satellite networks. The framework employs meta-optimization to automatically generate adaptive RL update rules, enabling agents to rapidly adjust their routing behaviors under heterogeneous and time-varying network conditions. Simulations reveal that the proposed approach effectively decreases energy consumption and delay while maintaining strong adaptability and scalability across diverse LEO sub-networks.