Low Earth Orbit (LEO) satellite constellations are emerging as an important platform for distributed dataflow execution in space-terrestrial integrated networks. Existing studies largely treat routing and processing separately, while next-generation LEO systems are expected to process and transform data in transit by leveraging on-board computing and software-defined infrastructures. However, jointly optimizing routing and in-network processing in dynamic LEO satellite networks remains challenging because of time-varying connectivity, limited on-board resources, and bandwidth constraints. In this paper, we formulate the Dynamic LEO In-network Processing Dataflow Optimization (DLIDO) problem, which aims to maximize the throughput of processed dataflows by jointly optimizing routing paths and processing-resource allocation over a dynamic flow network. We present an approximation algorithm with a proven (1−ϵ) approximation guarantee for 0 < ϵ ≤ 0.5, providing near-optimal throughput under dynamic processing and communication constraints. To further improve efficiency and practicality, we develop a 2-walk based iterative heuristic algorithm that substantially reduces runtime while maintaining strong empirical performance, and in some regimes provably optimal behavior. Extensive evaluations on realistic LEO network topologies show that both algorithms significantly outperform existing approaches in throughput and adaptability, highlighting a promising direction for dataflow-aware scheduling and optimization in dynamic satellite systems.
Efficient data transmission in low Earth orbit (LEO) satellite networks is critical for supporting real-time global communication, Earth observation, and numerous data-intensive space missions. A fundamental challenge in these networks involves solving the maximum flow problem, which determines the optimal data throughput across highly dynamic topologies with limited onboard energy and data processing capability. Traditional algorithms often fall short in these environments due to their high computational costs and inability to adapt to frequent topological changes or fluctuating link capacities. This paper introduces an accelerated maximum flow algorithm specifically designed for dynamic LEO networks, leveraging a prediction-enhanced approach to improve both speed and adaptability. The proposed algorithm integrates a novel energy-time expanded graph (e-TEG) framework, which jointly models satellite-specific constraints including time-varying inter-satellite visibility, limited onboard processing capacities, and dynamic link capacities. In addition, a learning-augmented warm-start strategy is introduced to enhance the Ford–Fulkerson algorithm. It generates near-optimal initial flows based on historical network states, which reduces the number of augmentation steps required and accelerates computation under dynamic conditions. Theoretical analyses confirm the correctness and time efficiency of the proposed approach. Evaluation results validate that the prediction-enhanced approach achieves up to a 32.2% reduction in computation time compared to conventional methods, particularly under varying storage capacity and network topologies. These results demonstrate the algorithm’s potential to support high-throughput, efficient data transmission in future satellite communication systems.
Given a natural language description, text-based person retrieval aims to identify images of a target person from a large-scale person image database. Existing methods generally face a color over-reliance problem, which means that the models rely heavily on color information when matching cross-modal data. Indeed, color information is an important decision-making accordance for retrieval, but the over-reliance on color would distract the model from other key clues (e.g. texture information, structural information, etc.), and thereby lead to a sub-optimal retrieval performance. To solve this problem, in this paper, we propose to Capture All-round Information Beyond Color (CAIBC) via a jointly optimized multi-branch architecture for text-based person retrieval. CAIBC contains three branches including an RGB branch, a grayscale (GRS) branch and a color (CLR) branch. Besides, with the aim of making full use of all-round information in a balanced and effective way, a mutual learning mechanism is employed to enable the three branches which attend to varied aspects of information to communicate with and learn from each other. Extensive experimental analysis is carried out to evaluate our proposed CAIBC method on the CUHK-PEDES and RSTPReid datasets in both supervised and weakly supervised text-based person retrieval settings, which demonstrates that CAIBC significantly outperforms existing methods and achieves the state-of-the-art performance on all the three tasks.
The rapid expansion of Low Earth Orbit (LEO) satellite constellations presents new challenges for maintaining efficient inter-satellite communication under dynamic network topologies. Traditional minimum spanning tree (MST) algorithms, designed for static or quasi-static networks, are inefficient in handling frequent topology changes and fluctuating link costs in LEO environments. Moreover, existing dynamic MST approaches often rely on localized updates, which lack global foresight, resulting in scalability issues and poor real-time responsiveness. To address these challenges, we propose a novel time-varying topology model based on the inter-satellite link (ISL) properties and orbital prediction, introducing the concepts of Stable ISL Period (SIP) and Stable Communication Period (SCP) to identify intervals of topological stability. Building on this model, we develop the All-Time Dynamic Minimum Spanning Tree (ATDMST) algorithm, which incorporates a filtering mechanism to avoid unnecessary recomputations and employs an "edge swapping" strategy for incremental updates. The ATDMST algorithm dynamically maintains optimal spanning trees as ISL communication costs evolve, significantly reducing computational overhead while preserving communication efficiency. Evaluation results show that, that ATDMST outperforms existing methods, achieving up to an 8.8% reduction in communication costs, 95% fewer topology updates, and a 72% decrease in response time, while maintaining robust performance.
Network traffic classification is crucial for monitoring network health, detecting malicious activities, and ensuring Quality-of-Service (QoS). The use of dynamic ports and encryption complicates the process, rendering traditional port-based or payload-based classification methods ineffective. Conventional machine learning and statistical approaches often depend on manual feature or pattern extraction by experts, leading to inefficiencies and potential inaccuracies. Deep learning offers a promising alternative, with its inherent capability to autonomously extract patterns and features from data. Nonetheless, the design of existing deep learning models often limits them to high-level semantic feature extraction, neglecting the rich multidimensional spatial and temporal information in network traffic. To address these limitations, this paper introduces STARNet, a deep learning-based model for encrypted traffic classification. STARNet incorporates a dual-stream pathway network architecture that optimizes feature extraction from each pathway. It also features a novel spatiotemporal multidimensional semantic feature recall mechanism, designed to enrich the model’s analytical depth by retaining important information that might be missed when focusing solely on high-level features. Evaluated on two public network traffic datasets, STARNet demonstrates superior accuracy in traffic classification tasks, highlighting its potential to enhance network monitoring and security.
Convolutional neural networks (ConvNets) have been widely used for feature extraction in various computer vision tasks, such as image classification, object detection, and instance segmentation. Recently, Vision Transformers (ViTs) have demonstrated exceptional performance on upstream vision tasks, such as image classification, due to their effectiveness in modeling long-range dependencies. However, ViTs' performance is limited by weak inductive biases in modeling two-dimensional data and under-explored multi-scale representation learning, both of which are fundamental strengths of ConvNets. In this paper, we introduce CT-Mixer, a novel architecture that combines convolution-based and transformer-based modules to exploit the advantages of both and compensate for their respective weaknesses. The CT-Mixer architecture cross-stacks convolution-based and transformer-based modules, where each module is assigned an order to better learn local information and global context. Additionally, we incorporate a dynamic mechanism into the convolution-based modules to model adaptive dependence. We also improve the multi-scale representation learning strategy by adopting the multi-branch structure of MPViT. Experimental results demonstrate that CT-Mixer achieves competitive performance compared to existing methods.
Recent mainstream masked distillation methods function by reconstructing selectively masked areas of a student network from the feature map of its teacher counterpart. In these methods, the masked regions need to be properly selected, such that reconstructed features encode sufficient discrimination and representation capability like the teacher feature. However, previous masked distillation methods only focus on spatial masking, making the resulting masked areas biased towards spatial importance without encoding informative channel clues. In this study, we devise a Dual Masked Knowledge Distillation (DMKD) framework which can capture both spatially important and channel-wise informative clues for comprehensive masked feature reconstruction. More specifically, we employ dual attention mechanism for guiding the respective masking branches, leading to reconstructed feature encoding dual significance. Furthermore, fusing the reconstructed features is achieved by self-adjustable weighting strategy for effective feature distillation. Our experiments on object detection task demonstrate that the student networks achieve performance gains of 4.1% and 4.3% with the help of our method when RetinaNet and Cascade Mask R-CNN are respectively used as the teacher networks, while outperforming the other state-of-the-art distillation methods.
It is challenging for either traditional modelling or experiments to capture the complex relationship between the strength of cement composites and the many influencing factors. Machine learning (ML) with powerful data analysis provides an elegant solution. This work adopts random forest (RF), AutoGluon-Tabular (AGT) and artificial neural network (ANN) to predict the 28-day compressive strength (CS) of carbon nanotube (CNT) reinforced cement composites (CNTRCCs), in which the effects of the specimen size is incorporated for the first time. In addition, this work introduces a Gaussian function to account for the distribution of CNT dimensions. An adaptive training strategy is proposed to improve the performance of the ML models. Specimen size is found to have significant effect on the 28-day CS of the CNTRCCs. ANN is evidenced to have the best performance whereas it requires intensive experience and computation. In contrast, AGT demonstrates improved efficiency and flexibility with satisfying predicted results.
Text passwords are the primary method for identity authentication. However, easy-to-remember passwords are vulnerable to password-guessing attacks. The study of password-guessing not only enhances the understanding of password security, but also promotes the improvement of password library security. Nowadays, deep learning based models have demonstrated their promising ability for password-guessing, e.g., Recurrent Neural Network (RNN) and its variants. However, RNNs are failed to parallelize and lack long memory, resulting in unsatisfactory guessing efficiency and effectiveness. Aiming to generate a high-quality password dictionary for password checking, we propose a temporal convolutional network-based password-guessing model named PGTCN. Specifically, in PGTCN we improve the password-guessing efficiency by introducing the Improved Temporal Convolutional Network (ITCN) which adopts residual learning and feature fusion techniques. PGTCN can automatically study the structure and characteristics of passwords and generate new passwords based on the learned knowledge. To verify the performance of PGTCN, we compare it with state-of-the-art models on six public password datasets. Evaluation results show that the password dictionary generated by the proposed PGTCN achieves the structure coverage rate up to 84%, and boosts the matching rate up to 22%. We conclude that the PGTCN could be regarded as a robust and efficient model for password guessing.
Application offloading plays a crucial role in application deployment in edge-cloud computing. However, finding the optimal solution for application offloading is challenging due to the heterogeneous resources, computing dependency of tasks, and complex network. Existing research on application offloading problems often assumes that the communication delay between the edge and the cloud (inter-side) is symmetrical or that within the cloud or edge (intra-side) can be omitted. However, this assumption is not practical considering the distinct features of the clouds and the edge clusters. Therefore, we study application offloading in the heterogeneous edge-cloud environment by considering both intra-side communication delay between tasks assigned to the same side and asymmetry inter-side communication delay between edge and cloud sides. We first focus on the specific circumstances with a boundary condition that lead to an optimal offloading solution in a heterogeneous edge-cloud environment. Then we study the general case by designing an iterative algorithm with maximum gain technique to solve it. Furthermore, considering various bandwidths within each side and resource capacities of physical nodes, we develop two efficient algorithms by combining minimum cut and maximum gain approaches. Both simulations and real trace-based evaluations are conducted to validate that the proposed algorithms outperform existing solutions.
As a general model compression paradigm, feature-based knowledge distillation allows the student model to learn expressive features from the teacher counterpart. In this paper, we mainly focus on designing an effective feature-distillation framework and propose a spatial-channel adaptive masked distillation (AMD) network for object detection. More specifically, in order to accurately reconstruct important feature regions, we first perform attention-guided feature masking on the feature map of the student network, such that we can identify the important features via spatially adaptive feature masking instead of random masking in the previous methods. In addition, we employ a simple and efficient module to allow the student network channel to be adaptive, improving its model capability in object perception and detection. In contrast to the previous methods, more crucial object-aware features can be reconstructed and learned from the proposed network, which is conducive to accurate object detection. The empirical experiments demonstrate the superiority of our method: with the help of our proposed distillation method, the student networks report 41.3%, 42.4%, and 42.7% mAP scores when RetinaNet, Cascade Mask-RCNN and RepPoints are respectively used as the teacher framework for object detection, which outperforms the previous state-of-the-art distillation methods including FGD and MGD.
Edge artificial intelligence (AI) has emerged as a promising paradigm catering to overwhelming explosions of smart applications, by offloading the computation-intensive deep neural network (DNN) inference to an edge network for processing. The surging of edge AI brings new vigor and vitality to shape the prospect of smart transportation. However, when considering the cooperation between heterogeneous edge devices and the operation precedence between DNN tasks, it is still challenging to decompose a DNN across multiple edge devices in an edge network with a general topology to minimize DNN inference delay. In this article, we devise a polynomial-time optimal solution to the DNN inference offloading problem for smart roadside applications, in which the roadside edge network is usually organized with chain topology. Specifically, the DNN inference offloading problem for the roadside edge network with chain topology is transformed into an equivalent graph optimization problem. Theoretical analysis and extensive evaluations validate the performance of the proposed solution in minimizing the total inference delay.
Vision Transformers are born with the property of data-dependent and long-range dependencies, accomplishing a number of astonishing results against their contemporary competitor CNNs. To alleviate the excessive computational burden, previous methods apply the local operation (e.g., convolution, local attention) in the high-resolution stages. Although these designs are efficient for local relations learning, especially for the high redundancy stages, they inevitably lead to the losses of non-locality and are constrained by the limited receptive field. In this paper, we present an effective hybrid-style vision backbone that is explicitly built with dynamic convolution and self-attention to respectively undertake both local and global interaction, dubbed C2SFormer. We adopt two homogeneous modules whose structure follows the typical Transformers. For local relations learning, we take the parallel multi-scale design and additive aggregation as simple but effective ideas, named MS-SCDC. For the global context modeling, we leverage the efficient factorized self-attention mechanism proposed in CoaT and apply it with the MS-SCDC in a cross-stacking manner over the high-resolution stages. Additionally, we further introduce a general approach for multi-scale learning of transformer-based modules, named MS-MHSA. The experiments conducted on a variety of general-purpose vision tasks demonstrate the superiority of the proposed model.
Due to extremely high temperatures, friction, and vibrations, aircraft engines tend to have various types of internal damage. To guarantee the safety of the aircraft, maintenance for aircraft engines is regularly performed by manual borescope inspection, which is time consuming and error prone. Existing studies adopted deep learning-based approaches to detect potential damage in borescope images to accelerate the process of engine maintenance. However, these approaches are designed for damage detection from static images and cannot be directly used to track damage in borescope videos due to the extremely high computational cost. To detect and track the damage in borescope videos in real time, we propose the deep fusion network (DFNET), which works along two parallel functional paths: i.e., the segmentation path and the spatial warping path. The segmentation path only runs on selected key frames to extract semantic features, which are then propagated to other frames through optical flows in the spatial warping path. The performance and efficiency of the DFNET are validated through extensive experiments using real borescope videos from a local air carrier.
Metal-organic frameworks (MOFs) have become an active topic because of their excellent carbon capture and storage (CCS) properties. However, it is quite challenging to identify MOFs with superior performance within a massive combinatorial search space. To this end, we propose a deep-learning-based end-to-end prediction model to rapidly and accurately predict the CO2 working capacity and CO2/N-2 selectivity of a given MOF under low-pressure conditions. Different from previous methods, our prediction model relies only on the data from the Crystallographic Information File (CIF) rather than handcrafted geometric descriptors and chemical descriptors. The model was developed, trained, and tested on a dataset of 342489 topologically diverse MOFs. Experimental results on the dataset show that the proposed model achieves high prediction performance, i.e., R-2 = 0.916 for predicting the CO2 working capacity and R-2 = 0.911 for predicting the CO2/N-2 selectivity. With regard to the identification of potential high-performing MOFs, 1020 of 1027 (top 3%) high-performance MOFs were recovered while screening only 12% of the entire dataset using our provided pretrained model, reducing the computation time by nearly an order of magnitude when the model was used to prescreen material prior to computationally intensive grand canonical Monte Carlo (GCMC) simulations while still capturing 99% of the high-performance MOFs. In the ab initio training task, the method can achieve R-2 = 0.85 with only 20% of the labeled data used for training and recover 995 of 1027 (top 3%) high-performance MOFs with only 12% of the entire dataset screened.
The core problem of text-based person retrieval is how to bridge the heterogeneous gap between multi-modal data. Many previous approaches contrive to learning a latent common manifold mapping paradigm following a cross-modal distribution consensus prediction (CDCP) manner. When mapping features from distribution of one certain modality into the common manifold, feature distribution of the opposite modality is completely invisible. That is to say, how to achieve a cross-modal distribution consensus so as to embed and align the multi-modal features in a constructed cross-modal common manifold all depends on the experience of the model itself, instead of the actual situation. With such methods, it is inevitable that the multi-modal data can not be well aligned in the common manifold, which finally leads to a sub-optimal retrieval performance. To overcome this CDCP dilemma, we propose a novel algorithm termed LBUL to learn a Consistent Cross-modal Common Manifold (C^3M) for text-based person retrieval. The core idea of our method, just as a Chinese saying goes, is to `san si er hou xing', namely, to Look Before yoU Leap (LBUL). The common manifold mapping mechanism of LBUL contains a looking step and a leaping step. Compared to CDCP-based methods, LBUL considers distribution characteristics of both the visual and textual modalities before embedding data from one certain modality into C^3M to achieve a more solid cross-modal distribution consensus, and hence achieve a superior retrieval accuracy. We evaluate our proposed method on two text-based person retrieval datasets CUHK-PEDES and RSTPReid. Experimental results demonstrate that the proposed LBUL outperforms previous methods and achieves the state-of-the-art performance.
1研究背景和意义 近年来,人们的出行需求随着国家经济的快速发展而急速增长,而航空出行也日益成为人民群众主要出行方式之一.中国民用航空局的发布的公告显示,我国民航2022年虽受新冠疫情影响,但国内航线旅客运输量预计将达到4.4亿人次.
In this study, we present a tunable metamaterial consisting of rotatable non-uniform Mie resonators (NMRs) with identical structures. The metamaterial can in real-time manipulate the direction of acoustic radiation and guarantee high transmission efficiency by simply changing the rotation angle of the NMR unit cells, which is induced by the anisotropic property of NMR. In addition, according to generalized Snell’s law, the arbitrarily direction-scanning capability is realized by tuning the phase shift distribution along the metamaterial. Our proposed anisotropic metamaterial could contribute to designing a device for the emission and reception of acoustic waves in real-time.
Text-based person re-identification aims to retrieve images of the corresponding person from a large visual database according to a natural language description. When it comes to visual local information extraction, most of the state-of-the-art methods adopt either a strict uniform strategy which can be too rough to catch local details properly, or pre-processing with external cues which may suffer from the deviations of the pre-trained model and the large computation consumption. In this paper, we proposed an Adversarial Self-aligned Part Detecting Network (ASPD-Net) model which extracts and combines multi-granular visual and textual features. A novel Self-aligned Part Mask Module was presented to autonomously learn the information of human body parts, and obtain visual local features in a soft-attention manner by using K Self-aligned Part Mask Detectors. Regarding the main model branches as a generator, a discriminator is employed to determine whether the representation vector comes from the visual modality or the textual modality. With Adversarial Loss training, ASPD-Net can learn more robust representations, as long as it successfully tricks the discriminator. Experimental results demonstrate that the proposed ASPD-Net outperforms the previous methods and achieves the state-of-the-art performance on the CUHK-PEDES and RSTPReid datasets.
Many previous methods on text-based person retrieval tasks are devoted to learning a latent common space mapping, with the purpose of extracting modality-invariant features from both visual and textual modality. Nevertheless, due to the complexity of high-dimensional data, the unconstrained mapping paradigms are not able to properly catch discriminative clues about the corresponding person while drop the misaligned information. Intuitively, the information contained in visual data can be divided into person information (PI) and surroundings information (SI), which are mutually exclusive from each other. To this end, we propose a novel Deep Surroundings-person Separation Learning (DSSL) model in this paper to effectively extract and match person information, and hence achieve a superior retrieval accuracy. A surroundings-person separation and fusion mechanism plays the key role to realize an accurate and effective surroundings-person separation under a mutually exclusion constraint. In order to adequately utilize multi-modal and multi-granular information for a higher retrieval accuracy, five diverse alignment paradigms are adopted. Extensive experiments are carried out to evaluate the proposed DSSL on CUHK-PEDES, which is currently the only accessible dataset for text-base person retrieval task. DSSL achieves the state-of-the-art performance on CUHK-PEDES. To properly evaluate our proposed DSSL in the real scenarios, a Real Scenarios Text-based Person Reidentification (RSTPReid) dataset is constructed to benefit future research on text-based person retrieval, which will be publicly available.