In aerial images, numerous small objects with low resolution and sparse external features present significant challenges for object detection. Existing methods typically focus on single improvement strategies to enhance detection performance, often overlooking the potential for collaborative optimization within the overall network structure. Moreover, complex network architectures considerably increase resource consumption, making efficient deployment on edge devices difficult. To address these challenges, this study proposes a lightweight Aggregation and Recalibration Network (ARNet) that optimizes the model construction across feature extraction, feature aggregation, and feature alignment. Specifically, we design three key components: the Adaptive Cross Stage Partial Network (ACSPNet), the Multi-Scale Aggregation Reconstruction (MSAR) module, and the Feature Recalibration (FR) module. ACSPNet effectively prevents the loss of critical information during feature extraction; the MSAR module captures semantic information to complement the details of objects with weak feature representations; and the FR module aligns offsets generated by feature interactions. Experimental results show that ARNet achieves notable improvements in mAP50 on the VisDrone2019-DET, AI-TOD, and TinyPerson datasets, with increases of 5.0%, 1.5%, and 6.3%, respectively, compared to baseline algorithms. Furthermore, ARNet reduces the number of parameters by 73.4%, reaching a compact size of only 2.6M. These results highlight that ARNet effectively balances detection accuracy and model size, making it particularly suitable for UAV-based object detection tasks.
This paper presents a novel image fusion method designed to enhance the integration of infrared and visible images through the use of a residual attention mechanism. The primary objective is to generate a fused image that effectively combines the thermal radiation information from infrared images with the detailed texture and background information from visible images. To achieve this, we propose a multi-level feature extraction and fusion framework that encodes both shallow and deep image features. In this framework, deep features are utilized as queries, while shallow features function as keys and values within a residual cross-attention module. This architecture enables a more refined fusion process by selectively attending to and integrating relevant information from different feature levels. Additionally, we introduce a dynamic feature preservation loss function to optimize the fusion process, ensuring the retention of critical details from both source images. Experimental results demonstrate that the proposed method outperforms existing fusion techniques across various quantitative metrics and delivers superior visual quality.
Genotype imputation is essential for medical genomics studies. Herein, we present the STROMICS imputation reference panel, constructed from high-depth whole-genome sequencing (WGS) data of 10,241 Chinese individuals. It includes 53,061,655 single-nucleotide variants and insertion-deletions, spanning 22 autosomes and the X chromosome. Imputation performance of the STROMICS and seven other reference panels was compared using WGS data from 159 individuals. STROMICS panel outperformed others in imputation quality, and in genome-wide population- and individual-level accuracy. Validation using 301 Chinese individuals from the 1000 Genomes Project demonstrated STROMICS achieving high imputation accuracy. Among the three Chinese subgroups, the STROMICS reference panel yielded the highest accuracy in Han Chinese in Beijing samples. Notably, STROMICS outperformed all the other panels for the insertion-deletion imputation. When imputing stroke-risk variants and their closely linked variants with STROMICS, only a small statistically significant difference in sensitivity was observed between diseased and healthy individuals for variants closely linked to stroke-risk variants. Furthermore, calculated using pruned variants, the genetic distances between diseased and healthy groups remained largely unchanged before and after imputation. Collectively, these findings indicate that the source material used to construct the STROMICS reference panel has minimal impact on its imputation performance. Finally, we demonstrated high accuracy of STROMICS for genotype imputation in the X-unique region. In conclusion, the STROMICS panel provides a high-quality reference for imputing genotypes across autosomes and the X chromosome in the Chinese population.
Scaling up model and data size has been quite successful for the evolution of LLMs. However, the scaling law for the diffusion based text-to-image (T2I) models is not fully explored. It is also unclear how to efficiently scale the model for better performance at reduced cost. The different training settings and expensive training cost make a fair model comparison extremely difficult. In this work, we empirically study the scaling properties of diffusion based T2I models by performing extensive and rigours ablations on scaling both denoising backbones and training set, including training scaled UNet and Transformer variants ranging from 0.4B to 4B parameters on datasets upto 600M images. For model scaling, we find the location and amount of cross attention distinguishes the performance of existing UNet designs. And increasing the transformer blocks is more parameter-efficient for improving text-image alignment than increasing channel numbers. We then identify an efficient UNet variant, which is 45% smaller and 28% faster than SDXL's UNet. On the data scaling side, we show the quality and diversity of the training set matters more than simply dataset size. Increasing caption density and diversity improves text-image alignment performance and the learning efficiency. Finally, we provide scaling functions to predict the text-image alignment performance as functions of the scale of model size, compute and dataset size.
We empirically study the scaling properties of various Diffusion Transformers (DiTs) for text-to-image generation by performing extensive and rigorous ablations, including training scaled DiTs ranging from 0.3B upto 8B parameters on datasets up to 600M images. We find that U-ViT, a pure self-attention based DiT model provides a simpler design and scales more effectively in comparison with cross-attention based DiT variants, which allows straightforward expansion for extra conditions and other modalities. We identify a 2.3B U-ViT model can get better performance than SDXL UNet and other DiT variants in controlled setting. On the data scaling side, we investigate how increasing dataset size and enhanced long caption improve the text-image alignment performance and the learning efficiency.
Retinex-based methods have achieved significant progress in enhancing low-light images benefits for its disentanglement property. However, existing methods either ignore the semantic priors or randomly leverage them only in the spatial domain, which leads to insufficient coupling and limits the performance gains. Considering the lightness mainly exists in the amplitude component and the rest is related to the phase component, making it optimal to combine the Retinex decomposition with the Fourier transform to achieve customized restoration. In this paper, we propose a novel method RetinexFour tailored for low-light image enhancement. Specifically, it consists of Phase-Guided Multi-head Self-Attention (PG-MSA) and Local Spatial Attention (LSA) to allow for the reconstruction of structure from spatial-frequency perspectives. To achieve exposure correction, we introduce Selective Amplitude feature Fusion (SAFF) by combining the original and complementary amplitude to achieve global lightness adjustment. Extensive experiments demonstrate the superiority of our method over other SOTA methods on four benchmark datasets.
Model selection is essential for reducing the search cost of the best pre-trained model over a large-scale model zoo for a downstream task. After analyzing recent hand-designed model selection criteria with 400+ ImageNet pre-trained models and 40 downstream tasks, we find that they can fail due to invalid assumptions and intrinsic limitations. The prior knowledge on model capacity and dataset also can not be easily integrated into the existing criteria. To address these issues, we propose to convert model selection as a recommendation problem and to learn from the past training history. Specifically, we characterize the meta information of datasets and models as features, and use their transfer learning performance as the guided score. With thousands of historical training jobs, a recommendation system can be learned to predict the model selection score given the features of the dataset and the model as input. Our approach enables integrating existing model selection scores as additional features and scales with more historical data. We evaluate the prediction accuracy with 22 pre-trained models over 40 downstream tasks. With extensive evaluations, we show that the learned approach can outperform prior hand-designed model selection methods significantly when relevant training history is available.
Ischemic strokes (IS) and transient ischemic attacks (TIA) account for approximately 80% of all strokes and are leading causes of death worldwide. Assessing the risk of recurrence or functional impairment in IS and TIA patients is essential to both acute phase treatment and secondary prevention. Current risk prediction systems that rely on clinical parameters alone without leveraging imaging data have only modest performance. Herein, a deep learning‐based risk prediction system (RPS) is developed to predict the probability of stroke recurrence or disability (i.e., deep‐learning stroke recurrence risk score, SRR score). Then, Kaplan–Meier analysis to evaluate the ability of SRR score to stratify patients at stroke recurrence risk is discussed. Using 15 166 Third China National Stroke Registry (CNSR‐III) cases, the RPS's receiver operating characteristic curve (AUC) values of 0.850 for 14 day TIA recurrence prediction and 0.837 for 3 month IS disability prediction are used. Among patients deemed high risk by SRR score, 22.9% and 24.4% of individuals with TIA and IS respectively have stroke recurrence within 1 year, which are significantly higher than the rates in low‐risk individuals. Deep learning‐based RPS can outperform conventional risk scores and has the potential to assist accurate prognostication in stroke patients to optimize management.
Adapting pre-trained models with broad capabilities has become standard practice for learning a wide range of downstream tasks. The typical approach of fine-tuning different models for each task is performant, but incurs a substantial memory cost. To efficiently learn multiple down-stream tasks we introduce Task Adaptive Parameter Sharing (TAPS), a simple method for tuning a base model to a new task by adaptively modifying a small, task-specific subset of layers. This enables multi-task learning while minimizing the resources used and avoids catastrophic forgetting and competition between tasks. TAPS solves a joint optimization problem which determines both the layers that are shared with the base model and the value of the task-specific weights. Further, a sparsity penalty on the number of active layers promotes weight sharing with the base model. Compared to other methods, TAPS retains a high accuracy on the target tasks while still introducing only a small number of task-specific parameters. Moreover, TAPS is agnostic to the particular architecture used and requires only minor changes to the training scheme. We evaluate our method on a suite of fine-tuning tasks and architectures (ResNet, DenseNet, ViT) and show that it achieves state-of-the-art performance while being simple to implement.
In view of the problem of slow speed and low accuracy of the vehicle recognition in advanced driver assistance systems, a vehicle recognition method based on pseudo invariant linear moment features and ELM is proposed. Target edge is extracted by the improved PCNN model, according to the characteristic of multiple target features, the pseudo invariant linear moment features are extracted, then ELM model is used to train and recognize the databases. The validity of the model is verified through experiments, compared with other algorithms, the recognition accuracy of pseudo invariant linear moment features and ELM vehicle recognition method is higher and the speed is faster, which provides a new way to identify the vehicle in real-time monitoring system of the vehicle.
Discovering and clustering similar trajectories is a cornerstone task for movement pattern analysis and location prediction in applications like ride-sharing, supply-chain, maps and autonomous driving. However, the existing distance computation is computationally expensive and is hard to parallelize, which makes the large-scale computation prohibitive. We propose TrajDistLearn, a unified learning-based approach for trajectory distance computation, in which the traditional point-based trajectories are converted into rasterized images, and the distance function is learned via Siamese Networks in an end-to-end way. The framework accurately learns various distance metrics for the trajectory similarity computation, including the widely used Fréchet distance, which is a computationally expensive distance metric. The efficiency gain with neural network approximation is significant. Our approach achieves at least a 3000x speed-up on GPU and a 40x speed-up on CPU in comparison with naive Fréchet distance computation. In addition, our approach's computational overhead is independent of the sampling rate of the trajectories. Extensive experiments on real-world trajectory datasets demonstrate the effectiveness and efficiency of TrajDistLearn.
Fine-tuning from a collection of models pre-trained on different domains (a"model zoo") is emerging as a technique to improve test accuracy in the low-data regime. However, model selection, i.e. how to pre-select the right model to fine-tune from a model zoo without performing any training, remains an open topic. We use a linearized framework to approximate fine-tuning, and introduce two new baselines for model selection -- Label-Gradient and Label-Feature Correlation. Since all model selection algorithms in the literature have been tested on different use-cases and never compared directly, we introduce a new comprehensive benchmark for model selection comprising of: i) A model zoo of single and multi-domain models, and ii) Many target tasks. Our benchmark highlights accuracy gain with model zoo compared to fine-tuning Imagenet models. We show our model selection baseline can select optimal models to fine-tune in few selections and has the highest ranking correlation to fine-tuning accuracy compared to existing algorithms.
Discovering and clustering similar trajectories is a cornerstone task for movement pattern analysis and location prediction in applications like ride-sharing, supply-chain, maps and autonomous driving. However, the existing distance computation is computationally expensive and is hard to parallelize, which makes the large-scale computation prohibitive. We propose TrajDistLearn, a unified learning-based approach for trajectory distance computation, in which the traditional point-based trajectories are converted into rasterized images, and the distance function is learned via Siamese Networks in an end-to-end way. The framework accurately learns various distance metrics for the trajectory similarity computation, including the widely used Fréchet distance, which is a computationally expensive distance metric. The efficiency gain with neural network approximation is significant. Our approach achieves at least a 3000x speed-up on GPU and a 40x speed-up on CPU in comparison with naive Fréchet distance computation. In addition, our approach's computational overhead is independent of the sampling rate of the trajectories. Extensive experiments on real-world trajectory datasets demonstrate the effectiveness and efficiency of TrajDistLearn.
Fine-tuning from pre-trained ImageNet models has become the de-facto standard for various computer vision tasks. Current practices for fine-tuning typically involve selecting an ad-hoc choice of hyperparameters and keeping them fixed to values normally used for training from scratch. This paper re-examines several common practices of setting hyperparameters for fine-tuning. Our findings are based on extensive empirical evaluation for fine-tuning on various transfer learning benchmarks. (1) While prior works have thoroughly investigated learning rate and batch size, momentum for fine-tuning is a relatively unexplored parameter. We find that the value of momentum also affects fine-tuning performance and connect it with previous theoretical findings. (2) Optimal hyperparameters for fine-tuning, in particular, the effective learning rate, are not only dataset dependent but also sensitive to the similarity between the source domain and target domain. This is in contrast to hyperparameters for training from scratch. (3) Reference-based regularization that keeps models close to the initial model does not necessarily apply for dissimilar datasets. Our findings challenge common practices of fine-tuning and encourages deep learning practitioners to rethink the hyperparameters for fine-tuning.
Detection of multi-class rotated objects is a challenging task in optical remote sensing images because of large-scale variations, arbitrary orientations and complex backgrounds, etc. Most of the state-of-the-art object detectors for natural images, that use horizontal bounding boxes, are not suitable for oriented objects in remote sensing images. In this paper, we propose an end-to-end cascade detector that can effectively detect rotated objects in complex remote sensing images. Specifically, a feature fusion block is designed to capture features with more details. Meanwhile, a supervised spatial attention mechanism is adopted to improve performance in detecting objects with complex backgrounds by weakening noise and enhancing object regions. Finally, to obtain more accurate object position, a cascade of multi-step detection subnet is implemented to refine anchors. Experiments using a publicly available remote sensing dataset DOTA show that our object detector achieves superior performance over other state-of-the-art approaches.
Global context modeling has been used to achieve better performance in various computer-vision-related tasks, such as classification, detection, segmentation and multimedia retrieval applications. However, most of the existing global mechanisms display problems regarding convergence during training. In this paper, we propose a novel gated global module (GGM) that is lightweight and yet effective in terms of achieving better integration of global information in relation to feature representation. Regarding the original structure of the network as a local block, our module infers global information in parallel with local information, and then a gate function is applied to generate global guidance which is applied to the output of the local module to capture representative information. The proposed GGM can be easily integrated with common CNN architectures and is training friendly. We used a classification task as an example to verify the effectiveness of the proposed GGM, and extensive experiments on ImageNet and CIFAR demonstrated that our method can be widely applied and is conducive to integrating global information into common networks.
Impressive progresses have been achieved in object detection for images by convolution neural networks. However, a robust multi-class object detection method is still one of the great challenges for remote sensing images. Due to the great diversity of scale, orientation, density and background of objects, most advanced object detection algorithms in natural scenes usually suffer a sharp decline in remote sensing images detection. To solve these problems, we proposed a Single-stage Multi-class Object Detection (SMOD) method, aiming at remote sensing images, which can be trained from scratch and detect multi-class objects quickly and precisely. The proposed method introduces a novel Feature Reuse and Attention (FRA) structure as a key module of feature extraction backbone, which combines SE Attention module and dense Feature Reuse connection. Especially, a multiclass detection structure is proposed to learn from multi-scale, multi-level feature map and get effective attention representation for multi-class remote sensing object detection. SMOD can be trained from scratch without pre-trained network stably and converge well simply by employing batch normalization throughout the network. Experiments show that our trainingfrom-scratch method can obtain better performance compared with some state-of-art algorithms on public multi-class remote sensing dataset AIIA2018-6.
Most video surveillance systems use both RGB and infrared cameras, making it a vital technique to re-identify a person cross the RGB and infrared modalities. This task can be challenging due to both the cross-modality variations caused by heterogeneous images in RGB and infrared, and the intra-modality variations caused by the heterogeneous human poses, camera views, light brightness, etc. To meet these challenges a novel feature learning framework, HPILN, is proposed. In the framework existing single-modality re-identification models are modified to fit for the cross-modality scenario, following which specifically designed hard pentaplet loss and identity loss are used to improve the performance of the modified cross-modality re-identification models. Based on the benchmark of the SYSU-MM01 dataset, extensive experiments have been conducted, which show that the proposed method outperforms all existing methods in terms of Cumulative Match Characteristic curve (CMC) and Mean Average Precision (MAP).
A new model acceleration algorithm of convolutional neural network (CNN) was proposed based on filters pruning in order to promote the compression and acceleration of the CNN model. The computational cost could be effectively reduced by calculating the standard deviation of filters in the convolutional layer to measure its importance and pruning filters with less influence on the accuracy of the neural network and its corresponding feature map. The algorithm did not cause the network to be sparsely connected unlike the method of pruning weight value, so there was no need of the support of special sparse convolution libraries. The experimental results based on the CIFAR-10 dataset show that the filters pruning algorithm can accelerate the VGG-16 and ResNet-110 models by more than 30%. Results can be close to or reach the accuracy of the original model by fine-tuning the inherited pre-training parameters.
Neural network training relies on our ability to find "good" minimizers of highly non-convex loss functions. It is well-known that certain network architecture designs (e.g., skip connections) produce loss functions that train easier, and well-chosen training parameters (batch size, learning rate, optimizer) produce minimizers that generalize better. However, the reasons for these differences, and their effects on the underlying loss landscape, are not well understood. In this paper, we explore the structure of neural loss functions, and the effect of loss landscapes on generalization, using a range of visualization methods. First, we introduce a simple "filter normalization" method that helps us visualize loss function curvature and make meaningful side-by-side comparisons between loss functions. Then, using a variety of visualizations, we explore how network architecture affects the loss landscape, and how training parameters affect the shape of minimizers.