Single image super-resolution (SISR) is still an important while challenging task. Existing methods usually ignore the diversity of generated Super-Resolution (SR) images. The fine details of the corresponding high-resolution (HR) images cannot be confidently recovered due to the degradation of detail in low-resolution (LR) images. To address the above issue, this paper presents a flow-based multi-scale learning network (FMLnet) to explore the diverse mapping spaces for SR. First, we propose a multi-scale learning block (MLB) to extract the underlying features of the LR image. Second, the introduced pixel-wise multi-head attention allows our model to map multiple representation subspaces simultaneously. Third, by employing a normalizing flow module for a given LR input, our approach generates various stochastic SR outputs with high visual quality. The trade-off between fidelity and perceptual quality can be controlled. Finally, the experimental results on five datasets demonstrate that the proposed network outperforms the existing methods in terms of diversity, and achieves competitive PSNR/SSIM results. Code is available at https://github.com/qianyuwu/FMLnet.
Creating high-dynamic videos such as motion-rich actions and sophisticated visual effects poses a significant challenge in the field of artificial intelligence. Unfortunately, current state-of-the-art video generation methods, primarily focusing on text-to-video generation, tend to produce video clips with minimal motions despite maintaining high fidelity. We argue that relying solely on text instructions is insufficient and suboptimal for video generation. In this paper, we introduce PixelDance, a novel approach based on diffusion models that incorporates image instructions for both the first and last frames in conjunction with text instructions for video generation. Comprehensive experimental results demonstrate that PixelDance trained with public data exhibits significantly better proficiency in synthesizing videos with complex scenes and intricate motions, setting a new standard for video generation.
Generating rich and controllable motion is a pivotal challenge in video synthesis. We propose Boximator, a new approach for fine-grained motion control. Boximator introduces two constraint types: hard box and soft box. Users select objects in the conditional frame using hard boxes and then use either type of boxes to roughly or rigorously define the object's position, shape, or motion path in future frames. Boximator functions as a plug-in for existing video diffusion models. Its training process preserves the base model's knowledge by freezing the original weights and training only the control module. To address training challenges, we introduce a novel self-tracking technique that greatly simplifies the learning of box-object correlations. Empirically, Boximator achieves state-of-the-art video quality (FVD) scores, improving on two base models, and further enhanced after incorporating box constraints. Its robust motion controllability is validated by drastic increases in the bounding box alignment metric. Human evaluation also shows that users favor Boximator generation results over the base model.
This paper reviews the NTIRE 2021 challenge on learning the super-Resolution space. It focuses on the participating methods and final results. The challenge addresses the problem of learning a model capable of predicting the space of plausible super-resolution (SR) images, from a single low-resolution image. The model must thus be capable of sampling diverse outputs, rather than just generating a single SR image. The goal of the challenge is to spur research into developing learning formulations and models better suited for the highly ill-posed SR problem. And thereby advance the state-of-the-art in the broader SR field. In order to evaluate the quality of the predicted SR space, we propose a new evaluation metric and perform a comprehensive analysis of the participating methods. The challenge contains two tracks: 4× and 8 scale factor. In total, 11 teams competed in the final testing× phase.
From March 2018 to February 2019, quantitative detection was made of 102 kinds of atmospheric volatile organic compounds (VOCs) using online gas chromatography in Ezhou City. We compared and analyzed the composition, seasonal variation, and diurnal variation of VOCs. Using maximum incremental reactivity (MIR), we estimated the ozone generation potential (OFP) of VOCs. The results show that the annual average volume fraction of atmospheric VOCs in Ezhou is (30.78±15.89)×10-9, and is overall higher in winter than summer, represented by alkane > oxygen > halogenated hydrocarbon > olefin > aromatic hydrocarbon > alkyne. The night volume fraction is higher than in the daytime, and overall the distribution is "double peak". The aromatic hydrocarbons, halogenated hydrocarbons, and OVOCs appear as a "third peak" at 00:00-02:00. Aromatic hydrocarbons and olefins contribute more to the OFP potential of VOCs, with contribution rates of 35.45% and 29.5%, respectively. The highest contribution rate to OFP is ethylene, reaching 24.217%. Analysis of VOC characteristic species found that vehicle exhaust fumes and solvent volatilization are the main sources of VOCs in Ezhou. Of these, motor vehicle emissions are the most important source. Controlling Ezhou's motor vehicle emissions helps to reduce the composition of atmospheric VOCs, thereby reducing ozone production.