Integrating watermarks into generative images is a critical strategy for protecting intellectual property and enhancing artificial intelligence security. This paper proposes Plug-in Generative Watermarking (PiGW) as a general framework for integrating watermarks into generative images. More specifically, PiGW embeds watermark information into the initial noise using a learnable watermark embedding network and an adaptive frequency spectrum mask. Furthermore, it optimizes training costs by gradually increasing timesteps. Extensive experiments demonstrate that PiGW enables embedding watermarks into the generated image with negligible quality loss while achieving true invisibility and high resistance to noise attacks. Moreover, PiGW can serve as a plugin for various commonly used generative structures and multimodal generative content types. Finally, we demonstrate how PiGW can also be utilized for detecting generated images, contributing to the promotion of secure AI development. The project code will be made available on GitHub.
Text image machine translation (TIMT) has been widely used in various real-world applications, which translates source language texts in images into another target language sentence. Existing methods on TIMT are mainly divided into two categories: the recognition-then-translation pipeline model and the end-to-end model. However, how to transfer knowledge from the pipeline model into the end-to-end model remains an unsolved problem. In this paper, we propose a novel Multi-Teacher Knowledge Distillation (MTKD) method to effectively distillate knowledge into the end-to-end TIMT model from the pipeline model. Specifically, three teachers are utilized to improve the performance of the end-to-end TIMT model. The image encoder in the end-to-end TIMT model is optimized with the knowledge distillation guidance from the recognition teacher encoder, while the sequential encoder and decoder are improved by transferring knowledge from the translation sequential and decoder teacher models. Furthermore, both token and sentence-level knowledge distillations are incorporated to better boost the translation performance. Extensive experimental results show that our proposed MTKD effectively improves the text image translation performance and outperforms existing end-to-end and pipeline models with fewer parameters and less decoding time, illustrating that MTKD can take advantage of both pipeline and end-to-end models.
Text image machine translation (TIMT) aims to translate texts embedded in images from one source language to another target language. Existing methods, both two-stage cascade and one-stage end-to-end architectures, suffer from different issues. The cascade models can benefit from the large-scale optical character recognition (OCR) and MT datasets but the two-stage architecture is redundant. The end-to-end models are efficient but suffer from training data deficiency. To this end, in our paper, we propose an end-to-end TIMT model fully making use of the knowledge from existing OCR and MT datasets to pursue both an effective and efficient framework. More specifically, we build a novel modal adapter effectively bridging the OCR encoder and MT decoder. End-to-end TIMT loss and cross-modal contrastive loss are utilized jointly to align the feature distribution of the OCR and MT tasks. Extensive experiments show that the proposed method outperforms the existing two-stage cascade models and one-stage end-to-end models with a lighter and faster architecture. Furthermore, the ablation studies verify the generalization of our method, where the proposed modal adapter is effective to bridge various OCR and MT models. (Our codes are available at: https://github.com/EriCongMa/E2TIMT )
Forward osmosis (FO) membrane fouling is one of the main reasons that hinder the further application of FO technology in the treatment of dye wastewater. To alleviate membrane fouling, a conductive coal carbon-based substrate and polydopamine nanoparticles (PDA NPs) interlayer composite FO membrane (CPFO) was prepared by interfacial polymerization (IP). CPFO-10 membrane prepared by depositing 10 mL of PDA NPs solution exhibited an optimum performance with water flux of 7.56 L/(m2h) for FO mode and 10.75 L/(m2h) for pressure retarded osmosis (PRO) mode, respectively. For rhodamine B and chrome black T dye wastewater treatment, the water flux losses were reduced by 21.6%, and 14.5% under the voltages of +1.5 V, and -1.5 V, respectively, compared with no voltage applied after the device was operated for 8 h. The applied voltage had little effect on the fouling mitigation performance of the CPFO membrane for neutral charged cresol red. After the device was operated for 4 cycles, the rejection rates of dyes wastewater treated by the CPFO membranes with applied voltage were close to 100%. The flux decline rate and flux recovery rate of CPFO membrane for rhodamine B and chrome black T wastewater treatment under application of +1.5 V and -1.5 V voltage after 4 cycles were 11.6%, 99.2%, and 16.7%, 98.9%, respectively. Therefore, the voltage-applied CPFO membrane still maintained good rejection and antifouling performance in long-term operation. This study provides a new insight into the preparation of conductive FO membranes for dye wastewater treatment and membrane fouling control.
Sodium tetraborate pentahydrate (STB) was intercalated into graphene oxide (GO) nanosheets to form a nano composite (STB@GO). Subsequently, it was self-assembled on a substrate membrane to prepare STB@GO nanofiltration membrane. The properties of the STB@GO powder samples and the nanofiltration membrane were studied using scanning electron microscopy (SEM), Fourier transform infrared spectroscopy (FTIR), X-ray photoelectron spectroscopy (XPS), X-ray diffraction (XRD), contact angle (CA), and zeta potential. When the STB concentration was 1.0 g/L in the cross-linking reaction, the membrane was described as the STB2@GO membrane and exhibited a large interlayer space (D-spacing = 1.347 nm), high hydrophilicity (CA = 22.2 degrees), and high negative potential (zeta =-18.0 mV). Meanwhile, the pure water flux of the membrane was significantly increased by 56.60% than that of the GO membrane. In addition, the STB2@GO membrane exhibited a favorable capability for dye rejection,98.52% for Evans blue (EB), 99.26% for Victoria blue B (VB), 91.94% for Alizarin yellow (AY), and 93.21% for Neutral red (NR). Furthermore, the STB2@GO membrane performed better in dye separation under various types and concentrations of dye, pH values, and ions in solution. Thus, this study provides a promising method for preparing laminated GO nanofiltration membranes for dye wastewater treatment.
Carbon dots (CDs) as a fluorescent nanomaterial plays an important role in bioimaging and light-emitting devices, and its multifunctionalization will be a promising evolution. Here, by a freely rotatable single bond to connect materials with complementary functions together, a dual-function aggregation-induced emission (AIE) CDs (V-CDs) is obtained to visually monitor the dead of cancerous cells and inhibit the production of vascular endothelial growth factor (VEGF). The fluorescence of V-CDs will change between blue and green as the changes of environment viscosity based on the restriction of intramolecular rotation (RIR) effect. For another, the VEGF content decreased by about 3 times compared with the blank group after V-CDs was introduced. The co-administration group shows the best anti-cancer effect among the rest of control groups after Taxol was involved, and the same result are achieved in vitro and vivo. The performances of V-CDs make it possible to realize the potential of visualizing anti-cancer process.
Labanotation is a professional dance notation system widely used in dance education and choreography preservation. Automatically generatingLabanotation dance scores from motion capture data can save a huge amount of manual time and effort. Recently, the sequence-to-sequence (seq2seq) model is applied to the automatic Labanotation generation. This model is based on an encoder-decoder structure, which encodes the input motion sequence to a fixed-length vector and then decodes it to generate the target sequence. However, the encoding of spatial skeleton structure of motion data is not considered in the existing work. Besides, it is challenging to align between the input motion data and the output Laban symbol sequences due to the severe imbalance of sequence lengths. Therefore, in this paper, we present a new seq2seq model for more effective Labanotation generation. In the encoder, we propose a new gesture-sensitive graph convolutional network with learned adaptive joint weights and non-physical connections to learn both spatial and temporal patterns from motion data sequences. In the decoder, we exploit motion rhythm information and propose a novel rhythm-aware attention mechanism to learn a good alignment between motion sequences and Laban symbol sequences, so that we can focus on relevant parts of the input motion sequence without searching in the whole sequence when predicting a target Laban symbol. Extensive experiments on two real-world datasets show that the proposed method achieves a better performance compared with the state-of-the-art approaches on the task of automatic Labanotation generation.
Deraining driven by semantic segmentation task is very important for autonomous driving because rain streaks and raindrops on the car window will seriously degrade the segmentation accuracy. As a pre-processing step of semantic segmentation network, a deraining network should be capable of not only removing rain in images but also preserving semantic-aware details of derained images. However, most of the state-of-the-art deraining approaches are only optimized for high PSNR and SSIM metrics without considering objective effect for high-level vision tasks. Not only that, there is no suitable dataset for such tasks. In this paper, we first design a new deraining network that contains a semantic refinement residual network (SRRN) and a novel two-stage segmentation aware joint training method. Precisely, our training method is composed of the traditional deraining training and the semantic refinement joint training. Hence, we synthesize a new segmentation-annotated rain dataset called Raindrop-Cityscapes with rain streaks and raindrops which makes it possible to test deraining and segmentation results jointly. Our experiments on our synthetic dataset and real-world dataset show the effectiveness of our approach, which outperforms state-of-the-art methods and achieves visually better reconstruction results and sufficiently good performance on semantic segmentation task.
Video object detection plays a vital role in a wide variety of computer vision applications. To deal with challenges such as motion blur, varying view-points/poses, and occlusions, we need to solve the temporal association across frames. One of the most typical solutions to maintain frame association is exploiting optical flow between consecutive frames. However, using optical flow alone may lead to poor alignment across frames due to the gap between optical flow and high-level features. In this paper, we propose an Attention-Based Temporal Context module (ABTC) for more accurate frame alignments. We first extract two kinds of features for each frame using the ABTC module and a Flow-Guided Temporal Coherence module (FGTC). Then, the features are integrated and fed to the detection network for the final result. The ABTC and FGTC are complementary to each other and can work together to obtain a higher detection quality. Experiments on the ImageNet VID dataset show that the proposed framework performs favorable against the state-of-the-art methods.
Labanotation is an important notation system for recording dances. Automatically generating Labanotation scores from motion capture data has attracted more interest in recent years. Current methods usually focus on individual movement segments and generate Labanotation symbols one by one. This requires segmenting the captured data sequence in advance. Manual segmentation will consume a lot of time and effort, while automatic segmentation may not be reliable enough. In this paper, we propose a sequence-to-sequence approach that can generate Labanotation scores from unsegmented motion data sequences. First, we extract effective features from motion capture data based on body skeleton analysis. Then, we train a neural network under the encoder-decoder architecture to transform the motion feature sequences to corresponding Labanotation symbols. As such, the dance score is generated. Experiments show that the proposed method performs favorably against state-of-the-art algorithms in the automatic Labanotation generation task.
Recent years have seen a flurry of activities in designing provably efficient nonconvex procedures for solving statistical estimation problems. Due to the highly nonconvex nature of the empirical loss, state-of-the-art procedures often require proper regularization (e.g. trimming, regularized cost, projection) in order to guarantee fast convergence. For vanilla procedures such as gradient descent, however, prior theory either recommends highly conservative learning rates to avoid overshooting, or completely lacks performance guarantees. This paper uncovers a striking phenomenon in nonconvex optimization: even in the absence of explicit regularization, gradient descent enforces proper regularization implicitly under various statistical models. In fact, gradient descent follows a trajectory staying within a basin that enjoys nice geometry, consisting of points incoherent with the sampling mechanism. This "implicit regularization" feature allows gradient descent to proceed in a far more aggressive fashion without overshooting, which in turn results in substantial computational savings. Focusing on three fundamental statistical estimation problems, i.e. phase retrieval, low-rank matrix completion, and blind deconvolution, we establish that gradient descent achieves near-optimal statistical and computational guarantees without explicit regularization. In particular, by marrying statistical modeling with generic optimization theory, we develop a general recipe for analyzing the trajectories of iterative algorithms via a leave-one-out perturbation argument. As a byproduct, for noisy matrix completion, we demonstrate that gradient descent achieves near-optimal error control — measured entrywise and by the spectral norm — which might be of independent interest.
Labanotation is a widely used dance notation system. The topic of generating Labanotation scores from captured dance motion data has attracted more research interest in recent years. Current methods usually generate Laban symbols via recognizing motion segments that each contains a single dance movement. However, they rely on the manual segmentation of raw dance motion sequences, which can cost a lot of time and effort. In this paper, we present a fully automatic framework to generate Labanotation scores from continuous motion data. First, we split the captured dance data to motion segments based on the Laban theory of body weight support transferring. Then, based on the segmented data, we utilize a network with both 1D-convolutional and recurrent layers to recognize body movements and generate Laban symbols. As such, the Labanotation score is created. Extensive experiments show that the proposed automatic framework performs favorably against previous solutions.
This paper describes the CASIA’s system for the IWSLT 2020 open domain translation task. This year we participate in both Chinese→Japanese and Japanese→Chinese translation tasks. Our system is neural machine translation system based on Transformer model. We augment the training data with knowledge distillation and back translation to improve the translation performance. Domain data classification and weighted domain model ensemble are introduced to generate the final translation result. We compare and analyze the performance on development data with different model settings and different data processing techniques.
Although there are many existing research works about the salient object detection (SOD) in RGB images, there are still many complex situations that regular RGB images cannot provide enough cues for the accurate SOD, such as the shadow effect, similar appearance between background and foreground, strong or insufficient illumination, etc. Because of the success of near-infrared spectrum in many computer vision tasks, we explore the multi-spectral SOD in the synchronized RGB images and near-infrared (NIR) images for the both simple and complex situations. We assume that the RGB SOD in the existing RGB image datasets could provide references for the multi-spectral SOD problem. In this paper, we first collect and will publicize a large multi-spectral dataset including 780 synchronized RGB and NIR image pairs for the multi-spectral SOD problem in the simple and complex situations. We model this research problem as an adversarial domain adaptation from the existing RGB image dataset (source domain) to the collected multi-spectral dataset (target domain). Experimental results show the effectiveness and accuracy of the proposed adversarial domain adaptation for the multi-spectral SOD.
Salient Object Detection (SOD) plays an important role in many image-related multimedia applications. Although there are many existing research works about the salient object detection in traditional RGB (visible-light spectrum) images, there are still many complex situations that regular RGB images cannot provide enough cues for the accurate SOD, such as the shadow effect, similar appearance between background and foreground, strong or insufficient illumination, etc. Because of the success of near-infrared spectrum in many computer vision tasks, we explore the multi-spectral SOD in the synchronized RGB images and near-infrared (NIR) images for the both simple and complex situations. We assume that the RGB SOD in the existing RGB image datasets could provide references for the multi-spectral SOD problem. In this paper, we mainly model this research problem as a deep learning based domain adaptation from the traditional RGB image data (source domain) to the multi-spectral data (target domain), and an adversarial deep domain adaptation model is proposed. We first collect and will publicize a large multi-spectral dataset, RGBN-SOD dataset, including 780 synchronized RGB and NIR image pairs for the multi-spectral SOD problem in the simple and complex situations. Intensive experimental results show the effectiveness and accuracy of the proposed deep domain adaptation for the multi-spectral SOD. Besides, due to the absence of research on the field of multi-spectral co-saliency detection, we also collect 200 synchronized RGB and NIR image pairs in addition to explore the multi-spectral co-saliency detection.
The purpose of video segmentation is to segment foreground objects from a video sequence. In this paper, we propose a CNN based method for the semi-supervised video object segmentation, where a hybrid encoder-decoder network is designed to generate pixel-wise foreground object segmentation in use of both spatial and temporal information. In order to minimize cumulative error of the network as much as possible, we develop a two-stage training scheme: alternate training and back-propagation-through-time training. Then the performances of our method and other state-of-the-art ones are compared on two annotated video segmentation databases. Furthermore, we also run an extensive ablation study to test the effects of different components from our method.
Detecting anomalies from trajectory data is an important task in video surveillance. However, it is difficult to give a precise definition of this term since trajectory data obtained from different camera views may vary in shape, direction, and spatial distribution. In this paper, we propose trajectory distance metrics based on a recurrent neural network to measure similarities and detect anomalies from trajectory data. First, we use an autoencoder to capture the dynamic features of a trajectory. The distance between two trajectories is defined by the reconstruction errors based on the learned models. We then detect anomalies based on the nearest neighbors using the proposed metric. As such, we can deal with various kinds of anomalies in different scenes and detect anomalous trajectories in either a supervised or unsupervised manner. Experiments show that the proposed algorithm performs favorably against the state-of-the-art anomaly detections on the benchmark datasets.
Automatic text summarization is a fundamental natural language processing (NLP) application that aims to condense a source text into a shorter version. The rapid increase in multimedia data transmission over the Internet necessitates multi-modal summarization (MMS) from asynchronous collections of text, image, audio, and video. In this work, we propose an extractive MMS method that unites the techniques of NLP, speech processing, and computer vision to explore the rich information contained in multi-modal data and to improve the quality of multimedia news summarization. The key idea is to bridge the semantic gaps between multi-modal content. Audio and visual are main modalities in the video. For audio information, we design an approach to selectively use its transcription and to infer the salience of the transcription with audio signals. For visual information, we learn the joint representations of text and images using a neural network. Then, we capture the coverage of the generated summary for important visual information through text-image matching or multi-modal topic modeling. Finally, all the multi-modal aspects are considered to generate a textual summary by maximizing the salience, non-redundancy, readability, and coverage through the budgeted optimization of submodular functions. We further introduce a publicly available MMS corpus in English and Chinese.1 The experimental results obtained on our dataset demonstrate that our methods based on image matching and image topic framework outperform other competitive baseline methods.
In this letter, we derive the basic conditions that should be met for the weighted signal suitable for the communication system and extend the weighted fractional Fourier transform-based classical hybrid carrier system to the generalized hybrid carrier system. Furthermore, a generalized hybrid carrier signal design is presented with equal component power. Compared with the classical hybrid carrier signal, the power distribution of the proposed system in the time-frequency plane is more uniform, which makes it more suitable for doubly selective channels. The simulation results show that over doubly selective channels, the generalized hybrid carrier signal with equal component power achieves better bit error rate performance than the classical hybrid carrier signal does.
Labanotation is a widely used notation system for recording body movements, especially dances. It has wide applications in choreography preservation, dance archiving, and so on. However, the manual creation of Labanotation scores is rather difficult and time-consuming. Therefore, research on the generation of Labanotation scores is of great interest. In this paper, we aim to generate Labanotation scores based on the motion-captured data obtained from real-world dance performances. First, to deal with challenges such as various dance movement patterns, different dancer shapes, and noises in the motion-captured data, we propose a novel feature that is invariant to anthropometric variation and body orientation. Then, we generate the notations of both lower-limb movements and upper-limb gestures. On the one hand, we utilize the hidden Markov model (HMM) to analyze the temporal dynamic characteristics of limb movements and map each lower limb movement to a corresponding dance notation. On the other hand, for upper limbs, we train a multi-class classifier based on the extremely randomized trees (Extra-Trees) to identify the notations for arm gestures. Finally, we generate the Labanotation symbols based on the above movement analysis and thus create Labanotation scores. The proposed methods can generate the spatial symbols describing directions and levels in both the support column and arm column based on motion-captured data. The generated scores are clear and reliable. Experimental results show an average recognizing accuracy of over 92 for the generated notations, which is significantly better than previous work.