Zero-watermarking is an emerging distortion-free copyright protection method for volumetric medical images. However, achieving both robustness against various malicious attacks and distinguishability between individual images remains challenging. In this article, we propose a novel attack-defending contrastive learning zero- watermarking (ADCL-ZW) scheme to tackle the above challenge using deep learning-based representations. In our approach, we design an attack-defending data enrichment mechanism to enhance the watermarking robustness by generating a large number of image samples under various watermarking attacks. Subsequently, features for both watermarking distinguishability and robustness are enhanced through application of a contrastive loss. In particular, we implement a dual-stream Siamese network architecture to effectively handle both signal attacks and geometric attacks in order to enhance the watermarking performance. Experimental results demonstrate that ADCL-ZW achieves stronger watermarking robustness and a better tradeoff between watermarking robustness and distinguishability compared with state-of-the art zero-watermarking methods. One of the highlighted metrics is that the false-negative rate of ADCL-ZW achieves 0.01 when a fixed false-positive rate is set to 1%, which is more than 13.3 times better than the benchmark methods.
Class activation maps (CAMs) have emerged as a popular technique to improve model interpretability of deep learning-based models. While existing CAM methods are able to extract salient semantic regions to provide high-confidence pseudo-labels for downstream tasks such as semantic segmentation, they are less effective when dealing with multi-object scenes. In this paper, we design a multi-channel weight assignment scheme that learns from both positive and negative regions to yield an improved CAM model for images comprising multiple objects. We demonstrate the effectiveness of our proposed method on two new data sets, a cat-and-dog dataset and a PASCAL VOC 2012-based multi-object dataset, and show it to compare favourably with other state-of-the-art CAM methods, outperforming them in terms of both mIoU and inter-object activation ratio (IAR), a new evaluation measure proposed to evaluate CAM performance in multi-object scenes.
Sign language translation (SLT) has attracted significant interest both from research and industry, enabling convenient communications with the deaf-mute community. While recent transformer-based models have shown improved sign translation performance, it is still under-explored how to design an efficient transformer-based deep network architecture that effectively extracts joint visual-text features by exploiting multi-level spatial and temporal contextual information. In this paper, we propose heterogeneous attention based transformer(HAT), a novel SLT model to generate attentions from diverse spatial and temporal contextual levels. Specifically, the proposed light dual-stream sparse attention-based module yields more effective visual-text representations compared to conventional transformers. Extensive experiments demonstrate that our HAT achieves state-of-the-art performance on the challenging PHOENIX2014T benchmark dataset with a BLEU-4 score of 25.33 on the test set.
Despite the success of watermarking technique for protecting depth image-based rendering (DIBR) 3-D videos, existing methods still can hardly ensure the robustness against geometric attacks, lossless video quality, and distinguishability between different videos simultaneously. In this article, we propose a novel zero-watermarking scheme to address this challenge. Specifically, we design CT-SVD features to ensure both distinguishability and robustness against signal processing and DIBR conversion attacks. In addition, a logistic–logistic chaotic system is utilized to encrypt features for the enhanced security. Moreover, a rectification mechanism based on salient map detection and SIFT matching is designed to resist geometric attacks. Finally, we establish an attention-based fusion mechanism to explore the complementary robustness of rectified and unrectified features. Experimental results demonstrate that our proposed method outperforms the existing schemes in terms of losslessness, distinguishability, and robustness against geometric attacks.
Photorealistic style transfer is the task of transferring the artistic style of an image onto a content target, producing a result that is plausibly taken with a camera. Recent approaches, based on deep neural networks, produce impressive results but are either too slow to run at practical resolutions, or still contain objectionable artifacts. We propose a new end-to-end model for photorealistic style transfer that is both fast and inherently generates photorealistic results. The core of our approach is a feed-forward neural network that learns local edge-aware affine transforms that automatically obey the photorealism constraint. When trained on a diverse set of images and a variety of styles, our model can robustly apply style transfer to an arbitrary pair of input images. Compared to the state of the art, our method produces visually superior results and is three orders of magnitude faster, enabling real-time performance at 4K on a mobile phone. We validate our method with ablation and user studies.
Machine-learning excels in many areas with well-defined goals. However, a clear goal is usually not available in art forms, such as photography. The success of a photograph is measured by its aesthetic value, a very subjective concept. This adds to the challenge for a machine learning approach. We introduce Creatism, a deep-learning system for artistic content creation. In our system, we break down aesthetics into multiple aspects, each can be learned individually from a shared dataset of professional examples. Each aspect corresponds to an image operation that can be optimized efficiently. A novel editing tool, dramatic mask, is introduced as one operation that improves dramatic lighting for a photo. Our training does not require a dataset with before/after image pairs, or any additional labels to indicate different aspects in aesthetics. Using our system, we mimic the workflow of a landscape photographer, from framing for the best composition to carrying out various post-processing operations. The environment for our virtual photographer is simulated by a collection of panorama images from Google Street View. We design a "Turing-test"-like experiment to objectively measure quality of its creations, where professional photographers rate a mixture of photographs from different sources blindly. Experiments show that a portion of our robot's creation can be confused with professional work.
We improves colorization-based image compression by sparsely sampling color points on a semi-regular grid and compressing them using JPEG. We generate variations of sampling locations based on extreme gray-scale values to to further improve PSNR.
We present MethMorph, a system for producing realistic simulations of how drug-free people would look if they used methamphetamine. Significant weight loss and facial lesions are common side effects of meth usage. MethMorph fully automates the process of thinning the face and applying lesions to healthy faces. We combine several recently-developed detection methods such as Viola-Jones based cascades and Lazy Snapping to localize facial features in healthy faces. We use the detected facial features in our method for thinning the face. We then synthesize a new facial texture, which contains lesions and major wrinkles. We apply this texture to the thinned face. We test MethMorph using a database of healthy faces, and we conclude that MethMorph produces realistic meth simulation images.
This paper introduces a practical approach to register large-scale GIS imagery to a database of road vectors automatically. The proposed approach breaks the global alignment problem into a set of localized domains (tiles). Within each tile, the displacement between imagery and vectors is approximated by a translation. Finally, a global thin-plate-spline warp based on these local approximations is applied to register the imagery to the vector data. The critical step in this approach is a fully automatic algorithm to compute the best imagery-to-vectors translation within a tile. The proposed algorithm performs vector-guided extraction of road features, aggregates features obtained in the neighborhood of multiple vectors, and then estimates the best translation through a least-squares optimization applied to a selected subset of the aggregated features. It also computes a confidence value for each processed image tile, so that a human operator can easily find out the places where the automatic approach has encountered difficulties, if necessary. The algorithm has been tested on hundreds of production satellite images of different countries. It has correctly registered over 80 percent of the imagery, and consistently reported low confidence values for the rest.
Shape deformation is a common practice in digital image editing, but can unrealistically stretch or compress texture detail. We propose an image editing system that decouples feature position from pixel color generation, by resynthesizing texture from the source image to preserve its detail and orientation around a new feature curve location. We introduce a new distortion to patch-based texture synthesis that aligns texture features with image features. A dense correspondence field between source and target images generated by the control curves then guides texture synthesis.
An MPEG-like method is developed for the compressed transmission of time-varying 3-D flow datasets that emphasizes on feature preservation. Key frames of the flow motion are compressed using a harmonic analysis of the flow. A novel bi-directional advection model is then used to approximate intermediate frames. Key features like vortical structures and shocks lost in compression are reconstructed during visualization. Textureshop is an image editing system that applies texture onto a surface in a photograph. Shape from shading is used to approximate normal field on the surface. Then texture synthesis is applied on that surface with texture coordinates deformed by the recovered normal field. The result is a texture that follows the undulation of the surface. Rototexture is a normal based video editing system that allows a user to apply a time-coherent texture to a surface depicted in the raw video from a single uncalibrated camera. Our system uses the recovered normal field to deform the texture so that it plausibly adheres to the undulations of the depicted surface. The texture mapping method uses a spring model to control the behavior of the texture image as it is deformed to match the evolving normal field through the video. The texture synthesis method uses a coarse optical flow to advect clusters of pixels corresponding to patches of similarly oriented surface points. We propose an image editing system that decouples feature position from pixel color generation to achieve a morph that preserves texture detail and orientation near the dragged silhouette, synthesized using the original image as an anisotropic texture. We introduce a new distortion to patch-based texture synthesis that aligns texture features with image features. A dense correspondence field between source and target images generated by the control curves then guides texture synthesis.
We propose a video editing system that allows a user to apply a time-coherent texture to a surface depicted in the raw video from a single uncalibrated camera, including the surface texture mapping of a texture image and the surface texture synthesis from a texture swatch. Our system avoids the construction of a 3D shape model and instead uses the recovered normal field to deform the texture so that it plausibly adheres to the undulations of the depicted surface. The texture mapping method uses the nonlinear least-squares optimization of a spring model to control the behavior of the texture image as it is deformed to match the evolving normal field through the video. The texture synthesis method uses a coarse optical flow to advect clusters of pixels corresponding to patches of similarly oriented surface points. These clusters are organized into a minimum advection tree to account for the dynamic visibility of clusters. We take a rather crude approach to normal recovering and optical flow estimation, yet the results are robust and plausible for nearly diffuse surfaces such as faces and t-shirts.
Material replacement has wide application throughout the entertainment industry, particularly for post-production make-up application or wardrobe adjustment. More generally, any low-cost mock-up object can be processed to have the appearance of expensive, high-quality materials. We demonstrate a new system that allows fast, intuitive material replacement in photographs. We extend recent work in object selection and fast texture synthesis, as well as develop a novel approach to shape-from-shading capable of handling objects with albedo changes. Each component of our system runs with interactive speed, allowing for easy experimentation and refinement of results.
We combine existing techniques for shape-from-shading and texture synthesis to create a new tool for texturing objects in photographs. Our approach clusters pixels with similar recovered normals into patches on which texture is synthesized. Distorting the texture based on the recovered normals creates the illusion that the texture adheres to the undulations of the photographed surface. Inconsistencies in the recovered surface are disguised by the graphcut blending of the individually textured patches. Further applications include the generation of detail on manually-shaded painting, extracting and synthesizing a displacement map from a texture swatch, and the embossed transfer of normals from one image to another, which would be difficult to create with current image processing packages.
We combine existing techniques for shape-from-shading and texture synthesis to create a new tool for texturing objects in photographs. Our approach clusters pixels with similar recovered normals into patches on which texture is synthesized. Distorting the texture based on the recovered normals creates the illusion that the texture adheres to the undulations of the photographed surface. Inconsistencies in the recovered surface are disguised by the graphcut blending of the individually textured patches. Further applications include the generation of detail on manually-shaded painting, extracting and synthesizing a displacement map from a texture swatch, and the embossed transfer of normals from one image to another, which would be difficult to create with current image processing packages.
Complex virtual environments can be simulated with physical or procedural motion. Physical motion is more realistic, but requires the integration of an ordinary differential equation from an initial state. Procedural motion has the advantage of being randomly accessible, though it is not physically based. We combine these techniques into a multiresolution model. Acoarse-level simulation is sampled to yield keyframes. Then a physical motion between keyframes is calculated and distorted to meet its boundary conditions. We demonstrate these ideas with a simple virtual world consisting of wind blowing trees and leaves that can beentered and experienced efficiently at any point in time.