Feature fusion is an effective solution for improving image retrieval performance. Although the more feature types, the better accuracy, complexity also increases. Applications in practice typically afford a limited number of feature types. Due to the strong complementarity, global and local features form an ideal combination for many fusion applications. However, the two kinds of features are intrinsically different in nature, thus cannot be fused in a straightforward way. In this work, we propose an integrated image retrieval and feature fusion framework for global and local features. It is based on inverted index fusion, a technique for efficient image retrieval. The core idea is to rank candidates by weighted voting during candidate selection, which is named pre-ranking. This procedure takes place before re-ranking, and is potentially superior to conventional late fusion. Extensive experiments on three public datasets show that the light-weight pre-ranking stage significantly contributes to accuracy, and brings substantial improvement when used together with re-ranking. Our method is robust and versatile, and can be applied to any scenario where inverted indexing is used. It is a promising technique for multimedia retrieval in the big data era.
SUMAC 2025 is the 7th edition of the workshop on analySis, Understanding and proMotion of heritAge Contents. It is held in Dublin, Ireland, on 27 October and is co-located with the 33rd ACM International Conference on Multimedia. The workshop's objective is to present and discuss the latest and most significant trends, challenges, and advances in the fields of machine learning, signal processing, multimodal techniques, and human-machine interaction. The workshop is dedicated to the valorization of cultural heritage, with an emphasis on unlocking and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field.
This work proposes a supervised machine learning algorithm for target localization in deep brain stimulation (DBS). DBS is a recognized treatment for movement disorders, such as essential tremor. The effects of DBS significantly depends on the precise choice of target location. Recent research on diffusion tensor imaging (DTI) shows that the optimal target is related to the dentato-rubro-thalamic tract (DRTT), thus DRTT analysis has become a promising approach. This technique is more accurate than conventional ones, but still too complicated for clinical scenarios, where only magnetic resonance imaging (MRI) data is available. In order to improve efficiency and utility, we consider target localization as a non-linear regression problem in a reduced-reference learning framework, and solve it with convolutional neural networks. The proposed method is light-weight, and consists of two image-based networks: one for classification and the other for localization. We model the basic workflow as an image retrieval process and define relevant performance metrics. Using DRTT analysis as groundtruths, we show that DTI-based optimal targets can be inferred from MRI data with high accuracy. For 280x220 (0.7 mm slice thickness) MRI input, our model achieves an average posterior localization error of 2.3 mm, and a median of 1.7 mm. The proposed framework is the first in the DBS domain. It is a successful application of reduced-reference learning, and may serve as a baseline for general target localization problems in DBS.
In this work, we propose a novel approach to the dating of cultural heritage objects using a bag-of-time (BoT) model, which combines visual and temporal information. The BoT model extends the traditional bag-of-words (BoW) representation, allowing multiple cultural elements in an image to be associated with different time points. We introduce the concept of cultural elements as stable and distinct data patterns, enabling the BoT model to encode time information efficiently. Furthermore, we present an aggregated bag-of-time (ABoT) model that captures the temporal distribution of image contents. Our experiments on a Buddha image dataset demonstrate the effectiveness of the proposed approach in predicting built years and enabling fine-grained dating. Additionally, we explore the applications of multi-modal image retrieval and time-based visualization, showcasing the versatility of our models in cultural heritage research and analysis.
Optical remote sensing images (RSIs) have been widely used in many applications, and one of the interesting issues about optical RSIs is the salient object detection (SOD). However, due to diverse object types, various object scales, numerous object orientations, and cluttered backgrounds in optical RSIs, the performance of the existing SOD models often degrade largely. Meanwhile, cutting-edge SOD models targeting optical RSIs typically focus on suppressing cluttered backgrounds, while they neglect the importance of edge information which is crucial for obtaining precise saliency maps. To address this dilemma, this article proposes an edge-guided recurrent positioning network (ERPNet) to pop-out salient objects in optical RSIs, where the key point lies in the edge-aware position attention unit (EPAU). First, the encoder is used to give salient objects a good representation, that is, multilevel deep features, which are then delivered into two parallel decoders, including: 1) an edge extraction part and 2) a feature fusion part. The edge extraction module and the encoder form a U-shape architecture, which not only provides accurate salient edge clues but also ensures the integrality of edge information by extra deploying the intraconnection. That is to say, edge features can be generated and reinforced by incorporating object features from the encoder. Meanwhile, each decoding step of the feature fusion module provides the position attention about salient objects, where position cues are sharpened by the effective edge information and are used to recurrently calibrate the misaligned decoding process. After that, we can obtain the final saliency map by fusing all position attention cues. Extensive experiments are conducted on two public optical RSIs datasets, and the results show that the proposed ERPNet can accurately and completely pop-out salient objects, which consistently outperforms the state-of-the-art SOD models.
Image hash algorithms generate compact binary representations that can be quickly matched by Hamming distance, thus become an efficient solution for large-scale image retrieval. This paper proposes RV-SSDH, a deep image hash algorithm that incorporates the classical VLAD (vector of locally aggregated descriptors) architecture into neural networks. Specifically, a novel neural network component is formed by coupling a random VLAD layer with a latent hash layer through a transform layer. This component can be combined with convolutional layers to realize a hash algorithm. We implement RV-SSDH as a point-wise algorithm that can be efficiently trained by minimizing classification error and quantization loss. Comprehensive experiments show this new architecture significantly outperforms baselines such as NetVLAD and SSDH, and offers a cost-effective trade-off in the state-of-the-art. In addition, the proposed random VLAD layer leads to satisfactory accuracy with low complexity, thus shows promising potentials as an alternative to NetVLAD.
Chapter 9 Perceptual Hashing for Large-Scale Multimedia Search Li Weng, Li WengSearch for more papers by this authorI-Hong Jhuo, I-Hong JhuoSearch for more papers by this authorWen-Huang Cheng, Wen-Huang ChengSearch for more papers by this author Li Weng, Li WengSearch for more papers by this authorI-Hong Jhuo, I-Hong JhuoSearch for more papers by this authorWen-Huang Cheng, Wen-Huang ChengSearch for more papers by this author Stefanos Vrochidis, Stefanos Vrochidis Information Technologies Institute, Centre for Research and Technology Hellas Thessaloniki, GreeceSearch for more papers by this authorBenoit Huet, Benoit Huet EURECOM, Sophia-Antipolis, FranceSearch for more papers by this authorEdward Chang, Edward Chang HTC Research & Healthcare, San Francisco, USASearch for more papers by this authorIoannis Kompatsiaris, Ioannis Kompatsiaris Information Technologies Institute, Centre for Research and Technology Hellas Thessaloniki, GreeceSearch for more papers by this author Book Author(s):Stefanos Vrochidis, Stefanos Vrochidis Information Technologies Institute, Centre for Research and Technology Hellas Thessaloniki, GreeceSearch for more papers by this authorBenoit Huet, Benoit Huet EURECOM, Sophia-Antipolis, FranceSearch for more papers by this authorEdward Chang, Edward Chang HTC Research & Healthcare, San Francisco, USASearch for more papers by this authorIoannis Kompatsiaris, Ioannis Kompatsiaris Information Technologies Institute, Centre for Research and Technology Hellas Thessaloniki, GreeceSearch for more papers by this author First published: 15 March 2019 https://doi.org/10.1002/9781119376996.ch9Citations: 1 AboutPDFPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShareShare a linkShare onFacebookTwitterLinked InRedditWechat Summary This chapter presents perceptual hashing technique together with a particular category of algorithms called perceptual hash algorithms. These algorithms are used for generating hash values from large-scale multimedia objects, such as images, audio, and video. The chapter focuses on unsupervised perceptual hash algorithms and supervised perceptual hash algorithms. Perceptual hashing is one of the approaches that seek compact representations of multimedia data. Perceptual hashing mainly consists of two parts: hash generation and hash verification. Hash generation is the focus of hash algorithm design. There are a few essential components: feature extraction, feature transformation, dimension reduction, quantization, and randomization. Hash verification is typically made simple in order to be fast. The basic properties of perceptual hashing are robustness and discrimination. Kernelized locality sensitive hashing is an extension of locality sensitive hashing. Semi-supervised hashing is a hash algorithm that takes both semantic relevance and "maximal bit variance" into account. Citing Literature Big Data Analytics for Large-Scale Multimedia Search RelatedInformation
When encryption and authentication techniques are applied to image or video data, sometimes it is advantageous to limit the operation to the DC DCT coefficient of each 8× 8 block in a picture. In this work, the performance of such an approach is evaluated. This problem is considered as an image quality problem, and the metric structural similarity is used to show that by authenticating the DC coefficient, about 60% of the information can be guaranteed; by encrypting the DC coefficient, about 80% of the information can be
Perceptual hashing is a promising tool for multimedia content authentication. Digital watermarking is a convenient way of data hiding. By combining the two, we get a more efficient and versatile solution. In a typical scenario, multimedia data is sent from a server to a client. The corresponding hash value is embedded in the data. The data might undergo incidental distortion and malicious modification. In order to verify the authenticity of the received content, the client can compute a hash value from the received data, and compare it with the hash value extracted from the data. The advantage is that no extra communication is required – the original hash value is always available and synchronized. However, on the other hand, image quality can be degraded due to watermark embedding. There is interesting interaction between hashing and watermarking. We investigate this issue by proposing a content authentication system. The hash algorithm and the watermarking algorithm are designed to have minimal interference. This is achieved by doing hashing and watermarking in different wavelet subbands. Through extensive experiments we show that the parameters of the watermarking algorithm have significant influence to the authentication performance. This work gives useful insights into the research and practice in this field.
We propose a novel image authentication system by combining perceptual hashing and robust watermarking. An image is divided into blocks. Each block is represented by a compact hash value. The hash value is embedded in the block. The authenticity of the image can be verified by re-computing hash values and comparing them with the ones extracted from the image. The system can tolerate a wide range of incidental distortion, and locate tampered areas as small as 1/64 of an image. In order to have minimal interference, we design both the hash and the watermark algorithms in the wavelet domain. The hash is formed by the sign bits of wavelet coefficients. The lattice-based QIM watermarking algorithm ensures a high payload while maintaining the image quality. Extensive experiments confirm the good performance of the proposal, and show that our proposal significantly outperforms a state-of-the-art algorithm.
This paper presents a robust and secure image hash algorithm. The algorithm extracts robust image features in the Radon transform domain. A randomization mechanism is designed to achieve good discrimination and security. The hash value is dependent on a secret key. We evaluate the performance of the proposed algorithm and compare the results with those of one existing Radon transform-based algorithm. We show that the proposed algorithm has good robustness against content-preserving distortion. It withstands JPEG compression, filtering, noise addition as well as moderate geometrical distortions. Additionally, we achieve improved performance in terms of discrimination, sensitivity to malicious tampering and receiver operating characteristics. We also analyze the security of the proposed algorithm using differential entropy and confusion/diffusion capabilities. Simulation shows that the proposed algorithm well satisfies these metrics.
Perceptual hashing is conventionally used for content identification and authentication. In this work, we explore a new application of image hashing techniques. By comparing the hash values of original images and their compressed versions, we are able to estimate the distortion level. A particular image hash algorithm is proposed for this application. The distortion level is measured by the signal to noise ratio (SNR). It is estimated from the bit error rate (BER) of hash values. The estimation performance is evaluated by experiments. The JPEG, JPEG2000 compression, and additive white Gaussian noise are considered.We show that a theoretical model does not work well in practice. In order to improve estimation accuracy, we introduce a correction term in the theoretical model. We find that the correction term is highly correlated to the BER and the uncorrected SNR. Therefore it can be predicted using a linear model. A new estimation procedure is defined accordingly. New experiment results are much improved.
Perceptual hashing is a promising solution to image content authentication. However, conventional image hash algorithms only offer a limited authentication level for the protection of overall content. In this work, we propose an image hash algorithm with block level content protection. It extracts features from DFT coefficients of image blocks. Experiments show that the hash has strong robustness against JPEG compression, scaling, additive white Gaussian noise, and Gaussian smoothing. The hash value is compact, and highly dependent on a key. It has very efficient trade-offs between the false positive rate and the true positive rate.
In this paper, a watermarking scheme for 3D stereo images is presented. The target application is 3D Digital Cinema. The watermarking is based on a dependent stereo image coding scheme, where the watermark is embedded in the JPEG2000 decoding pipeline after the inverse quantization and prior to the inverse discrete wavelet transform (IDWT). A perceptual mask is designed to cope with the particularity of the 3D perception and is tuned by the disparity map and the wavelet properties related to the Human Visual System (HVS). In this scheme, the watermark is inherited in the target image during the decoding process and hence, both the reference and the target image in the stereo pair are watermarked. Our results show improvements due to the integration of the 3D visual mask in terms of Structural Similarity Metric (SSM) and Bit Error Rate (BER).
Perceptual hashing is a technique for content identification and authentication. In this work, a frame hash based video hash construction framework is proposed. This approach reduces a video hash design to an image hash design, so that the performance of the video hash can be estimated without heavy simulation. Target performance can be achieved by tuning the construction parameters. A frame hash algorithm and two performance metrics are proposed.
Perceptual hashing is an emerging solution for identification and authentication of multimedia content. In this work, a video hash algorithm is proposed. This algorithm computes a 180-bit hash value for videos of arbitrary lengths. The hash value can resist common signal processing and slight geometric distortion. The basic mechanism of the algorithm is to compute and accumulate frame hash values. A frame hash algorithm is designed by combining semi-global and local features. Semi-global features are extracted by computing several statistics from image blocks. Local features are extracted by computing a compact edge density map around stable feature points. The good performance of the new algorithm has been demonstrated by experiments.
Perceptual hashing is a solution for identification and authentication of multimedia content. The key of this technique is the extraction of proper features. In this paper, two features are proposed for natural image hashing. They are based on the description of shapes, in terms of contours and regions. The contour-based feature is formed by edge detection. The region-based feature is formed by the angular radial transform. Simulation results show that they have good robustness and discriminability. Compared to some other features, better ROC performance is achieved.
Buyer-seller watermarking protocols integrate multimedia watermarking and fingerprinting with cryptography, for copyright protection, piracy tracing, and privacy protection. We propose an efficient buyer-seller watermarking protocol based on dynamic group signatures and additive homomorphism, to provide all the required security properties, namely traceability, anonymity, unlinkability, dispute resolution, non-framing, and non-repudiation. Another distinct feature is the improvement of the protocol's utility, such that the double watermark insertion mechanism is avoided; the final quality of the distributed content is improved; the communication expansion ratio and computation complexity are reduced, comparing with conventional schemes.
Perceptual hashing is an emerging solution for multimedia content authentication. Due to their robustness, such techniques might not work well when malicious attack is perceptually insignificant. We designed an experiment and verified that some state-of-the-art image hash algorithms could not distinguish small malicious distortion and some authentic distortion. We proposed an enhancement framework as a remedy. It suggests extracting information from the content and combining it with the secret key to generate the perceptual hash, so that perceptually insignificant information can be protected.