Investigating the solar magnetic field is crucial to understand the physical processes in the solar interior as well as their effects on the interplanetary environment. We introduce a novel method to predict the evolution of the solar line-of-sight (LoS) magnetogram using image-to-image translation with Denoising Diffusion Probabilistic Models (DDPMs). Our approach combines "computer science metrics" for image quality and "physics metrics" for physical accuracy to evaluate model performance. The results indicate that DDPMs are effective in maintaining the structural integrity, the dynamic range of solar magnetic fields, the magnetic flux and other physical features such as the size of the active regions, surpassing traditional persistence models, also in flaring situation. We aim to use deep learning not only for visualisation but as an integrative and interactive tool for telescopes, enhancing our understanding of unexpected physical events like solar flares. Future studies will aim to integrate more diverse solar data to refine the accuracy and applicability of our generative model.
In this paper, we consider a problem of self-supervised learning for small-scale datasets based on contrastive loss between multiple views of the data, which demonstrates the state-of-the-art performance in classification task. Despite the reported results, such factors as the complexity of training requiring complex architectures, the needed number of views produced by data augmentation, and their impact on the classification accuracy are understudied problems. To establish the role of these factors, we consider an architecture of contrastive loss system such as SimCLR, where baseline model is replaced by geometrically invariant "hand-crafted" network ScatNet with small trainable adapter network and argue that the number of parameters of the whole system and the number of views can be considerably reduced while practically preserving the same classification accuracy. In addition, we investigate the impact of regularization strategies using pretext task learning based on an estimation of parameters of augmentation transform such as rotation and jigsaw permutation for both traditional baseline models and ScatNet based models. Finally, we demonstrate that the proposed architecture with pretext task learning regularization achieves the state-of-the-art classification performance with a smaller number of trainable parameters and with reduced number of views.
We propose a scheme for multi-layer representation of images. The problem is first treated from an information-theoretic viewpoint where we analyze the behavior of different sources of information under a multi-layer data compression framework and compare it with a single-stage (shallow) structure. We then consider the image data as the source of information and link the proposed representation scheme to the problem of multi-layer dictionary learning for visual data. For the current work we focus on the problem of image compression for a special class of images where we report a considerable performance boost in terms of PSNR at high compression ratios in comparison with the JPEG2000 codec.
Micro-structures provide unique identifiers for physical objects, and since they can be optically detected with a mobile device, they can be applied in fields from security and forensics to anti-counterfeiting.
We address the content-identification problem by modeling it as a multi-class classification problem. The goal is to pave the way and establish a general framework to incorporate the powerful algorithms of the machine learning literature in learning from data into this problem. Through this end, a particular successful approach, linked with the coding theory known as ECOC is considered and studied. We argue that the conventional codings used in this approach are suboptimal by analyzing the problem from an information-theoretic viewpoint. We then advise the use of our recently proposed method for this problem. The ECOC approach converts the multi-class problem to several binary problems. We consider these equivalent binary classification tasks in more details and use the Gaussian Mixture Models instead of SVM’s. This latter brings significant reduction in complexity by having an assumption on the distributions.
In this work, we address the problem of content identification. We consider content identification as a special case of multiclass classification. The conventional approach towards identification is based on content fingerprinting where a short binary content description known as a fingerprint is extracted from the content. We propose an alternative solution based on elements of machine learning theory and digital communications. Similar to binary content fingerprinting, binary content representation is generated based on a set of trained binary classifiers. We consider several training/encoding strategies and demonstrate that the proposed system can achieve the upper theoretical performance limits of content identification. The experimental results were carried out both on a synthetic dataset with different parameters and the FAMOS dataset of microstructures from consumer packages.
Content identification based on digital fingerprinting attracts a lot of attention in different emerging applications. In this paper, we consider digital identification based on the sign-magnitude decomposition of fingerprint codewords and analyze the achievable rates for each component. We introduce a channel splitting approach and reveal certain interesting phenomena related to channel polarization. It is demonstrated that under certain conditions almost all rate in the sign channel is concentrated in reliable components, this can be of interest for complexity and security in various content identification applications. The envisioned extensions cover applications where the input and output alphabets of the channel are different at the encoding and decoding stages. Additionally, the reduction of the input data dimensionality at the encoding/enrollment stage can increase the cryptographic protection in terms of privacy leakage and simplify the decoding algorithms in biometric applications.
In this paper we analyze the problem of object identification in channels with desynchronization. In our analysis we assume that the identification system is designed using a pilot-based re-synchronization mechanism that assists desynchronization compensation with a certain accuracy. We demonstrate how the accuracy of re-synchronization impacts the information-theoretic limits of identification system performance.
In this paper, we consider an information-theoretic formulation of the content identification under search complexity constrain. The proposed framework is based on soft fingerprinting, i.e., joint consideration of sign and magnitude of fingerprint coefficients. The fingerprint magnitude is analyzed in the scope of communications with side information that results in channel decomposition, where all bits of fingerprints are classified to be communicated via several channels with distinctive characteristics. We demonstrate that under certain conditions the channels with low identification capacity can be neglected without considerable rate loss. This is a basis for the analysis of fast identification techniques trading-off theoretical performance in terms of achievable rate and search complexity.
We study the problem of multiple hypothesis testing (HT) in view of a rejection option. That model of HT has many different applications. Errors in testing of M hypotheses regarding the source distribution with an option of rejecting all those hypotheses are considered. The source is discrete and arbitrarily varying (AVS). The tradeoffs among error probability exponents/reliabilities associated with false acceptance of rejection decision and false rejection of true distribution are investigated, the optimal decision strategies are outlined. The special case of discrete memoryless source (DMS) is also discussed. An interesting insight that the analysis implies is the phenomenon (comprehensible in terms of supervised/unsupervised learning) that in optimal discrimination within M hypothetical distributions one permits always lower error than in deciding to decline the set of hypotheses. Geometric interpretations of the optimal decision schemes and bounds in multi-HT for AVS's are given.
In this paper we consider the problem of reversible information hiding in the case when the attacker uses only discrete memoryless channels (DMC), the decoder knows only the class of channels, but not the DMC chosen by the attacker, the attacker knows the information-hiding strategy, probability distributions of all random variables, but not the side information. We introduce the notion of reversible information hiding Incapacity, which expresses the dependence of the information hiding rate on the error probability exponent E and the distortion levels for the information hider, for the attacker and for the host data approximation. The random coding bound for reversible information hiding B-capacity is found. We obtain the lower bound for reversibility information hiding capacity for E rarr 0. In particular, we have analyzed two special cases of the general problem formulation, pure reversibility and pure message communications.