A new approach to the problem of error correction in communication channels is proposed, in which the input sequence is transformed in such a way that the interdependence of symbols is significantly increased. Then, after the sequence is transmitted over the channel, this property is used for error correction so that the remaining error rate is significantly reduced. The complexity of encoding and decoding is linear.
We investigate the impact of (possible) deviations of the probability distribution of key values from a uniform distribution for the information-theoretic strong, or perfect, message authentication code. We found a simple expression for the decrease in security as a function of the statistical distance between the real key probability distribution and the uniform one. In a sense, a perfect message authentication code is robust to small deviations from a uniform key distribution.
Entropy codes1, i.e. methods of lossless data compression, are widely used in various systems for transmitting and storing audio and video data and for many other information systems. However, despite numerous works, the choice of an entropy code for a specific task requires extensive theoretical and experimental research. The purpose of this paper is to analyze known entropy codes, allowing to simplify the choice of a method for a specific data compression system, paying special attention to the transmission of video and audio signals. We show that letter-wise entropy codes can be useful only for small blocks of letters.So we focus our analysis on stream codes, namely, arithmetic coding, range coding and asymmetric numeral systems coding. The Huffman code is used as a point to comparison.The carried out analysis allows us to develop recommendations that significantly simplify the practical application of entropy codes and their subsequent assessment. We start with binary codes which are widely used in audio and video compression and then show thatlarge alphabet codes can provide a noticeable increase in coding speed, at least for the case of small source memory.
In this paper, we apply an information-theoretic method proposed by Ryabko and Savina (therefore called the RS-method), based on the use of data compression, to recognize the individual author’s style of a writer across four languages from different language groups and families. In this paper, the presented method was used to study fiction texts in Russian (East Slavic group of languages of the Indo-European language family), Amharic (South Ethiosemitic group of the Semitic language family), Chinese (Sinitic group of the Sino-Tibetan language family) and English (West Germanic language group of the Indo-European language family). It was found that the amount of data necessary for recognizing an author’s style is almost the same for all four languages, i.e., the amount of data is invariant across different language groups. The results obtained are of interest to computer science, literary studies, linguistics and, in particular, computational linguistics.
Nowadays there are several classes of constrained codes intended for different applications. The following two large classes can be distinguished. The first class contains codes with local constraints; for example, the source data must be encoded by binary sequences containing no sub-words 00 and 111. The second class contains codes with global constraints; for example, the code-words must be binary sequences of certain even length where half of the symbols are zeros and half are ones. It is important to note that often the necessary codes must fulfill some requirements of both classes. In this paper we propose a general polynomial complexity method for constructing codes for both classes, as well as for combinations thereof. The proposed method uses the Cover enumerative code, but calculates all the parameters on the fly with polynomial complexity, unlike the known applications of that code which employ combinatorial formulae. The main idea of the paper is to use dynamic programming to perform calculations like: how many sequences with a given prefix and a given suffix length satisfying constraints exist. For the constraints under consideration, we do not need to know the entire prefix, but much less knowledge about the prefix is sufficient. That is, we only need a brief description of the prefix.
The modern era of globalisation has led to the expansion of scientific and cultural ties between countries and has aroused interest in the problem of quality translation. For everyday use, there are now many translation programs, for example, Yandex Translate or Google Translate. When watching movies, people also most often choose a dubbed translation. Moreover, in messengers, messages can be automatically translated into the desired language to overcome the language barrier. The transition of translation to the digital dimension affects the quality of the translated text and its perception by a person. Therefore, in the first decades of the current century, the problem of quantifying translation from one language to another becomes very relevant and attractive to researchers. Our method is based on A. Kolmogorov's concept of the sources of textual variability: 1) information content, 2) form, 3) unconscious author's style. Obviously, the task of preserving the first two sources when translating a text is quite simple for professionals. Any bilingual reader can easily verify that. However, the third source - the unconscious “author's” style - determines the quality of the translation, because the translator's influence on the author's style. In this regard, literary critics need accurate non-subjective methods to determine how adequately the author's style is reproduced. In this work, we use a method based on data compression, which makes it possible to quantify the quality of translation of texts by different translators. The study was carried out on the texts of translations of classical English-language works of fiction into Russian and, conversely, Russian fiction into English. Our method is intended for comparative analysis of different translations of the same text. As a result of the study, we received an unexpected discovery. Translations from English into Russian preserve the style of the author much better than translations from Russian into English. We think this result is an additional demonstration of the power of information theory.
The problem of testing random number generators is considered and a new method for comparing the power of different statistical tests is proposed. It is based on the definitions of random sequence developed in the framework of algorithmic information theory and allows comparing the power of different tests in some cases when the available methods of mathematical statistics do not distinguish between tests. In particular, it is shown that tests based on data compression methods using dictionaries should be included in test batteries.
We describe and explore so-called linear hash functions and show how they can be used to build error detection and correction codes. The method can be applied for different types of errors (for example, burst errors). When the method is applied to a model where number of distorted letters is limited, the obtained estimate of its performance is slightly better than the known Varshamov-Gilbert bound. We also describe random code whose performance is close to the same boundary, but its construction is much simpler. In some cases the obtained methods are simpler and more flexible than the known ones. In particular, the complexity of the obtained error detection code and the well-known CRC code is close, but the proposed code, unlike CRC, can detect with certainty errors whose number does not exceed a predetermined limit.
Perfect ciphers have been a very attractive cryptographic tool ever since C. Shannon described them. Note that, by definition, if a perfect cipher is used, no one can get any information about the encrypted message without knowing the secret key. We consider the problem of reducing the key length of perfect ciphers, because in many applications the length of the secret key is a crucial parameter. This paper describes a simple method of key length reduction. This method gives a perfect cipher and is based on the use of data compression and randomisation, and the average key length can be made close to Shannon entropy (which is the key length limit). It should be noted that the method can effectively use readily available data compressors (archivers).
We consider the problem of constructing an unconditionally secure cipher with a short key for the case where the probability distribution of encrypted messages is unknown. Note that unconditional security means that an adversary with no computational constraints can only obtain a negligible amount of information ("leakage") about an encrypted message (without knowing the key). Here, we consider the case of a priori (partially) unknown message source statistics. More specifically, the message source probability distribution belongs to a given family of distributions. We propose an unconditionally secure cipher for this case. As an example, one can consider constructing a single cipher for texts written in any of the languages of the European Union. That is, the message to be encrypted could be written in any of these languages.
We consider the problem of constructing an unconditionally secure cipher for the case when the key length is less than the length of the encrypted message. (Unconditional security means that a computationally unbounded adversary cannot obtain information about the encrypted message without the key). In this article, we propose a cipher based on data compression and randomisation in combination with entropically-secure encryption and apply it to the following two cases: (i) the statistics of encrypted messages are known; and (ii) statistics are unknown, but messages are generated by a Markov chain with known memory (or connectivity). In both cases, the length of the secret key is negligible compared to the length of the message.
In recent years, the task of translating from one language to another has attracted wide attention from researchers due to numerous practical uses, ranging from the translation of various texts and speeches, including the so-called "machine" translation, to the dubbing of films and numerous other video materials. To study this problem, we propose to use the information-theoretic method for assessing the quality of translations. We based our approach on the classification of sources of text variability proposed by A.N. Kolmogorov: information content, form, and unconscious author's style. It is clear that the unconscious "author's" style is influenced by the translator. So researchers need special methods to determine how accurately the author's style is conveyed, because it, in a sense, determines the quality of the translation. In this paper, we propose a method that allows us to estimate the quality of translation from different translators. The method is used to study translations of classical English-language works into Russian and, conversely, Russian classics into English. We successfully used this method to determine the attribution of literary texts.
We consider the problem of constructing an unconditionally secure cipher for the case when the key length is less than the length of the encrypted message. (Unconditional security means that a computationally unbounded adversary cannot obtain information about the encrypted message without the key.) In this article, we propose data compression and randomisation techniques combined with entropically-secure encryption for the case when message statistics are known. The resulting cipher can be used for encryption in such a way that the key length does not depend on the entropy or the length of the encrypted message; instead, it is determined by the required security level.
The problem of constructing effective statistical tests for random number generators (RNG) is considered. Currently, statistical tests for RNGs are a mandatory part of cryptographic information protection systems, but their effectiveness is mainly estimated based on experiments with various RNGs. We find an asymptotic estimate for the p-value of an optimal test in the case where the alternative hypothesis is a known stationary ergodic source, and then describe a family of tests each of which has the same asymptotic estimate of the p-value for any (unknown) stationary ergodic source.
In 2002, Russell and Wang proposed a definition of entropically security that was developed within the framework of secret key cryptography. An entropically-secure system is unconditionally secure, that is, unbreakable, regardless of the enemy's computing power. In 2004, Dodis and Smith developed the results of Russell and Wang and, in particular, stated that the concept of an entropy-protected symmetric encryption scheme is extremely important for cryptography, since it is possible to construct entropy-protected symmetric encryption schemes with keys much shorter than the keys. the length of the input data, which allows you to bypass the famous lower bound on the length of the Shannon key. In this report, we propose an entropy-protected scheme for the case where the encrypted message is generated by a Markov chain with unknown statistics. The length of the required secret key is proportional to the logarithm of the length of the message (as opposed to the length of the message itself for the one-time pad).
We consider the issue of direct access to any letter of a sequence encoded with a variable length code and stored in the computer's memory, which is a special case of the random access problem to compressed memory. The characteristics according to which methods are evaluated are the access time to one letter and the memory used. The proposed methods, with various trade-offs between the characteristics, outperform the known ones.
Pseudo-random number generators (PRNGs) are widely used in computer simulation, cryptography, and many other fields. In this paper, we describe a PRNG class, which, firstly, has been successfully tested using the most powerful modern test batteries, and secondly, is proved to consist of generators that generate normal sequences. The latter property means that, for any generated sequence [Formula: see text] and any binary word [Formula: see text], we have [Formula: see text] where [Formula: see text] is the number of occurrences of [Formula: see text] in the sequence [Formula: see text], [Formula: see text].
Currently, statistical tests for random number generators (RNGs) are widely used in practice, and some of them are even included in information security standards. But despite the popularity of RNGs, consistent tests are known only for stationary ergodic deviations of randomness (a test is consistent if it detects any deviations from a given class when the sample size goes to infinity). However, the model of a stationary ergodic source is too narrow for some RNGs, in particular, for generators based on physical effects. In this article, we propose computable consistent tests for some classes of deviations more general than stationary ergodic and describe some general properties of statistical tests. The proposed approach and the resulting test are based on the ideas and methods of information theory.
We propose an information-theoretical method of attribution of literary texts, developed within the framework of information theory and mathematical statistics. Using the proposed method, the following two problems of disputed authorship in Russian and Soviet literature were investigated: i) the problem of false attribution of some novels to Nekrasov and ii) the problem of dubious attribution of two novels to Bulgakov. The research has shown the high efficiency of the data-compression method for attribution of literary texts.
Currently, statistical tests for random number generators (RNGs) are widely used in practice, and some of them are even included in information security standards. But despite the popularity of RNGs, consistent tests are known only for stationary ergodic deviations of randomness (a test is consistent if it detects any deviations from a given class when the sample size goes to ∞). However, the model of a stationary ergodic source is too narrow for some RNGs, in particular, for generators based on physical effects. In this article, we propose computable consistent tests for some classes of deviations more general than stationary ergodic and describe some general properties of statistical tests. The proposed approach and the resulting test are based on the ideas and methods of information theory.