
Online learning is a foundational paradigm underlying applications from recommendation systems to the continual learning of modern AI models. Yet much of its theory centers on either fully adversarial or purely stochastic settings. However, real-world environments typically fall between these extremes, making classical models inadequate for describing practical behavior. This monograph develops a unified perspective for analyzing online learning under more nuanced and realistic environments. The authors approach the problem through the lens of universality from information theory and extend tools such as the Shtarkov sum, covering numbers and packing arguments to the online setting, revealing deeper structural connections between these two fields. Building on this viewpoint, they characterize minimax regret for logarithmic and Lipschitz losses, analyze expected regret under i.i.d. and more general stochastic processes and study hybrid adversarial–stochastic scenarios. The authors further develop constructive algorithms that achieve near-optimal regret guarantees, yielding a coherent and fine-grained information-theoretic framework of online universal learning.
Directed information (DI) is an information measure that attempts to capture directionality in the flow of information from one random process to another. The definition of DI, in its present form, is due to Massey (1990), but the origins can be traced back to the work of Marko in the 1960s. It is closely related to other causal influence measures, such as transfer entropy and Granger causality. This monograph provides an overview of DI and its main application in information theory, namely, characterizing the capacity of channels with feedback and memory. The authors begin by reviewing the definitions and basic properties of DI, particularly its relation to mutual information, the classic information measure introduced by Shannon. The authors then explain how DI is related to transfer entropy, Granger causality and Pearl’s causality. DI is often used to identify causal relationships in natural processes, such as neural spike trains. Several methods have been developed in the literature to estimate the DI between two processes from time series data. The authors provide a survey, ranging from classic plug-in estimators to modern neural-network-based estimators. Considering the information-theoretic application of channel capacity estimation, the authors describe how such estimators numerically optimize DI rate over a class of joint distributions on input and output processes. A significant part of the monograph is devoted to techniques to compute the feedback capacity of finite-state channels (FSCs). The feedback capacity of a strongly connected FSC is given by a multi-letter expression involving the maximization of the DI rate from the channel input process to the output process, the maximization being performed over the class of causal conditioned probability distributions on the input process. When the FSC is also unifilar, i.e., the next state is given by a time-invariant function of the current state and the new input-output symbol pair, the feedback capacity is the optimal average reward of an appropriately formulated Markov decision process (MDP). This MDP formulation has been exploited to develop several methods to compute exactly, or at least estimate closely, the feedback capacity of a unifilar FSC. This monograph describes these methods, starting from the classic value iteration algorithm, moving on to Q-graph methods, and ending with reinforcement learning algorithms that can handle channels with large input and output alphabets.
Current wireless networks are designed to optimize spectral efficiency for human users, who typically require sustained connections for high-data-rate applications like file transfers and video streaming. However, these networks are increasingly inadequate for the emerging era of machine-type communications (MTC). With a vast number of devices exhibiting sporadic traffic patterns consisting of short packets, the grant-based multiple access procedures utilized by existing networks lead to significant delays and inefficiencies. To address this issue the unsourced random access (URA) paradigm has been proposed. This paradigm assumes the devices to share a common encoder thus simplifying the reception process by eliminating the identification procedure. The URA paradigm not only addresses the computational challenges but it also considers the random access (RA) as a coding problem, i.e., takes into account both medium access protocols and physical layer effects. In this monograph we provide a comprehensive overview of the URA problem in noisy channels, with the main task being to explain the major ideas rather than to list all existing solutions.
Twenty Questions originated as a parlor game between two players. The game starts from a player named an oracle, who privately thinks of a secret. The other player, called the questioner, tries to guess the secret by querying the oracle with at most twenty questions having Yes/No answers. Early versions of the game can be traced to ancient Greece and ancient Rome. Motivated by the Hungarian version of this game, in the middle of the twentieth century, R & eacute;nyi formulated the game as a mathematical problem of guessing an integer from a finite set, where the oracle could lie either randomly to each question or lie to a finite number of questions. The mathematical study of Twenty Questions is motivated by current applications in many domains: communications; ma-chine learning; and computer vision. The game with an oracle who is allowed a fixed number of lies was also studied by Ulam and Berlekamp and is known as the R & eacute;nyi-Ulam-Berlekamp game. In contrast, the setting where the oracle lies randomly is less understood. In this monograph, we summarize recent advances in the information theoretical analysis of Twenty Questions with random error. In particular, focusing on the practical application of sensor network target localization, we study a query-dependent channel to model oracle's noisy response behavior, such as providing a wrong answer or declining to answer a question. We concentrate on non-adaptive query procedures where all questions are designed prior to posing questions. We cover settings relevant to estimating a single target, a single moving target, and multiple targets over the unit cube of a finite dimension. We also consider adaptive querying for a single target to illustrate the benefit of adaptivity. In adaptive querying, each question is designed sequentially using responses to all previous questions. All of our theoretical results are illustrated using numerical examples. Finally, we discuss future research directions. These include geometry constraints for query sets, low-complexity query procedures, connections to group testing, and practical applications in machine learning and communications.
Power control is often used to ensure efficient resource utilization in communication systems. Its role becomes even more critical in the emerging paradigm of energy harvesting communications due to the intermittency and randomness of ambient energy sources. This monograph provides a review of the fundamental power control policies and their performance analysis in the basic setting of a discrete-time battery-limited energy harvesting communication system with independent and identically distributed energy arrivals. Three different settings, namely, offline power control, online power control, and power control with lookahead, are considered, corresponding respectively to the cases with non-causal, causal, and partial non-causal knowledge of the energy arrival process. A complete characterization of the optimal offline power control policy is presented. In the online setting, the focus is placed on the greedy policy, which is optimal in the low-battery-capacity regime, and universally near-optimal policies, which include the maximin optimal policy, the fixed fraction policy, the two-piece fixed fraction policy, and the locally fixed fraction policy. Finally, power control with lookahead is introduced to bridge offline and online power control, the entire spectrum of optimal policies is characterized for Bernoulli energy arrivals, and the extension beyond the Bernoulli case is also discussed.
This monograph offers a toolbox of mathematical techniques that have been effective and widely applicable in information- theoretic analyses. The first tool is a generalization of the method of types to Gaussian settings, and then to general exponential families. The second tool is Laplace and saddle- point integration, which allow to refine the results of the method of types, and can obtain various precise asymptotic results. The third is the type class enumeration method, a principled method to evaluate the exact random-coding exponent of coded systems, which results in the best known exponent in various problems. The fourth is a subset of tools aimed at evaluating the expectation of non-linear functions of random variables, either via integral representations, by a refinement of Jensen's inequality via change-of-measure, by complementing Jensen's inequality with a reversed inequality, or by a class of generalized Jensen's inequalities that are applicable for functions beyond convex/concave. Various examples of all these tools are provided throughout the monograph.
One-shot channel simulation (or channel synthesis) has seen increasing applications in lossy compression, differential privacy and machine learning. In this setting, an encoder observes a source X, and transmits a description to a decoder, so as to allow it to produce an output Y with a desired conditional distribution PY|X. In other words, the encoder and the decoder are simulating the noisy channel PY|X using noiseless communication. This can also be seen as a lossy compression scheme with a stronger guarantee on the joint distribution of X and Y. This monograph gives an overview of the theory and applications of the channel simulation problem. We will present a unifying review of various one-shot and asymptotic channel simulation techniques that have been proposed in different areas, namely dithered quantization, rejection sampling, minimal random coding, likelihood encoder, soft covering, Poisson functional representation, and dyadic decomposition.
The usual answer to the question "What probability distribution maximizes entropy or differential entropy of a random variable X subject to the constraint that the expected value of a real-valued function g applied to X has a specified value mu ?" is an exponential distribution (probability mass or probability density function), with g(x) in the exponent multiplied by a parameter \ , and with the parameter chosen so the exponential distribution causes the expected value of g(X) to equal mu . The latter is called moment matching. While it is well-known that, when there are multiple expected value constraints, there are functions and expected value specifications for which moment matching is not possible, it is not well-known that this can happen when there is a single expected value constraint and a single parameter. This motivates the present monograph, whose goal is to reexammine the question posed above, and to derive its answer in an accessible, self-contained and complete fashion. It also derives the maximum entropy/differential entropy when there is a constraint on the support of the probability distributions, when there is only a bound on expected value and when there is a variance constraint. Properties of the resulting maximum entropy/differential entropy as a function of mu are derived, such as its convexity and its monotonicities. Example functions are presented, including many for which moment matching is possible for all relevant values of mu, and some for which it is not. Indeed, there can be only subtle differences between the two kinds of functions. As one-parameter exponential probability distributions play a dominant role, one section of this monograph provides a self-contained discussion and derivation of their properties, such as the finiteness and continuity of the exponential normalizing constant (sometimes called the partition function) as ,\ varies, the finiteness, continuity, monotonicity and limits of the expected value of g(X) under the exponential distribution as ,\ varies, and similar issues for entropy and differential entropy. Most of these are needed in deriving the maximum entropy/differential entropy or the properties of the resulting function of mu. Aside from addressing the question posed initially, this monograph can be viewed as a warmup for discussions of maximizing entropy/differential entropy with multiple expected value constraints and of multiparameter exponential families. It also provides a small taste of information geometry.
Over the last 70 years, information theory and coding has enabled communication technologies that have had an astounding impact on our lives. This is possible due to the match between encoding/decoding strategies and corresponding channel models. Traditional studies of channels have taken one of two extremes: Shannon-theoretic models are inherently average-case in which channel noise is governed by a memoryless stochastic process, whereas coding-theoretic (referred to as “Hamming”) models take a worst-case, adversarial, view of the noise. However, for several existing and emerging communication systems the Shannon/average-case view may be too optimistic, whereas the Hamming/worstcase view may be too pessimistic. This monograph takes up the challenge of studying adversarial channel models that lie between the Shannon and Hamming extremes.
We develop an information theoretic framework for addressing feature selection in applications where the inference task is not specified in advance and the data is from a large alphabet. We introduce a natural notion of universality for such problems, and show that locally optimal solutions are straightforward to obtain, admit natural interpretations via information geometry, have computationally efficient implementations, and represent a practically useful learning methodology. Our development also reveals the key role of Hirschfeld-Gebelein-Renyi maximal correlation and the alternating conditional expectations (ACE) algorithm in such problems.
Wireless communication has traditionally been designed to connect human users. The main design goal was to maximize the data rate while guaranteeing moderate reliability and latency targets dictated by the limitations of human senses. The application of wireless connectivity for machine to machine communications, typically known as machinetype communications (MTC), has been growing in the past decade due to its flexibility, scalability and ease of use. It is also driven by the proliferation of Internet of Things (IoT) nodes and applications, with several billions of connected devices expected by the next decade. The fifth-generation (5G) New Radio (NR) wireless system has introduced two distinct services classes to support MTC, namely massive machine-type communications (mMTC) and the ultra-reliable low-latency communications (URLLC). Out of these, designing URLLC solutions is the most challenging given that it aims to provide dependable connectivity for mission-critical applications in industrial scenarios, process engineering and other similar verticals. URLLC aims to guarantee very high reliability and very low latency, and therefore the outage performance replaces the average performance as the main design criterion. This calls for a new approach to the communication- and information-theoretic fundamentals of wireless system design. Different theoretic foundations of URLLC have so far been treated in individual and disconnected works that fail to provide a meta-level understanding of this topic. This monograph aims at filling this gap by presenting a comprehensive coverage of the topic including the motivation, theory, practical enablers and future evolution. The unified level of details in this monograph is aimed at providing a balanced coverage between its fundamental communication- and information-theoretic background and its practical enablers, including 5G NR system design aspects. Finally, this monograph offers an outlook on URLLC evolution in the sixth-generation (6G) era towards dependable and resilient wireless communications.
Probabilistic amplitude shaping (PAS) proposed in B & ouml;cherer, Steiner, Schulte [24] is a practical architecture for combining non-uniform distributions on higher-order constellations with off-the-shelf forward error correction (FEC) codes. PAS consists of a distribution matcher (DM) that imposes a desired distribution on the signal point amplitudes, followed by systematic FEC encoding, preserving the amplitude distribution. FEC encoding generates additional parity bits, which select the signs of the signal points. At the receiver, FEC decoding is followed by an inverse DM. PAS quickly had a large industrial impact, in particular in fiber-optic communications. This monograph details the practical considerations that led to the invention of PAS and provides an information-theoretic assessment of the PAS architecture. Because of the separation into a shaping layer and an FEC layer, the theoretic analysis of PAS requires new tools. On the shaping layer, the cost penalty and rate loss of finite length DMs is analyzed. On the FEC layer, achievable FEC rates are derived. Using mismatched decoding, achievable rates are studied for decoding metrics of practical importance. Combining the findings, it is shown that PAS with linear codes is capacity-achieving on a class of discrete input channels. Open questions for future study are discussed.
In this monograph, we review recent advances in second-order asymptotics for lossy source coding, which provides approximations to the finite blocklength performance of optimal codes. The monograph is divided into three parts. In part I, we motivate the monograph, present basic definitions, introduce mathematical tools and illustrate the motivation of non-asymptotic and second-order asymptotics via the example of lossless source coding. In part II, we first present existing results for the rate-distortion problem with proof sketches. Subsequently, we present five generations of the rate-distortion problem to tackle various aspects of practical quantization tasks: noisy source, noisy channel, mismatched code, Gauss-Markov source and fixed-to-variable length compression. By presenting theoretical bounds for these settings, we illustrate the effect of noisy observation of the source, the influence of noisy transmission of the compressed information, the effect of using a fixed coding scheme for an arbitrary source and the roles of source memory and variable rate. In part III, we present four multiterminal generalizations of the rate-distortion problem to consider multiple encoders, decoders or source sequences: the Kaspi problem, the successive refinement problem, the Fu-Yeung problem and the Gray-Wyner problem. By presenting theoretical bounds for these multiterminal problems, we illustrate the role of side information, the optimality of stop and transmit, the effect of simultaneous lossless and lossy compression, and the tradeoff between encoders' rates in compressing correlated sources. Finally, we conclude the monograph, mention related results and discuss future directions.
Reed-Muller (RM) codes were introduced in 1954 and have long been conjectured to achieve Shannon's capacity on symmetric channels. The activity on this conjecture has recently been revived with the emergence of polar codes. RM codes and polar codes are generated by the same matrix G m = [ 1 1 1 0 ] ⊗m but using different subset of rows. RM codes 1 1 select simply rows having largest weights. Polar codes select instead rows having the largest conditional mutual information proceeding top to down in Gm; while this is a more elaborate and channel-dependent rule, the top-to-down ordering allows Arıkan to show that the conditional mutual information polarizes, and this gives directly a capacity-achieving code on any symmetric channel. RM codes are yet to be proved to have such a property, despite the recent success for the erasure channel. In this article, we connect RM codes to polarization theory. We show that proceeding in the RM code ordering, i.e., not top-to-down but from the lightest to the heaviest rows in G m, the conditional mutual information again polarizes. Here “polarization” means that almost all the conditional mutual information becomes either very close to 0 or very close to 1. Polarization itself is a necessary condition for RM codes to achieve capacity on symmetric channels while polarization together with a strong order on the conditional mutual information gives a sufficient condition, where strong order means that rows with larger weight always correspond to larger conditional mutual information. Although we are not able to prove the strong order, we establish a partial order on the conditional mutual information, which is a subset of the strong order. While the main results of this article-polarization together with the partial order-provide some advances on the capacity-achieving conjecture of RM codes, we emphasize that our results do not allow us to prove the conjecture.
This monograph reviews a class of univariate piecewise polynomial functions known as discrete splines, which share properties analogous to the better-known class of spline functions, but where continuity in derivatives is replaced by (a suitable notion of) continuity in divided differences. As it happens, discrete splines bear connections to a wide array of developments in applied mathematics and statistics, from divided differences and Newton interpolation (dating back to over 300 years ago) to trend filtering (from the last 15 years). We survey these connections, and contribute some new perspectives and new results along the way.
Due to its longevity and enormous information density, DNA is an attractive medium for archival data storage. Natural DNA more than 700.000 years old has been recovered, and about 5 grams of DNA can in principle hold a Zetabyte of digital information, orders of magnitude more than what is achieved on conventional storage media. Thanks to rapid technological advances, DNA storage is becoming practically feasible, as demonstrated by a number of experimental storage systems, making it a promising solution for our society's increasing need of data storage. While in living things, DNA molecules can consist of millions of nucleotides, due to technological constraints, in practice, data is stored on many short DNA molecules, which are preserved in a DNA pool and cannot be spatially ordered. Moreover, imperfections in sequencing, synthesis, and handling, as well as DNA decay during storage, introduce random noise into the system, making the task of reliably storing and retrieving information in DNA challenging. This unique setup raises a natural information-theoretic question: how much information can be reliably stored on and reconstructed from millions of short noisy sequences? The goal of this monograph is to address this question by discussing the fundamental limits of storing information on DNA. Motivated by current technological constraints on DNA synthesis and sequencing, we propose a probabilistic channel model that captures three key distinctive aspects of the DNA storage systems: (1) the data is written onto many short DNA molecules that are stored in an unordered fashion; (2) the molecules are corrupted by noise and (3) the data is read by randomly sampling from the DNA pool. Our goal is to investigate the impact of each of these key aspects on the capacity of the DNA storage system. Rather than focusing on coding-theoretic considerations and computationally efficient encoding and decoding, we aim to build an information-theoretic foundation for the analysis of these channels, developing tools for achievability and converse arguments.
Codes in the sum-rank metric have attracted significant attention for their applications in distributed storage systems, multishot network coding, streaming over erasure channels, and multi-antenna wireless communication. This monograph provides a tutorial introduction to the theory and applications of sum-rank metric codes over finite fields. At the heart of the monograph is the construction of linearized Reed-Solomon codes, a general construction of maximum sum-rank distance (MSRD) codes with polynomial field sizes. Linearized Reed-Solomon codes specialize to classical Reed-Solomon and Gabidulin code constructions in the Hamming and rank metrics, respectively, and they admit an efficient Welch-Berlekamp decoding algorithm. Applications of these codes in distributed storage systems, network coding, and multi-antenna communication are developed. Other families of codes in the sum-rank metric, including convolutional codes and subfield subcodes are described, and recent results in the general theory of codes in the sum-rank metric are surveyed.
We focus on some specific problems in distribution testing, taking goodness-of-fit as a running example. In particular, we do not aim to provide a comprehensive summary of all the topics in the area; but will provide self-contained proofs and derivations of the main results, trying to highlight the unifying techniques.