Neural compression is currently dominated by Nonlinear Transform Coding (NTC), which maps data to real-valued latents via continuous transforms. Despite its success, NTC suffers from train-test mismatch due to non-differentiable quantization, a “smoothness bias" inherent in continuous transforms that precludes optimality for certain sources, and a loss of “shaping gain" due to the complexity of including high-dimensional vector quantization. We propose SoftBinary Coding (SBC), an end-to-end learning paradigm that bypasses these limitations by using a stochastic binary latent space. In the spirit of vector quantization, SBC employs discrete representations and compresses them through a novel fast binary channel simulation scheme, for which we provide a proof of rate optimality. Experimental gains on information-theoretic sources provide both theoretical and practical closure to NTC's limitations, establishing discrete binary structures as a viable path toward reaching optimal rate–distortion bounds. Surprisingly, SBC also achieves state-of-the-art performance on vector quantization of i.i.d. sources, exceeding Trellis Coded Quantization of the Gaussian source.
Through refined asymptotic analysis based on the normal approximation, we study how higher-order coding performance depends on the mean power as well as on finer statistics of the input power. We introduce a multifaceted power model in which the expectation of an arbitrary (but finite) number of arbitrary functions of the normalized average power is constrained. The framework generalizes existing models, recovering the standard maximal and expected power constraints and the recent mean and variance constraint as special cases. Under certain growth and continuity assumptions on the functions, our main theorem gives an exact characterization of the minimum average error probability for Gaussian channels as a function of the first- and second-order coding rates. The converse proof reduces the code design problem to minimization over a compact (under the Prokhorov metric) set of probability distributions, characterizes the extreme points of this set and invokes the Bauer’s maximization principle. Our results for the multifaceted power model serve as more precise benchmarks for practical modulation schemes with multiple amplitude levels, probabilistic shaping and nonuniform constellation geometries.
Suppose X,Y are independent random variables with values in a compact abelian group (G,+). We examine the following two entropy power-type inequalities: h(X+Y)≥1/2h(X)+1/2h(Y) and h(X+Y)≥max{h(X),h(Y)}, where the entropy h(Z) of a G-valued random variable Z is defined in terms of its density with respect to Haar measure on G. For groups that are either connected or finite with no nontrivial subgroups, we precisely characterize the cases of equality and establish explicit, quantitative stability estimates in terms of relative entropy for these two inequalities. The main tools are a generalization of an entropic inequality obtained by Green, Manners and Tao (2023) for discrete entropy, and a harmonic-analytic estimate for the chi-squared contraction coefficient in connected compact groups. As an application, we derive exponential convergence rates to the uniform distribution in relative entropy for random walks on connected compact abelian groups.
Neural compression is currently dominated by Nonlinear Transform Coding (NTC), which maps data to real-valued latents via continuous transforms. Despite its success, NTC suffers from train-test mismatch due to non-differentiable quantization, a ''smoothness bias'' inherent in continuous transforms that precludes optimality for certain sources, and a loss of ''shaping gain" due to the complexity of including high-dimensional vector quantization. We propose (SBC), an end-to-end learning paradigm that bypasses these limitations by using a stochastic binary latent space. In the spirit of vector quantization, SBC employs discrete representations and compresses them through a novel fast binary channel simulation scheme, for which we provide a proof of rate optimality. Experimental gains on information-theoretic sources provide both theoretical and practical closure to NTC's limitations, establishing discrete binary structures as a viable path toward reaching optimal rate--distortion bounds. Surprisingly, SBC also achieves state-of-the-art performance on vector quantization of i.i.d. sources, exceeding Trellis Coded Quantization of the Gaussian source.
In image compression, with recent advances in generative modeling, existence of a trade-off between rate and perceptual quality has been brought to light, where perceptual quality is measured by the closeness of the output and source distributions. We consider the compression of a memoryless source sequence Xn = (X1, . . . ,Xn) in the presence of memoryless side information Zn = (Z1, . . . ,Zn), originally studied by Wyner and Ziv, but elucidate the impact of a strong perfect realism constraint, which requires the joint distribution of output symbols Yn = (Y1, ..., Yn) to match the distribution of the source sequence. We consider two cases: when Zn is available only at the decoder, or at both the encoder and decoder, and characterize the information theoretic limits in both scenarios. Previous works show the superiority of randomized codes under strong perceptual quality constraints. When Zn is available at both terminals, we characterize its dual role, as a source of common randomness, and as a second look on the source for the receiver. We also study different notions of strong perfect realism, which we call marginal realism, joint realism and near-perfect realism. We derive explicit solutions when X and Z are jointly Gaussian under the squared error distortion measure. In traditional lossy compression, having Z only at the decoder imposes no rate penalty in the Gaussian scenario. We show that, when strong perfect realism constraints are imposed this holds only when sufficient common randomness is available.
For variable-length coding with an almost-sure distortion constraint, Zhang et al. show that for discrete sources the redundancy is upper bounded by log n/n and lower bounded (in most cases) by log n/(2n), ignoring lower order terms. For a uniform source with a distortion measure satisfying certain symmetry conditions, we show that log n/(2n) is achievable and that this cannot be improved even if one relaxes the distortion constraint to be in expectation rather than with probability one.
Realism constraints (or constraints on perceptual quality) have received considerable recent attention within the context of lossy compression, particularly of images. Theoretical studies of lossy compression indicate that high-rate common randomness between the compressor and the decompressor is a valuable resource for achieving realism. On the other hand, the utility of significant amounts of common randomness has not been noted in practice. We offer an explanation for this discrepancy by considering a realism constraint that requires satisfying a universal critic that inspects realizations of individual compressed reconstructions, or batches thereof. We characterize the optimal rate-distortion trade-off under such a realism constraint, and show that it is asymptotically achievable without any common randomness, unless the batch size is impractically large.
We consider channel coding for Gaussian channels with the recently introduced mean and variance cost constraints. Through matching converse and achievability bounds, we characterize the optimal first- and second-order performance. The main technical contribution of this paper is an achievability scheme which uses random codewords drawn from a mixture of three uniform distributions on (n-1) -spheres of radii R-1,R-2 and R-3 , where R-i=O(root n) and |R-i-R-j|=O(1) . To analyze such a mixture distribution, we prove a lemma giving a uniform O(logn) bound, which holds with high probability, on the log ratio of the output distributions Q(i)(c) and Q(j)(cc ), where Q(i)(c) is induced by a random channel input uniformly distributed on an (n-1) -sphere of radius R-i . To facilitate the application of the usual central limit theorem, we also give a uniform O(logn) bound, which holds with high probability, on the log ratio of the output distributions Q(i)(c) and Q(i)(& lowast;) , where Q(i)(& lowast;) is induced by a random channel input with i.i.d. components.
We consider channel coding for discrete memoryless channels (DMCs) with a novel cost constraint that constrains both the mean and the variance of the cost of the codewords. We show that the maximum (asymptotically) achievable rate under the new cost formulation is equal to the capacity-cost function; in particular, the strong converse holds. We further characterize the optimal second-order coding rate of these cost-constrained codes; in particular, the optimal second-order coding rate is finite. We then show that the second-order coding performance is strictly improved with feedback using a new variation of timid/bold coding, significantly broadening the applicability of timid/bold coding schemes from unconstrained compound-dispersion channels to all cost-constrained channels. Equivalent results on the minimum average probability of error are also given.
Presents corrections to the paper, (Errata to “Channel Coding With Mean and Variance Cost Constraints”).
Channel coding for discrete memoryless channels (DMCs) with mean and variance cost constraints has been recently introduced. We show that there is an improvement in coding performance due to cost variability, both with and without feedback. We demonstrate this improvement over the traditional almost-sure (per-codeword) cost constraint that prohibits any cost variation above a fixed threshold. Our result simultaneously shows that feedback does not improve the second-order coding rate of simple-dispersion DMCs under the almost-sure cost constraint. This finding parallels similar results for unconstrained simple-dispersion DMCs, additive white Gaussian noise (AWGN) channels and parallel Gaussian channels.
Channel simulation is an alternative to quantization and entropy coding for performing lossy source coding. Recently, channel simulation has gained significant traction in both the machine learning and information theory communities, as it integrates better with machine learning-based data compression algorithms and has better rate-distortion-perception properties than quantization. As the practical importance of channel simulation increases, it is vital to understand its fundamental limitations. Recently, Sriramu and Wagner provided an almost complete characterisation of the redundancy of channel simulation algorithms. In this paper, we complete this characterisation. First, we significantly extend a result of Li and El Gamal, and show that the redundancy of any instance of a channel simulation problem is lower bounded by the channel simulation divergence. Second, we give two proofs that the asymptotic redundancy of simulating iid non-singular channels is lower-bounded by 1/2: one using a direct approach based on the asymptotic expansion of the channel simulation divergence and one using large deviations theory.
We consider the redundancy of the exact channel synthesis problem under an i.i.d. assumption. Existing results provide an upper bound on the unnormalized redundancy that is logarithmic in the block length. We show, via an improved scheme, that the logarithmic term can be halved for most channels and eliminated for all others. For full-support discrete memoryless channels, we show that this is the best possible.
We introduce a distortion measure for images, Wasserstein distortion, that simultaneously generalizes pixel-level fidelity on the one hand and realism or perceptual quality on the other. We discuss its metric properties. Pairs of images that are close under Wasserstein distortion illustrate its utility. In particular, we generate random images that have high fidelity to a reference image in one location of the image and smoothly transition to an independent realization as one moves away from this point. Wasserstein distortion represents a generalization and synthesis of prior work on texture generation, image realism and distortion, and models of the early human visual system, in the form of an optimizable metric in the mathematical sense.
We consider the design of practically-implementable schemes for the task of channel simulation. Existing methods do not scale with the number of simultaneous uses of the channel and are therefore unable to harness the amortization gains associated with simulating many uses of the channel at once. We show how techniques from the theory of error-correcting codes can be applied to achieve scalability and hence improved performance. As an exemplar, we focus on how polar codes can be used to efficiently simulate i.i.d. copies of a class of binary-output channels.
We characterize the growth of the Sibson mutual information, of any order that is at least unity, between a random variable and an increasing set of noisy, conditionally independent observations of the random variable. The Sibson mutual information increases to an order-dependent limit exponentially fast, with an exponent that is order-independent. The result is contrasted with composition theorems in differential privacy.
In image compression, with recent advances in generative modeling, the existence of a trade-off between the rate and the perceptual quality (realism) has been brought to light, where the realism is measured by the closeness of the output distribution to the source. It has been shown that randomized codes can be strictly better under a number of formulations. In particular, the role of common randomness has been well studied. We elucidate the role of private randomness in the compression of a memoryless source $X^n=(X_1,...,X_n)$ under two kinds of realism constraints. The near-perfect realism constraint requires the joint distribution of output symbols $(Y_1,...,Y_n)$ to be arbitrarily close the distribution of the source in total variation distance (TVD). The per-symbol near-perfect realism constraint requires that the TVD between the distribution of output symbol $Y_t$ and the source distribution be arbitrarily small, uniformly in the index $t.$ We characterize the corresponding asymptotic rate-distortion trade-off and show that encoder private randomness is not useful if the compression rate is lower than the entropy of the source, however limited the resources in terms of common randomness and decoder private randomness may be.
Wasserstein distortion is a one-parameter family of distortion measures that was recently proposed to unify fidelity and realism constraints. After establishing continuity results for Wasserstein distortion in the extreme cases of pure fidelity and pure realism, we prove the first coding theorems for compression under Wasserstein distortion focusing on the regime in which both the rate and the distortion are small.
We show the existence of variable-rate rate-distortion codes that meet the distortion constraint almost surely and are minimax, i.e., strongly, universal with respect to an unknown source distribution and a distortion measure that is revealed only to the encoder and only at runtime. If we only require minimax universality with respect to the source distribution and not the distortion measure, then we provide an achievable $\tilde {O}(1/\sqrt {n})$ redundancy rate, which we show is optimal. This is in contrast to prior work on universal lossy compression, which provides $O(\log n/n)$ redundancy guarantees for weakly universal codes under various regularity conditions. We show that either eliminating the regularity conditions or upgrading to strong universality while keeping these regularity conditions entails an inevitable increase in the redundancy to $\tilde {O}(1/\sqrt {n})$ . Our construction involves random coding with non-i.i.d. codewords and a zero-rate uncoded transmission scheme. The proof uses exact asymptotics from large deviations, acceptance-rejection sampling, and the VC dimension of distortion measures.
We show that a variation of the timid/bold coding scheme for feedback communication can be used to improve the second-order coding rate for most discrete memoryless channels with a cost constraint, even in the simple-dispersion case for which the optimal input distribution is unique.