
Water pollution from organic dyes, antibiotics, and other chemicals poses a significant threat to water quality and availability.
We investigate numerically the blocking of two-dimensional bistable reaction diffusion fronts by geometric obstacles. Our goal is to derive quantitative criteria for front propagation in the presence of spatial heterogeneities. Using a conservation law approach, we show that the integral of the reaction term acts as an effective driving force for the front. Combining this insight with the exact one-dimensional traveling wave solution, we construct a reduced analytical model that predicts blocking thresholds. In particular, we obtain explicit conditions for front propagation in a waveguide connected to a conical region of angle theta, valid for widths w less than 4. The model captures the influence of both geometry and nonlinearity, and shows good agreement with numerical simulations. Finally, we extend the analysis to more complex geometries, including checkerboard-like obstacles, and derive simple heuristic rules governing front propagation.
Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks. Behind their ease of use lies a complex process: words become tokens, tokens become vectors, and vectors ultimately give rise to streams of bits that flow through High-Performance Computing (HPC) systems. As modern LLMs grow to billions or trillions of parameters, this path increasingly unfolds across thousands of interconnected accelerators, making the underlying communication fabric a critical and often opaque component of model training. This tutorial aims to walk the reader through the journey from words to network traffic, shedding light on how language is translated into communication flows within HPC training systems. Using concrete examples from Dante's Divine Comedy, we illustrate how model architecture, tokenization, embeddings, and parallelization strategies shape the volume, structure, and timing of data exchanged across the network. We combine architectural analysis with analytical traffic models and numerical examples to characterize the communication requirements of LLM training. We try to demystify how words travel across the network and provide practical insights into the network requirements needed to support the journey from text to trained model.
Extending a recent effective theory formulation for the dynamics of kinks in the sine-Gordon model [T. Dobrowolski et al., Phys. Rev. E 111, 024203 (2025)2470-004510.1103/PhysRevE.111.024203], we propose an analogous effective description of ϕ^{4} kinks. Three different reduced models based on the kink position, width and internal mode amplitude are introduced and compared systematically with the numerical solution of the equation with space- and time-dependent perturbations. In all cases considered, the model based on the kink position and width agrees the best with the full numerical solution. As long as the external driving frequency of the perturbation remains moderate, it captures with remarkable accuracy the intricate dynamical processes taking place in the system.
Knowledge Distillation (KD) and mixup have proven effective at inducing smoothness in class boundaries; KD captures inherent class relationships in probability distributions, and mixup enforces them through convex combinations of inputs. Their interaction, however, remains poorly understood, particularly when mixup is applied only during student training. In this setting, the teacher is queried on inputs drawn from a vicinal distribution it never saw during training, a controlled mismatch whose effect on knowledge transfer has not been characterised. We show that this mismatch causes the teacher's supervisory signal to be dominated by distributional confusion rather than inter-class structure. Despite it, the student does not merely imitate the teacher: it independently acquires greater linearity in the vicinal region, a structural property that the teacher lacks, and goes beyond dark-knowledge transfer. KD with mixup consistently improves student accuracy and reduces overconfidence by an order of magnitude relative to the baseline, across CIFAR and ImageNet with varying-capacity teachers. Crucially, calibration propagates from teacher to student independently of accuracy transfer, and temperature scaling governs a measurable accuracy-calibration trade-off that becomes more pronounced under vicinal training. These results reframe mixup distillation not as a degraded version of standard KD, but as a richer transfer channel that simultaneously shapes discriminative performance, uncertainty estimation, and representational geometry.