Deep linear networks have been extensively studied, as they provide simplified models of deep learning. However, little is known in the case of finite-width architectures with multiple outputs and convolutional layers. In this manuscript, we provide rigorous results for the statistics of functions implemented by the aforementioned class of networks, thus moving closer to a complete characterization of feature learning in the Bayesian setting. Our results include: (i) an exact and elementary non-asymptotic integral representation for the joint prior distribution over the outputs, given in terms of a mixture of Gaussians; (ii) an analytical formula for the posterior distribution in the case of squared error loss function (Gaussian likelihood); (iii) a quantitative description of the feature learning infinite-width regime, using large deviation theory. From a physical perspective, deep architectures with multiple outputs or convolutional layers represent different manifestations of kernel shape renormalization, and our work provides a dictionary that translates this physics intuition and terminology into rigorous Bayesian statistics.
Feature extraction - the ability to identify relevant properties of data - is a key factor underlying the success of deep learning. Yet, it has proved difficult to elucidate its nature within existing predictive theories, to the extent that there is no consensus on the very definition of feature learning. A promising hint in this direction comes from previous phenomenological observations of quasi-universal aspects in the training dynamics of neural networks, displayed by simple properties of feature geometry. We address this problem within a statistical-mechanics framework for Bayesian learning in one hidden layer neural networks with standard parameterization. Analytical computations in the proportional limit (when both the network width and the size of the training set are large) can quantify fingerprints of feature learning, both collective ones (related to manifold geometry) and microscopic ones (related to the weights). In particular, (i) the distance between different class manifolds in feature space is a nonmonotonic function of the temperature, which we interpret as the equilibrium counterpart of a phenomenon observed under gradient descent (GD) dynamics, and (ii) the microscopic learnable parameters in the network undergo a finite data-dependent displacement with respect to the infinite-width limit, and develop correlations. These results indicate that nontrivial feature learning is at play in a regime where the posterior predictive distribution is that of Gaussian process regression with a trivially rescaled prior.
Empirical data, on which deep learning relies, has substantial internal structure, yet prevailing theories often disregard this aspect. Recent research has led to the definition of structured data ensembles, aimed at equipping established theoretical frameworks with interpretable structural elements, a pursuit that aligns with the broader objectives of spin glass theory. We consider a one-parameter structured ensemble where data consists of correlated pairs of patterns, and a simplified model of unsupervised learning, whereby the internal representation of the training set is fixed at each layer. A mean field solution of the model identifies a set of layer-wise recurrence equations for the overlaps between the internal representations of an unseen input and of the training set. The bifurcation diagram of this discrete-time dynamics is topologically inequivalent to the unstructured one, and displays transitions between different phases, selected by varying the load (the number of training pairs divided by the width of the network). The network's ability to resolve different patterns undergoes a discontinuous transition to a phase where signal processing along the layers dissipates differential information about an input's proximity to the different patterns in a pair. A critical value of the parameter tuning the correlations separates regimes where data structure improves or hampers the identification of a given pair of patterns.
To achieve near-zero training error in a classification problem, the layers of a feed-forward network have to disentangle the manifolds of data points with different labels to facilitate the discrimination. However, excessive class separation can lead to overfitting because good generalization requires learning invariant features, which involve some level of entanglement. We report on numerical experiments showing how the optimization dynamics finds representations that balance these opposing tendencies with a non-monotonic trend. After a fast segregation phase, a slower rearrangement (conserved across datasets and architectures) increases the class entanglement. The training error at the inversion is stable under subsampling and across network initializations and optimizers, which characterizes it as a property solely of the data structure and (very weakly) of the architecture. The inversion is the manifestation of tradeoffs elicited by well-defined and maximally stable elements of the training set called 'stragglers', which are particularly influential for generalization. Feed-forward neural networks have become powerful tools in machine learning, but their behaviour during optimization is still not well understood. Ciceri and colleagues find that during optimization, class representations first separate and then rejoin, prompted by specific elements of the training set.
Despite the practical success of deep neural networks, a comprehensive theoretical framework that can predict practically relevant scores, such as the test accuracy, from knowledge of the training data is currently lacking. Huge simplifications arise in the infinite-width limit, in which the number of units Nℓ in each hidden layer (ℓ = 1, …, L, where L is the depth of the network) far exceeds the number P of training examples. This idealization, however, blatantly departs from the reality of deep learning practice. Here we use the toolset of statistical mechanics to overcome these limitations and derive an approximate partition function for fully connected deep neural architectures, which encodes information on the trained models. The computation holds in the thermodynamic limit, where both Nℓ and P are large and their ratio αℓ = P/Nℓ is finite. This advance allows us to obtain: (1) a closed formula for the generalization error associated with a regression task in a one-hidden layer network with finite α1; (2) an approximate expression of the partition function for deep architectures (via an effective action that depends on a finite number of order parameters); and (3) a link between deep neural networks in the proportional asymptotic limit and Student’s t-processes. Theoretical frameworks aiming to understand deep learning rely on a so-called infinite-width limit, in which the ratio between the width of hidden layers and the training set size goes to zero. Pacelli and colleagues go beyond this restrictive framework by computing the partition function and generalization properties of fully connected, nonlinear neural networks, both with one and with multiple hidden layers, for the practically more relevant scenario in which the above ratio is finite and arbitrary.
Compelling evidence shows that cancer persister cells represent a major limit to the long-term efficacy of targeted therapies. However, the phenotype and population dynamics of cancer persister cells remain unclear. We developed a quantitative framework to study persisters by combining experimental characterization and mathematical modeling. We found that, in colorectal cancer, a fraction of persisters slowly replicates. Clinically approved targeted therapies induce a switch to drug-tolerant persisters and a temporary 7- to 50-fold increase of their mutation rate, thus increasing the number of persister-derived resistant cells. These findings reveal that treatment may influence persistence and mutability in cancer cells and pinpoint inhibition of error-prone DNA polymerases as a strategy to restrict tumor recurrence.
S. Ariosto, 2 R. Pacelli, F. Ginelli, 2 M. Gherardi, 2 and P. Rotondo 2 Dipartimento di Scienza e Alta Tecnologia and Center for Nonlinear and Complex Systems, Università degli Studi dell’Insubria, Via Valleggio 11, 22100 Como, Italy I.N.F.N. Sezione di Milano, Via Celoria 16, 20133 Milano, Italy Dipartimento di Scienza Applicata e Tecnologia, Politecnico di Torino, 10129 Torino, Italy Università degli Studi di Milano, Via Celoria 16, 20133 Milano, Italy
When cancer cells are exposed to lethal doses of targeted therapies, the emergence of a subpopulation of drug-tolerant persister cells (DTPs) is often observed. We previously reported that colorectal cancer (CRC) cells exposed to targeted therapies activate an adaptive mutability stress response, involving DNA damage induction and a switch to low-fidelity DNA replication. Therefore, targeted treatment might lead to increased mutation rate in DTPs, but mutation rates of cancer cells during treatment have not been quantitatively assessed. Here, we combined biological experiments and mathematical modelling to characterize emergence and dynamics of DTPs. From this, we extrapolated parameters governing dynamics of cancer cells populations and designed a modified Luria-Delbrück assay on mammalian cells (MC-LD) to quantify mutations rates of CRC cells under standard growth conditions and during exposure to targeted therapy. We selected mismatch repair proficient CRC cell lines sensitive to different clinically used therapeutic agents, and derived clones to be used for experiments. By monitoring cell dynamics in drug-response growth assays, we found that CRC cells exposed to targeted therapy display a biphasic killing curve reaching a stable plateau, a pattern indicative of emergence of DTPs. By fitting model estimation to population growth assays, we predicted that, even if a subgroup of DTPs predated treatment, the majority of them emerged only upon exposure to targeted therapies. We also observed that DTPs slowly replicate under treatment, as shown by Carboxy fluorescein succinimidyl ester (CFSE) analysis and staining with 5-ethynyl-2’-deoxyuridine (EdU). We used these population dynamics parameters to design the MC-LD assay. CRC clones were plated in several 96-multiwell plates each, and after an expansion phase in standard culture conditions, treatment was added. After 3-4 weeks, a minority of wells showed growth of resistant colonies: based on the measured growth rates, we could predict that the resistant cells arose before treatment by spontaneous mutation. The remaining wells contained a homogenous population of DTPs. After several weeks of treatment, when pre-treatment resistant clones would have already emerged, late-emerging resistant colonies appeared in a subset of wells in which DTPs had previously been detected. Using the number of residual DTPs and resistant colonies to infer mutation rates, we found a 7- to 50-fold increase (depending on the cell line) in DTPs’ mutation rate compared to sensitive cells.In conclusion, we developed a new assay which allows quantitative comparisons of spontaneous and drug-induced mutation rates in cancer cells and showed that adaptive mutability in DTPs leads to increased mutation rates. This approach could be used to measure whether and how a wide range of environmental conditions affect DTP phenotype and mutation rates in mammalian cells. Citation Format: Alberto Sogari, Mariangela Russo, Simone Pompei, Mattia Corigliano, Giovanni Crisafulli, Andrea Bertotti, Marco Gherardi, Federica Di Nicolantonio, Marco Cosentino Lagomarsino, Alberto Bardelli. A modified Luria-Delbrück assay allows quantification of colorectal cancer persister cells’ mutation rate [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 2613.
Modern deep neural networks (DNNs) represent a formidable challenge for theorists: according to the commonly accepted probabilistic framework that describes their performance, these architectures should overfit due to the huge number of parameters to train, but in practice they do not. Here we employ results from replica mean field theory to compute the generalization gap of machine learning models with quenched features, in the teacher-student scenario and for regression problems with quadratic loss function. Notably, this framework includes the case of DNNs where the last layer is optimized given a specific realization of the remaining weights. We show how these results-combined with ideas from statistical learning theory-provide a stringent asymptotic upper bound on the generalization gap of fully trained DNN as a function of the size of the dataset P. In particular, in the limit of large P and N_{out} (where N_{out} is the size of the last layer) and N_{out}≪P, the generalization gap approaches zero faster than 2N_{out}/P, for any choice of both architecture and teacher function. Notably, this result greatly improves existing bounds from statistical learning theory. We test our predictions on a broad range of architectures, from toy fully connected neural networks with few hidden layers to state-of-the-art deep convolutional neural networks.
Physics can be seen as a conceptual approach to scientific problems, a method for discovery, but teaching this aspect of our discipline can be a challenge. We report on a first-time remote teaching experience for a computational physics third-year physics laboratory class taught in the first part of the 2020 COVID-19 pandemic (March-May 2020). To convey a 'physics of data' approach to data analysis and data-driven physical modeling we used interdisciplinary data sources, with an openended 'COVID-19 data challenge' project as the core of the course. COVID-19 epidemiological data provided an ideal setting for motivating the students to deal with complex problems, where there is no unique or preconceived solution. Our results indicate that such problems yield qualitatively different improvements compared to close-ended projects, as well as point to critical aspects in using these problems as a teaching strategy. By breaking the students' expectations of unidirectionality, remote teaching provided unexpected opportunities to promote active work and active learning.
Marco Cosentino Lagomarsino, 2 Guglielmo Pacifico, Valerio Firmano, Edoardo Bella, Pietro Benzoni, Jacopo Grilli, Federico Bassetti, Fabrizio Capuani, Pietro Cicuta, and Marco Gherardi IFOM Foundation, FIRC Institute of Molecular Oncology, Milan, Italy Dipartimento di Fisica, Università degli Studi di Milano, and I.N.F.N, via Celoria 16 Milano, Italy∗ Dipartimento di Fisica, Università degli Studi di Milano, via Celoria 16 Milano, Italy The Abdus Salam International Centre for Theoretical Physics (ICTP), Trieste, Italy. Politecnico di Milano, Piazza Leonardo da Vinci 32, Milano, Italy I.N.F.N. Sezione di Roma Sapienza. Piazzale Aldo Moro 2, Roma, Italy Cavendish Laboratory, Cambridge University, Cambridge, United Kingdom
In critical systems, the effect of a localized perturbation affects points that are arbitrarily far from the perturbation location. In this paper, we study the effect of localized perturbations on the solution of the random dimer problem in two dimensions. By means of an accurate numerical analysis, we show that a local perturbation of the optimal covering induces an excitation whose size is extensive with finite probability. We compute the fractal dimension of the excitations and scaling exponents. In particular, excitations in random dimer problems on nonbipartite lattices have the same statistical properties of domain walls in spin glass. Excitations produced in bipartite lattices, instead, are compatible with a loop-erased self-avoiding random walk process. In both cases, we find evidence of conformal invariance of the excitations that is compatible with SLE_{κ} with parameter κ depending on the bipartiteness of the underlying lattice only.
Compelling evidence shows that cancer persister cells limit the efficacy of targeted therapies. However, it is unclear whether persister cells are induced by anticancer drugs, and if their mutation rate quantitatively increases during treatment. Here, combining experimental characterization and mathematical modeling, we show that, in colorectal cancer, persisters are induced by drug treatment and show a 7- to 50-fold increase of mutation rate when exposed to clinically approved targeted therapies. These findings reveal that treatment may influence persistence and mutability in cancer cells and pinpoints new strategies to restrict tumor recurrence.
Linear separability, a core concept in supervised machine learning, refers to whether the labels of a data set can be captured by the simplest possible machine: a linear classifier. In order to quantify linear separability beyond this single bit of information, one needs models of data structure parameterized by interpretable quantities, and tractable analytically. Here, I address one class of models with these properties, and show how a combinatorial method allows for the computation, in a mean field approximation, of two useful descriptors of linear separability, one of which is closely related to the popular concept of storage capacity. I motivate the need for multiple metrics by quantifying linear separability in a simple synthetic data set with controlled correlations between the points and their labels, as well as in the benchmark data set MNIST, where the capacity alone paints an incomplete picture. The analytical results indicate a high degree of “universality”, or robustness with respect to the microscopic parameters controlling data structure.
Recent results comparing the temporal program of genome replication of yeast species belonging to the Lachancea clade support the scenario that the evolution of the replication timing program could be mainly driven by correlated acquisition and loss events of active replication origins. Using these results as a benchmark, we develop an evolutionary model defined as birth-death process for replication origins and use it to identify the evolutionary biases that shape the replication timing profiles. Comparing different evolutionary models with data, we find that replication origin birth and death events are mainly driven by two evolutionary pressures, the first imposes that events leading to higher double-stall probability of replication forks are penalized, while the second makes less efficient origins more prone to evolutionary loss. This analysis provides an empirically grounded predictive framework for quantitative evolutionary studies of the replication timing program.
Background: DNA bridging promoted by the H-NS protein, combined with the compaction induced by cellular crowding, plays a major role in the structuring of the E. coli genome. However, only few studies consider the effects of the physical interplay of these two factors in a controlled environment. Methods: We apply a single molecule technique (Magnetic Tweezers) to study the nanomechanics of compaction and folding kinetics of a 6 kb DNA fragment, induced by H-NS bridging and/or PEG crowding. Results: In the presence of H-NS alone, the DNA shows a step-wise collapse driven by the formation of multiple bridges, and little variations in the H-NS concentration-dependent unfolding force. Conversely, the DNA collapse force observed with PEG was highly dependent on the volume fraction of the crowding agent. The two limit cases were interpreted considering the models of loop formation in a pulled chain and pulling of an equilibrium globule respectively. Conclusions: We observed an evident cooperative effect between H-NS activity and the depletion of forces induced by PEG. General Significance: Our data suggest a double role for H-NS in enhancing compaction while forming specific loops, which could be crucial in vivo for defining specific mesoscale domains in chromosomal regions in response to environmental changes.
Many machine learning algorithms used for dimensional reduction and manifold learning leverage on the computation of the nearest neighbors to each point of a data set to perform their tasks. These proximity relations define a so-called geometric graph, where two nodes are linked if they are sufficiently close to each other. Random geometric graphs, where the positions of nodes are randomly generated in a subset of R^{d}, offer a null model to study typical properties of data sets and of machine learning algorithms. Up to now, most of the literature focused on the characterization of low-dimensional random geometric graphs whereas typical data sets of interest in machine learning live in high-dimensional spaces (d≫10^{2}). In this work, we consider the infinite dimensions limit of hard and soft random geometric graphs and we show how to compute the average number of subgraphs of given finite size k, e.g., the average number of k cliques. This analysis highlights that local observables display different behaviors depending on the chosen ensemble: soft random geometric graphs with continuous activation functions converge to the naive infinite-dimensional limit provided by Erdös-Rényi graphs, whereas hard random geometric graphs can show systematic deviations from it. We present numerical evidence that our analytical results, exact in infinite dimensions, provide a good approximation also for dimension d≳10.
The traditional approach of statistical physics to supervised learning routinely assumes unrealistic generative models for the data: Usually inputs are independent random variables, uncorrelated with their labels. Only recently, statistical physicists started to explore more complex forms of data, such as equally labeled points lying on (possibly low-dimensional) object manifolds. Here we provide a bridge between this recently established research area and the framework of statistical learning theory, a branch of mathematics devoted to inference in machine learning. The overarching motivation is the inadequacy of the classic rigorous results in explaining the remarkable generalization properties of deep learning. We propose a way to integrate physical models of data into statistical learning theory and address, with both combinatorial and statistical mechanics methods, the computation of the Vapnik-Chervonenkis entropy, which counts the number of different binary classifications compatible with the loss class. As a proof of concept, we focus on kernel machines and on two simple realizations of data structure introduced in recent physics literature: k-dimensional simplexes with prescribed geometric relations and spherical manifolds (equivalent to margin classification). Entropy, contrary to what happens for unstructured data, is nonmonotonic in the sample size, in contrast with the rigorous bounds. Moreover, data structure induces a transition beyond the storage capacity, which we advocate as a proxy of the nonmonotonicity, and ultimately a cue of low generalization error. The identification of a synaptic volume vanishing at the transition allows a quantification of the impact of data structure within replica theory, applicable in cases where combinatorial methods are not available, as we demonstrate for margin learning.
Background: The validation of prognostic indices might prove instrumental for optimal management of motoneuron disease (MND) patients. Aim: To test the role of PaO2, HCO3-, and vital capacity (VC) at diagnosis in predicting the prognosis of MND patients, both in terms of time to the start of non-invasive ventilation (NIV) and to death. Patients and Methods: We retrospectively analyzed the records of 192 MND patients followed-up in our dedicated MND clinic. All patients underwent blood gas analysis and spirometry (VC, both in the sitting and supine position) at the time of diagnosis; time to non-invasive ventilation (NIV) and death was recorded. Kaplan-Meier curves were built and differences between groups based on predefined thresholds were analyzed by the log-rank test using GraphPad Prism software; p-values<0.05 were considered statistically significant. Results: The figure shows the Kaplan-Meier curves for time to NIV. All curves were significantly different for patients above and below the predefined thresholds. Median time to death was different for patients with HCO3- above and below the threshold of 26 mmol/L (29 and 68 months, respectively; p<0.05). Conclusions: As expected, PaO2 and VC (particularly in the supine position) at diagnosis are relevant prognostic determinants in MND patients. HCO3-, that can be easily measured in venous blood, is emerging as a reliable prognostic index in these patients.