In this work, we propose a novel layerwise adaptive construction method for neural network architectures. Our approach is based on a goal--oriented dual-weighted residual technique for the optimal control of neural differential equations. This leads to an ordinary differential equation constrained optimization problem with controls acting as coefficients and a specific loss function. We implement our approach on the basis of a DG(0) Galerkin discretization of the neural ODE, leading to an explicit Euler time marching scheme. For the optimization we use steepest descent. Finally, we apply our method to the construction of neural networks for the classification of data sets, where we present results for a selection of well known examples from the literature.
We consider an optimal control problem for the heat equation as a prototypical parabolic partial differential equation with a non-convex control mechanism of the form continuous-or-off. We model this fundamental switching mechanism as the product of a classically continuous and a binary control both in the control term of the dynamics and in the objective. A total variation regularization is added to the cost in order to restrict the number of switching times. This leads to a mixed-integer non-linear PDE-constrained problem. We discuss well-posedness of the problem and present an exact relaxation result for a linearized and a trust-region type penalized problem. The exactness result is constructive and provides a way to numerically compute mixed-integer optimal solutions from the optimality conditions of an associated PDE-constrained problem without integer restrictions. It lays a foundation for a new class of sequential relaxation algorithms to solve the considered class of mixed-integer control problems. This is demonstrated numerically by showcasing a descent step in the presence of binary restrictions.
We propose constrained neural parameterization schemes for several classes of constraints arising in optimization problems in function spaces. This is achieved by constructing smooth neural parameterizations whose image lies entirely in the admissible set while remaining asymptotically dense. In this way, the original constrained optimization problem is transformed into a smooth unconstrained problem in parameter space, enabling efficient gradient-based optimization without penalty parameters or Lagrange multipliers. We develop geometric constructions for polyhedral constraint sets in Hilbert spaces, propose smooth neural architectures for some pointwise constraints, and introduce an exact reduced neural formulation for PDE constraints that admit a separable structure. Numerical experiments demonstrate the effectiveness of the proposed methods.
Primal-dual active set strategies (PDAS) are popular iterative solvers for mixed complementarity problems such as constrained optimization problems with pointwise inequality constraints. Examples include the reduced-space active set algorithm vinewtonrsls found in PETSc. When applied to discretized infinite-dimensional problems, PDAS exhibit local superlinear convergence thanks to their equivalence to a semismooth Newton method (SSN). However, for many problem classes the number of iterations, to reach convergence, grows without bound under mesh refinement. In this paper we numerically study PDAS iteration counts on uniformly refined meshes for obstacle problems, Signorini problems, and related models. As the mesh size tends to zero, PDAS applied to Signorini-type problems lose their local superlinear convergence, resulting in linear growth of the iteration count (adding some iterations with each refinement). For obstacle problems, PDAS stagnates, leading to exponential iteration growth (asymptotically doubling with each refinement). We explain these phenomena by (i) proving that, for obstacle problems, nodal degrees of freedom only peel away from the obstacle layer-by-layer during the deactivation phase, (ii) deriving a general global convergence rate for PDAS that depends on the magnitude of dual feasibility violation, and (iii) demonstrating why, in the infinite-dimensional setting, this leads to a well-defined solver, but without local superlinear convergence, for some problems yet divergence for others.
In this work, we study physics-informed neural networks (PINNs) constrained by partial differential equations (PDEs) and their application in approximating PDEs with two characteristic scales. From a continuous perspective, our formulation corresponds to a non-standard PDE-constrained optimization problem with a PINN-type objective. From a discrete standpoint, the formulation represents a hybrid numerical solver that utilizes both neural networks and finite elements. For the problem analysis, we introduce a proper function space, and we develop a numerical solution algorithm. The latter combines an adjoint-based technique for the efficient gradient computation with automatic differentiation. This new multiscale method is then applied exemplarily to a heat transfer problem with oscillating coefficients. In this context, the neural network approximates a fine-scale problem, and a coarse-scale problem constrains the associated learning process. We demonstrate that incorporating coarse-scale information into the neural network training process via a weak convergence-based reg-ularization term is beneficial. Indeed, while preserving upscaling consistency, this term encourages non-trivial PINN solutions and also acts as a preconditioner for the low-frequency component of the fine-scale PDE, resulting in improved convergence properties of the PINN method. The relevance of our approach to material science is discussed.
This study introduces a hybrid machine learning-based scale-bridging framework for predicting the permeability of fibrous textile structures. By addressing the computational challenges inherent to multiscale modeling, the proposed approach evaluates the efficiency and accuracy of different scale-bridging methodologies combining traditional surrogate models and even integrating physics-informed neural networks (PINNs) with numerical solvers, enabling accurate permeability predictions across micro- and mesoscales. Four methodologies were evaluated: fully resolved models (FRM), numerical upscaling method (NUM), scale-bridging method using data-driven machine learning methods (SBM) and a hybrid dual-scale solver incorporating PINNs. The FRM provides the highest fidelity model by fully resolving the micro- and mesoscale structural geometries, but requires high computational effort. NUM is a fully numerical dual-scale approach that considers uniform microscale permeability but neglects the microscale structural variability. The SBM accounts for the variability through a segment-wise assigned microscale permeability, which is determined using the data-driven ML method. This method shows no significant improvements over NUM with roughly the same computational efficiency and modeling runtimes of 45 min per simulation. The newly developed hybrid dual-scale solver incorporating PINNs shows the potential to overcome the problem of data scarcity of the data-driven surrogate approaches, as well as incorporating data from both scales via the hybrid loss function. The hybrid framework advances permeability modeling by balancing computational cost and prediction reliability, laying the foundation for further applications in fibrous composite manufacturing, while its full potential awaits realization as physics-informed machine learning approaches continue to mature.
We present a novel model of a coupled hydrogen and electricity market on the intraday time scale, where hydrogen gas is used as a storage device for the electric grid. Electricity is produced by renewable energy sources or by extracting hydrogen from a pipeline that is shared by non-cooperative agents. The resulting model is a generalized Nash equilibrium problem. Under certain mild assumptions, we prove that an equilibrium exists. Perspectives for future work are presented.
We study the existence of equilibria for state constrained, convex generalized Nash equilibrium problems (GNEPs) coupled with hyperbolic partial differential equations (PDEs). Analogous problems have been addressed for state constrained GNEPs coupled with elliptic and parabolic PDEs, respectively, but the problematic regularity of the set-valued constraint maps has been a barrier for development of an existence theory in the hyperbolic case. This is mainly due to compactness issues with the strategy-to-state maps. Beyond existence, we also provide first-order optimality conditions for a class of fairly general linear hyperbolic PDEs and show its relevance for applications such as the wave equation, advertising dynamics, and the linearized isothermal Euler system on networks.
In this paper we study the optimal control of a class of semilinear elliptic partial differential equations which have nonlinear constituents that are only accessible by data and are approximated by nonsmooth ReLU neural networks. The optimal control problem is studied in detail. In particular, the existence and uniqueness of the state equation are shown, and continuity as well as directional differentiability properties of the corresponding control-to-state map are established. Based on approximation capabilities of the pertinent networks, we address fundamental questions regarding approximating properties of the learning-informed control-to-state map and the solution of the corresponding optimal control problem. Finally, several stationarity conditions are derived based on different notions of generalized differentiability.
In the field of quantitative imaging, the image information at a pixel or voxel in an underlying domain entails crucial information about the imaged matter. This is particularly important in medical imaging applications, such as quantitative Magnetic Resonance Imaging (qMRI), where quantitative maps of biophysical parameters can characterize the imaged tissue and thus lead to more accurate diagnoses. Such quantitative values can also be useful in subsequent, automatized classification tasks in order to discriminate normal from abnormal tissue, for instance. The accurate reconstruction of these quantitative maps is typically achieved by solving two coupled inverse problems which involve a (forward) measurement operator, typically ill-posed, and a physical process that links the wanted quantitative parameters to the reconstructed qualitative image, given some underlying measurement data. In this review, by considering qMRI as a prototypical application, we provide a mathematically-oriented overview on how data-driven approaches can be employed in these inverse problems eventually improving the reconstruction of the associated quantitative maps.
We develop a semismooth Newton framework for the numerical solution of fixed-point equations that are posed in Banach spaces. The framework is motivated by applications in the field of obstacle-type quasi-variational inequalities and implicit obstacle problems. It is discussed in a general functional analytic setting and allows for inexact function evaluations and Newton steps. Moreover, if a certain contraction assumption holds, we show that it is possible to globalize the algorithm by means of the Banach fixed-point theorem and to ensure q-superlinear convergence to the problem solution for arbitrary starting values. By means of a localization technique, our Newton method can also be used to determine solutions of fixed-point equations that are only locally contractive and not uniquely solvable. We apply our algorithm to a quasi-variational inequality which arises in thermoforming and which not only involves the obstacle problem as a source of nonsmoothness but also a semilinear PDE containing a nondifferentiable Nemytskii operator. Our analysis is accompanied by numerical experiments that illustrate the mesh-independence and q-superlinear convergence of the developed solution algorithm.
A generalized Nash equilibrium problem (GNEP) in Banach space consists of N>1 optimal control problems with couplings in both the objective functions and, most importantly, in the feasible sets. We address the existence of equilibria for convex GNEPs in Banach space. We show that the standard assumption of lower semicontinuity of the set-valued constraint maps - foundational in the current literature on GNEPs - can be replaced by graph convexity or the so-called Knaster-Kuratowski-Mazurkiewicz (KKM) property. Lower semicontinuity is often essential for obtaining upper semicontinuity of best response maps, crucial for the existence theory based on Kakutani-Fan fixed-point arguments. However, in function spaces or PDE-constrained settings, verifying lower semicontinuity becomes much more challenging (even in convex cases), whereas graph convexity, for example, is often straightforward to check. Our results unify several existence theorems in the literature and clarify the structural role of constraint maps. We also extend Rosen's uniqueness condition to Banach spaces using a multiplier bias framework. Additionally, we present a geometric counterpart to our analytic framework using preference maps. This geometric is intended as a complement to, rather than a replacement for, the analytic theory developed in the main body of the paper.
We consider a risk-averse optimal control problem governed by an elliptic variational inequality (VI) subject to random inputs. By deriving KKT-type optimality conditions for a penalised and smoothed problem and studying convergence of the stationary points with respect to the penalisation parameter, we obtain two forms of stationarity conditions. The lack of regularity with respect to the uncertain parameters and complexities induced by the presence of the risk measure give rise to new challenges unique to the stochastic setting. We also propose a path-following stochastic approximation algorithm using variance reduction techniques and demonstrate the algorithm on a modified benchmark problem.
Quasi-variational inequalities (QVIs) of obstacle type in many cases have multiple solutions that can be ordered. We study a multitude of properties of the operator mapping the source term to the minimal or maximal solution of such QVIs. We prove that the solution maps are locally Lipschitz continuous and directionally differentiable and show existence of optimal controls for problems that incorporate these maps as the control-to-state operator. We also consider a Moreau–Yosida-type penalisation for the QVI wherein we show that it is possible to approximate the minimal and maximal solutions by sequences of minimal and maximal solutions (respectively) of certain PDEs, which have a simpler structure and offer a convenient characterisation in particular for computation. For solution mappings of these penalised problems, we prove a number of properties including Lipschitz and differential stability. Making use of the penalised equations, we derive (in the limit) C-stationarity conditions for the control problem, in addition to the Bouligand stationarity we get from the differentiability result.
In this article, we propose a novel regularization method for a class of nonlinear inverse problems that is inspired by an application in quantitative magnetic resonance imaging (qMRI). The latter is a special instance of a general dynamical image reconstruction technique, wherein a radio-frequency pulse sequence gives rise to a time-discrete physics-based mathematical model which acts as a side constraint in our inverse problem. To enhance reconstruction quality, we employ dictionary learning as a data-adaptive regularizer, capturing complex tissue structures beyond handcrafted priors. For computing a solution of the resulting non-convex and non-smooth optimization problem, we alternate between updating the physical parameters of interest via a Levenberg-Marquardt approach and performing several iterations of a dictionary learning algorithm. This process falls under the category of nested alternating optimization schemes. We develop a general overall algorithmic framework whose convergence theory is not directly available in the literature. Global sub-linear and local strong linear convergence in infinite dimensions under certain regularity conditions for the sub-differentials are investigated based on the Kurdyka-Lojasiewicz inequality. Eventually, numerical experiments demonstrate the practical potential and unresolved challenges of the method.
A distributed elliptic control problem with control constraints is considered, which is formulated as a three field problem and consists of two variational equations for the state and the co-state variables as well as of a variational inequality for the control variable. The adjoint control is associated with the residual of the variational inequality but does not appear in the weak formulation. Each of the three variables is discretized independently by hp-finite elements. In particular, the non-penetration condition of the control variable is relaxed to a finite set of quadrature points. Sufficient conditions for the unique existence of a discrete solution are stated. Also a priori error estimates and guaranteed convergence rates are derived in terms of the mesh size as well as of the polynomial degree. Moreover, reliable and efficient a posteriori error estimates are presented, which enable hp-adaptive mesh refinements. Several numerical experiments demonstrate the applicability of the discretization with hp-finite elements, the efficiency of the a posteriori error estimates and the improvements with respect to the convergence order resulting from the application of hp-adaptivity. In particular, the hp-adaptive schemes lead to superior convergence properties.
Many large-scale optimization problems arising in science and engineering are naturally defined at multiple levels of discretization or model fidelity. Multilevel methods exploit this hierarchy to accelerate convergence by combining coarse- and fine-level information, a strategy that has proven highly effective in the numerical solution of partial differential equations and related optimization problems. It turns out that many applications in PDE-constrained optimization and data science require minimizing the sum of smooth and nonsmooth functions. For example, training neural networks may require minimizing a mean squared error plus an L^1-regularization to induce sparsity in the weights. Correspondingly, we introduce a multilevel proximal trust-region method to minimize the sum of a nonconvex, smooth and a convex, nonsmooth function. Exploiting ideas from the multilevel literature allows us to reduce the cost of the step computation, which is a major bottleneck in single level procedures. Our work unifies theory behind the proximal trust-region methods and multilevel recursive strategies. We prove global convergence of our method in finite dimensional space and provide an efficient nonsmooth subproblem solver. We show the efficiency and robustness of our algorithm by means of numerical examples in PDE constrained optimization and machine-learning.
We develop a weak adversarial approach to solving obstacle problems using neural networks. By employing (generalised) regularised gap functions and their properties we rewrite the obstacle problem (which is an elliptic variational inequality) as a minmax problem, providing a natural formulation amenable to learning. Our approach, in contrast to much of the literature, does not require the elliptic operator to be symmetric. We provide an error analysis for suitable discretisations of the continuous problem, estimating in particular the approximation and statistical errors. Parametrising the solution and test function as neural networks, we apply a modified gradient descent ascent algorithm to treat the problem and conclude the paper with various examples and experiments. Our solution algorithm is in particular able to easily handle obstacle problems that feature biactivity (or lack of strict complementarity), a situation that poses difficulty for traditional numerical methods.
We propose and analyze a numerical algorithm for solving a class of optimal control problems for learning-informed semilinear partial differential equations. The latter is a class of PDEs with constituents that are in principle unknown and are approximated by nonsmooth ReLU neural networks. We first show that a direct smoothing of the ReLU network with the aim to make use of classical numerical solvers can have certain disadvantages, namely potentially introducing multiple solutions for the corresponding state equation. This motivates us to devise a numerical algorithm that treats directly the nonsmooth optimal control problem, by employing a descent algorithm inspired by a bundle-free method. Several numerical examples are provided and the efficiency of the algorithm is shown.
The analysis and boundary optimal control of the nonlinear transport of gas on a network of pipelines is considered. The evolution of the gas distribution on a given pipe is modeled by an isothermal semilinear compressible Euler system in one space dimension. On the network, solutions satisfying (at nodes) the Kirchhoff flux continuity conditions are shown to exist in a neighborhood of an equilibrium state. The associated nonlinear optimization problem then aims at steering such dynamics to a given target distribution by means of suitable (network) boundary controls while keeping the distribution within given (state) constraints. The existence of local optimal controls is established and a corresponding Karush–Kuhn–Tucker (KKT) stationarity system with an almost surely non-singular Lagrange multiplier is derived.