This paper addresses a multi-agent Homicidal Chauffeur reach-avoid game, where M simple-motion evaders aim to reach the goal region, while N Dubins-car pursuers aim to prevent that by capturing evaders, and it is assumed that N ≥ M and only incomplete information can be obtained by pursuers. We begin with one-pursuer-one-evader situation, and the optimal path types of the agents are inferred by the Hamiltonian of the differential game based on Homicidal Chauffeur dynamical model. Further, the optimal strategies of the agents are deduced from the engagement line we construct. Afterwards, the situation is expanded to two-pursuer-one-evader, where the intersection of engagement lines is introduced to describe the cooperation of two pursuers. Then, we analyze the N-pursuer-M-evader game by the combination of one-pursuer-one-evader and two-pursuer-one-evader games, and a potential game-based matching approach is proposed to achieve the combination with incomplete information. Simulation and reality experiments are provided to illustrate the theoretical results, where Joint Strategy Fictitious Play (JSFP) with inertia algorithm is executed to obtain the pure strategy Nash equilibrium. It is verified that the optimal pursuer-evader assignment is switchable and our scheme is feasible.
We study a finite-horizon linear–quadratic mean field game with heterogeneous observations of the initial mean field. These observations may be erroneous and are used by agents to construct their feedback laws. At the population level, the heterogeneous error profile is generally infinite-dimensional. We derive exact linear sensitivity formulas showing that its closed-loop effects nevertheless admit a finite-dimensional closure: for each agent, the deviations are governed by only two n-dimensional error channels, its private error E_i and the population-average error E̅. The population-average error determines the displacement of the actual mean field, whereas centered private errors determine agent-specific deviations and the additional cross-sectional covariance. The same representation yields quadratic cost identities and a uniform O(N^-1) mean-square approximation of the empirical aggregate. In the deterministic model, the private state history induces a linear observation problem for the initial errors. Nonsingularity of the associated observability Gramian is necessary and sufficient for exact recovery. The recovered errors reconstruct the actual mean-field state at the revision time and initialize a revised correct-information continuation from that state. We obtain an exact continuation-cost comparison; when the population-average error vanishes, revision does not increase the continuation cost. For stochastic dynamics, we take a two-component revision signal as given and study the resulting one-shot feedback relative to an oracle continuation initialized with the actual current mean field. The actual mean-field deviation is driven only by the population-average signal error, while individual deviations also retain the private signal error. Exact covariance and quadratic continuation-cost identities quantify these effects.
Neural network-based methods have demonstrated effectiveness in solving high-dimensional Mean-Field Games (MFG) equilibria, yet ensuring mathematically consistent density-coupled evolution remains a major challenge. This paper proposes the NF-MKV Net, a neural network approach that integrates process-regularized normalizing flow (NF) with state-policy-connected time-series neural networks to solve MKV FBSDEs and their associated fixed-point formulations of MFG equilibria. The method first reformulates MFG equilibria as MKV FBSDEs, embedding density evolution into the equation coefficients within a probabilistic framework. Neural networks are then employed to approximate value functions and their gradients. To enforce volumetric invariance and temporal continuity, NF architectures impose loss constraints on each density transfer function. Theoretical analysis establishes the algorithm's validity, while numerical experiments across various scenarios including traffic flow, crowd motion, and obstacle avoidance, demonstrate its capability in maintaining density consistency and temporal smoothness.
The paper studies a nonlinear mean field game (MFG) model based on McKean-Vlasov (MV) approximation, which involves multiple major agents and multiple populations of minor agents. Since the interactions occur not only between major-minor agents but also among major-major agents and minor-minor populations, complicating equilibrium existence analysis, we first cast the MFG as two sets of adjoint stochastic McKean-Vlasov equations coupled via the mean field behavior. Subsequently, a new norm on the product probability measure space is constructed to prove that one set of equations exists solutions, while the existence of solutions to the other set is proved based on Pontryagin stochastic maximum principle. Then, within the established product normed space, Banach's fixed point theorem is utilized to prove the existence of equilibrium for this MFG, which is manifested as solutions to two sets of coupled equations. Finally, an epsilon-Nash equilibrium is proved for the finite agent situation, in which epsilon -> 0 while all population sizes go to infinity, and a numerical experiment under certain settings is carried out.
With the vigorous development of online media, the prevalence of key opinion leaders and water armies has led to unexpected evolutions of users' opinions. Therefore, it is valuable and interesting to investigate the opinion evolution problem for large-scale groups with the coexistence of influential individuals and strongly organized groups. For the above problem, based on the mean-field game theory, this article innovatively proposes a multi-leader multi-population-follower Stackelberg mean field game (MLMPF-SMFG) model to describe the opinion evolution scenario, in which influential individuals are regarded as leaders, normal and strongly organized groups are regarded as follower populations. Moreover, for generality, the types of strongly organized individuals are classified into three typical types: propagandists, spies, and neutrals. Then, the optimal strategies are derived via the adjoint method and solved by forward-backward stochastic differential equations. Sufficient conditions for the existence and uniqueness of the Stackelberg equilibrium (SE) are given, and the approximate SE of the finite system is proven. Finally, simulation experiments on the opinion evolutions of two influential individuals and two ordinary groups are performed to demonstrate the feasibility and effectiveness of the proposed MLMPF-SMFG model.
In this paper, we consider the social optimal problem of discrete time finite state space mean field games (referred to as finite mean field games [1]). Unlike the individual optimization of their own cost function in competitive models, in the problem we consider, individuals aim to optimize the social cost by finding a fixed point of the state distribution to achieve equilibrium in the mean field game. We provide a sufficient condition for the existence and uniqueness of the individual optimal strategies used to minimize the social cost. According to the definition of social optimum and the derived properties of social optimal cost, the existence and uniqueness conditions of equilibrium solutions under initial-terminal value constraints in the finite horizon and the existence and uniqueness conditions of stable solutions in the infinite horizon are given. Finally, two examples that satisfy the conditions for the above solutions are provided.
Cooperation and non-cooperation are two fundamental topics in games, but in multi-population mean-field games (MPMFGs), cooperation may occur at the individual level (intra-population) or at the population level (inter-population) or at the both levels, which leads to different interesting results. In this paper, model frameworks based on linear quadratic mean-field games are developed for the above three possible hierarchical cooperation cases. The optimal controls are derived via the adjoint method and solved by forward-backward ordinary differential equations. Based on the analysis of asymmetric Riccati differential equations, sufficient conditions for the existence and uniqueness of optimal solutions are given. Then, the price of anarchy in different situations are defined to measure the efficiency of the proposed MPMFG models. Finally, simulation experiments on the opinion evolutions of two populations in social networks are performed to demonstrate the feasibility and effectiveness of the above models. Specially, it is presented that the experimental results of the intra-population non-cooperative and inter-population cooperative MPMFG model provide an explanation of the behavior of coalition governments over electoral cycles studied in the existing literature.
In this paper, the initial error affection and strategy modification in multi-population linear quadratic mean field games (MPLQMFGs) under erroneous initial distribution information are investigated. First, a MPLQMFG model is developed where agents in different populations are coupled by dynamics and cost functions. Next, by studying the evolutionary of MPLQMFGs under erroneous initial distribution information, the predicted and the actual evolutionary of mean field states are given. Furthermore, assume that each agent maintains observations only of its own state and control as well as the mean field terms of its own population, and agents are allowed to modify their strategies at an intermediate moment, two sufficient conditions are provided, where the game will reach the Nash equilibrium under correct information. Besides, the affection of the initial error on the game is discussed. Finally, simulations on the opinion evolutions of two groups are performed to verify above conclusions.
This paper proposes a two-level hierarchical control approach based on pursuit-evasion game and mean field game for the problem of large-scale pursuers with multi-population against a single evader, which implements that the evader is surrounded by pursuers. At the upper layer, we model the pursuit-evasion game between the centers of the pursuer populations and the single evader, which is formulated as a linear quadratic differential game (LQDG) to obtain the optimal control of each player. Then the optimal trajectories derived from the optimal controls are input to the lower layer as the reference trajectories. At the lower layer, we formulate the tracking of reference trajectories and terminal surrounding to the evader of large-scale pursuers with multi-population as a multi-population mean-field game (MPMFG), which solves the communication and computing difficulties caused by large-scale agents. Then, we derive the variational primal–dual formulation of the proposed MPMFG model and solve it with CA-Net, a coupled alternating neural network approach. Finally, simulation experiments are performed under various pursuit-evasion scenarios, and it is verified that the proposed game-based two-level hierarchical control approach is feasible and effective.
The mechanical LiDAR sensor is crucial in autonomous vehicles. After projecting a 3D point cloud onto a 2D plane and employing a deep learning model for computation, accurate environmental perception information can be supplied to autonomous vehicles. Nevertheless, the vertical angular resolution of inexpensive multi-beam LiDAR is limited, constraining the perceptual and mobility range of mobile entities. To address this problem, we propose a point cloud super-resolution model in this paper. This model enhances the density of sparse point clouds acquired by LiDAR, consequently offering more precise environmental information for autonomous vehicles. Firstly, we collect two datasets for point cloud super-resolution, encompassing CARLA32-128in simulated environments and Ruby32-128 in real-world scenarios. Secondly, we propose a novel temporal and spatial feature-enhanced point cloud super-resolution model. This model leverages temporal feature attention aggregation modules and spatial feature enhancement modules to fully exploit point cloud features from adjacent timestamps, enhancing super-resolution accuracy. Ultimately, we validate the effectiveness of the proposed method through comparison experiments, ablation studies, and qualitative visualization experiments conducted on the CARLA32-128 and Ruby32-128 datasets. Notably, our method achieves a PSNR of 27.52 on CARLA32-128 and a PSNR of 24.82 on Ruby32-128, both of which are better than previous methods.
This paper investigates the use of reinforcement learning for autonomous exploration in an unknown environment. Autonomous exploration is crucial in many situations, such as urban search, security inspection, environmental mapping, etc. Traditional approaches focused on frontiers are unlikely to span a variety of enormously complex scenarios. Convergence is a little more difficult for learning-based approaches, which can adapt to many different environments. Consequently, a hierarchical exploration framework is built using frontier information. We propose a reinforcement learning-based local decision exploration model that uses deep neural networks to learn the optimal strategy from the environment. To prevent falling into local optimization, we also suggest a global rescue module to assist the robot in returning to the proper exploration track. Compared with other hierarchical methods, the framework is more effective and resilient in many contexts, greatly decreasing the total completion time and path length.
As a branch of statistical latent variable modeling, multidimensional item response theory (MIRT) plays an important role in psychometrics. Multidimensional graded response model (MGRM) is a key model for the development of multidimensional computerized adaptive testing (MCAT) with graded-response data and multiple traits. This paper explores how to automatically identify the item-trait patterns of replenished items based on the MGRM in MCAT. The problem is solved by developing an exploratory pattern recognition method for graded-response items based on the least absolute shrinkage and selection operator (LASSO), which is named LPRM-GR and facilitates the subsequent parameter estimation of replenished items and helps maintaining the effectiveness of item replenishment in MCAT. In conjunction with the proposed approach, the regular BIC and weighted BIC are applied, respectively, to select the optimal item-trait patterns. Simulation for evaluating the LPRM-GR in pattern recognition accuracy of replenished items and the corresponding item estimation accuracy is conducted under multiple conditions across different numbers with respect to dimensionality, response-category numbers, latent trait correlation, stopping rules, and item selection criteria. Results show that the proposed method with the two types of BIC both have good performance in pattern recognition for item replenishment in the two- to four-dimensional MCAT with the MGRM, for which the weighted BIC is generally superior to the regular BIC. The proposed method has relatively high accuracy and efficiency in identifying the patterns of graded-response items, and has the advantages of easy implementation and practical feasibility.
Trajectory planning of massive unmanned aerial vehicles (UAVs) is very difficult in an environment with static and dynamic obstacles. This is mainly due to the huge number of UAVs, which pose challenges to their interaction and collision avoidance with companions and obstacles. In this paper, we propose a trajectory planning algorithm for a massive number of UAVs based on the mean field game (MFG). First, a differential game of N UAVs in a 3D environment is constructed, and the collision avoidance with static and dynamic obstacles is considered in the cost functional of each UAV. Then, when the number of UAVs is very large, the above differential game is transformed into a MFG using the mean field approximation. The existence and uniqueness of the equilibrium solution are proved. Finally, we derive the variational primal-dual formulation of the proposed MFG model and solve it with APAC-Net. The performance of the proposed algorithm is validated in an environment with multiple static obstacles and two different types of dynamic obstacles.
Dynamic scene understanding based on LiDAR point clouds is one of the critical perception tasks for self-driving vehicles. Among these tasks, point cloud semantic segmentation is highly challenging. Some existing work ignores the loss of crucial information caused by sampling and projecting. Others use modules with high computational complexity because of the pursuit of precision, challenging to deploy in the vehicle platform with limited computing power. This paper proposes Fusedown/Fuse-up modules for efficient down-sampling/up-sampling feature extraction. The modules combine the transformer in vision integrating the global information of the feature map with the CNN extracting local feature information. Based on these two modules, we built the transformer and CNN fusion network called TCFNet for point cloud semantic segmentation. Experiments on the SemanticKITTI show that our suitable combination of transformer and CNN is necessary for semantic segmentation accuracy, and the mIoU of our model can reach 82.7% at 10 FPS. The code can be accessed at https://github.com/donkeyofking/TCFNet.git.
The traditional optimization and control technologies deal with the dynamic interactions between individuals separately, with the increase in the agents' number, the modeling process of cooperative attack-defense problems tends to be complex, and the difficulty of solving the optimal strategy will increase significantly. Moreover, to carry out more accurate real-time control of agents, the state variables used to characterize their kinematics are usually high-dimensional. To overcome these challenges, we formulate the cooperative attack-defense evolution of large-scale agents as a multi-population high-dimensional stochastic mean-field game (MPHD-MFG). Numerical methods for MPHD-MFGs are practically non-existent, because, the heterogeneity of the multi-population model increases the complexity of sequential games, and grid-based spatial discretization leads to dimension explosion. Thus, we propose a generative adversarial network-based method, where we use a coupled alternating neural network composed of multiple generators and multiple discriminators, to tractably solve MPHD-MFGs. Simulation experiments are carried out for various attack-defense scenarios, the results verify the feasibility and effectiveness of our proposed model and algorithm.
The traditional optimization and control technologies deal with the dynamic interactions between individuals separately, with the increase in the agents' number, the modeling process of cooperative attack-defense problems tends to be complex, and the difficulty of solving the optimal strategy will increase significantly. Moreover, to carry out more accurate real-time control of agents, the state variables used to characterize their kinematics are usually high-dimensional. To overcome these challenges, we formulate the cooperative attack-defense evolution of large-scale agents as a multi-population high-dimensional stochastic mean-field game (MPHD-MFG). Numerical methods for MPHD-MFGs are practically non-existent, because, the heterogeneity of the multi-population model increases the complexity of sequential games, and grid-based spatial discretization leads to dimension explosion. Thus, we propose a generative adversarial network-based method, where we use a coupled alternating neural network composed of multiple generators and multiple discriminators, to tractably solve MPHD-MFGs. Simulation experiments are carried out for various attack-defense scenarios, the results verify the feasibility and effectiveness of our proposed model and algorithm.