Large language models and autonomous AI agents have evolved rapidly, resulting in a diverse array of evaluation benchmarks, frameworks, and collaboration protocols. Driven by the growing need for standardized evaluation and integration, we systematically consolidate these fragmented efforts into a unified framework. However, the landscape remains fragmented and lacks a unified taxonomy or comprehensive survey. Therefore, we present a side-by-side comparison of benchmarks developed between 2019 and 2025 that evaluate these models and agents across multiple domains. In addition, we propose a taxonomy of approximately 60 benchmarks that cover general and academic knowledge reasoning, mathematical problem-solving, code generation and software engineering, factual grounding and retrieval, domain-specific evaluations, multimodal and embodied tasks, task orchestration, and interactive assessments. Furthermore, we review AI-agent frameworks introduced between 2023 and 2025 that integrate large language models with modular toolkits to enable autonomous decision-making and multi-step reasoning. Moreover, we present real-world applications of autonomous AI agents in materials science, biomedical research, academic ideation, software engineering, synthetic data generation, chemical reasoning, mathematical problem-solving, geographic information systems, multimedia, healthcare, and finance. We then survey key agent-to-agent collaboration protocols, namely the Agent Communication Protocol (ACP), the Model Context Protocol (MCP), and the Agent-to-Agent Protocol (A2A). Finally, we discuss recommendations for future research, focusing on advanced reasoning strategies, failure modes in multi-agent LLM systems, automated scientific discovery, dynamic tool integration via reinforcement learning, integrated search capabilities, and security vulnerabilities in agent protocols.
Autonomous AI agents powered by large language models (LLMs) with structured function-calling interfaces have greatly expanded capabilities for real-time data retrieval, computation, and multi-step orchestration. However, the rapid growth of plugins, connectors, and inter-agent protocols has outpaced security practices, leading to brittle integrations — plugin APIs and protocol adapters that rely on ad-hoc authentication, inconsistent schemas, and weak validation — making them vulnerable to failures and exploitation. This survey introduces a unified end-to-end threat model for LLM-agent ecosystems, spanning host-to-tool and agent-to-agent communications, and catalogs over thirty attack techniques across Input Manipulation, Model Compromise, System and Privacy Attacks, and Protocol Vulnerabilities. For each category, we provide a formal mathematical formulation of the underlying threat model, defining attacker capabilities, objectives, and affected layers to enable systematic analysis. Representative examples include Prompt-to-SQL (P2SQL) injections and the Toxic Agent Flow exploit in GitHub’s MCP server. For each category, we assess feasibility, review defenses, and outline mitigation strategies such as dynamic trust management, cryptographic provenance tracking, and sandboxed agentic interfaces. The framework was validated through expert review and cross-mapping with real-world incidents and public vulnerability repositories (e.g., CVE, NIST NVD) to ensure practical relevance. Compared to prior surveys, this work provides the first integrated taxonomy bridging input-level exploits and protocol-layer vulnerabilities in LLM-agent ecosystems while introducing formal system definitions for each threat class. Ultimately, it offers actionable insights for securing next-generation AI agents through layered defense and continuous verification. Our work provides a comprehensive reference to guide the design of secure and resilient LLM-agent workflows.
The conventional processes to tackle solutions of linear Fredholm integral equations over real intervals of substantial but finite lengths is based on two steps: First, we discretize the equation which gives a huge linear algebraic system. Second, we need to use some iterative schemes to approach the solution of this algebraic system. In this paper, we propose a new process denoted by iterate-discretize, which leads to the development of the iterative theory tackling this type of Fredholm integral equations over such intervals. This new process is dependent on two parts: We transform the Fredholm equation into a matrix of bounded linear operators defined over small subintervals, whence we construct a new functional variant of the iterative SOR scheme adapted to this operators’ matrix. As a second part, we don’t need to discretize the whole operators’ matrix but only its diagonal part. Within the context of study, some properties are provided to make sure the convergence of the generalized SOR scheme. Ultimately, a regular and weakly singular equation are studied and we obtain positive results that ground the new process. Moreover, we provide graphical approximations to the optimal relaxation parameter ω for which convergence is optimized.
Causal inference is a crucial framework in a variety of fields, such as economics, healthcare, and social science. In this context, data-driven machine learning models have become more popular for estimating the effects of treatments. One fundamental technique for causal inference is structural equation modeling (SEM), which makes it possible to estimate the relationships between variables. However, conventional SEM methods need help with the complexities of structural causal models (SCM) in real-world data, including latent confounders, nonlinear relationships, and challenges in accurately specifying model structures. These limitations could cause the interpretation or biased causal effects. To address these challenges, we introduce a new proposed method combining the piecewise structural equation modeling (PSEM) with the backdoor criterion, named PSEMBC. The main innovation of PSEMBC is the estimation of causal effects with the use of a linear SCM model, which captures complex relationships and interactions in the data. We demonstrate the value of PSEMBC for precisely and reliably identifying the average treatment effect (ATE) in simulated and real-world datasets utilizing a comparative study with current causal inference approaches.
We investigate a stochastic eco-epidemiological framework where disease transmission occurs within the prey population and interacts with the predators under environmental noise. By employing stochastic Lyapunov methods and Khasminskii’s theory, we establish persistence-extinction conditions and confirm the existence of an ergodic stationary distribution. Furthermore, we determine the threshold parameters that govern the long-term dynamics of the system. Theoretical results are supported through numerical simulations which illustrate how variations in key biological and noise parameters affect the coexistence and extinction of species. The study highlights the significant influence of stochastic perturbations on disease dynamics and provides useful insights into the stability of eco-epidemiological systems under random environmental effects.