Polyhydroxyalkanoates (PHAs) have emerged as a promising alternative to conventional plastics due to their biodegradable and generally favorable biocompatible profile, allowing their application in medical fields, such as drug delivery systems and surgical implants. However, the toxicity assessment of these materials is complex, time-consuming, and costly. Currently, toxicity data for PHAs are limited, dispersed across various studies, and insufficiently reported, which hinders comparative analysis and the development of predictive models. In response to these challenges, recent developments in predictive toxicology have incorporated machine learning-based approaches to estimate toxicological endpoints while reducing the reliance on in vivo experimentation. The present study aims to construct comprehensive, standardized data libraries for the cytotoxicity and ecotoxicity of PHAs and to develop and evaluate polymer-specific machine learning models that link polymer composition to toxicological outcomes. Several computational workflows were designed for this research, with Extra Trees Classifier and Gradient Boosting Classifier being the primary predictive algorithms. Furthermore, Shapley Additive Explanations (SHAP) analysis was performed to identify the descriptors that most strongly influence the predicted cytotoxicity and ecotoxicity. The cytotoxicity model achieved a test Matthews Correlation Coefficient (MCC) of 0.678 and a balanced accuracy of 0.901, with additive type, exposure conditions and particle morphology identified as the most influential descriptors. The ecotoxicity model reached a test MCC of 0.639 and balance accuracy of 0.818 within its applicability domain, with organism- and exposure-level descriptors dominating the predictions. Polymer composition contributed comparatively little, supporting the established biocompatibility of bulk PHB and PHBV. These results, however, should be interpreted in light of the modest dataset size, experimental protocols heterogeneity, and lack of independent experimental validation.
The pharmacokinetic literature is rich in aggregated concentration data that contain valuable information, yet tools to extract this information remain limited. This work introduces distributional physics-informed neural networks (D-PINNs), a novel algorithm designed to enable statistical modelling within the PINN framework, allowing recovery of pharmacokinetic parameter distributions at the population level from published concentration means and variances. Unlike traditional PINNs, which often focus on point estimates, D-PINNs incorporate distributional assumptions directly into the optimisation process. The framework utilises neural networks for predicting the mean and variance of the concentration over time. These predictions are then incorporated into a sampling-based procedure within the residual network, which uses the governing ordinary differential equation (ODE) system to compute the physics-informed loss term. The methodology accounts for both interindividual variability through the parameter distribution and measurement noise through a residual error model. The capability of D-PINNs to infer population-level parameter distributions from concentration summary statistics was demonstrated through a simple proof-of-concept using simulated data from a one-compartment pharmacokinetic model of intravenous drug administration. The model achieved high accuracy in estimating both the parameter distribution and the residual error. Hyperparameter tuning highlighted important aspects of model development. The modelling framework was then applied to real-world data to demonstrate its ability to recover information on the distribution of kinetic parameters in the studied population. Specifically, a minimal physiologically-based pharmacokinetic (mPBPK) model for monoclonal antibodies (mAbs) was fitted to aggregated plasma concentration data reported in the literature using D-PINNs. The same aggregated data were also analysed using a Markov chain Monte Carlo (MCMC) analogue to benchmark the proposed methodology.
Per- and polyfluoroalkyl substances (PFAS) are a large family of persistent environmental contaminants. Some PFAS are known to bioaccumulate, are frequently detected in human serum, and are associated with several adverse effects on the immune system, the endocrine system, and the liver. PFAS-mediated activation of the peroxisome proliferator-activated receptor alpha (PPARα), which plays a key role in lipid and cholesterol homeostasis, is suggested to be an important molecular initiating event triggering PFAS toxicity. The aim of this study was to evaluate the PPARα activation potential of a diverse panel of 34 PFAS congeners, consisting of both legacy and novel compounds, using a PPARα-dependent transactivation assay in transfected HEK293T cells. The resulting concentration-response data were analyzed using benchmark dose (BMD) modelling to quantify PPARα activation potency. A key finding was that PFAS with a sulfonic acid group showed a lower potency compared to those with a carboxylic group. The most potent activators belonged to the perfluoroalkylether carboxylic acid (PFECA) subgroup. Computational descriptors were generated to characterize each congener, and quantitative structure-activity relationship (QSAR) modelling was applied to relate molecular features to in vitro PPARα activation potency, as expressed by BMD estimates. For prioritization purposes in the context of PFAS hazard characterization, the QSAR model was used to screen about 10,000 PFAS congeners. Of these, roughly 10
Artificial Intelligence (AI) is increasingly influencing chemical risk assessment, enabling faster, more comprehensive, and potentially more ethical assessments. The application of AI in chemical risk assessment refers to both generative and predictive algorithms encompassing machine learning, to analyse complex chemical, biological, and environmental data and provide insights into adverse effect potential for humans and ecosystems. AI systems support the prediction of chemical hazards, exposure levels, and adverse effects by learning from experimental results, mechanistic models, and regulatory datasets, thereby enhancing the efficiency of safety evaluations.In October 2024, ECETOC held an international workshop, with experts from academia, industry, and regulatory bodies, to reflect upon the historical challenges in integrating multidimensional omics technologies into chemical regulation and explore the current capabilities and future potential of AI in toxicology and regulatory science. Discussions emphasised that implementation of Findable, Accessible, Interoperable, and Reusable (FAIR) data principles is not just a best practice but rather a prerequisite for building transparent, reliable, and unbiased AI systems. The reliability of AI in producing scientifically valid and socially responsible outcomes depends fundamentally on the availability of FAIR data. However, ensuring trustworthiness also requires robust governance frameworks that go beyond data and human oversight. Critical enablers of responsible AI in chemical risk assessment are rigorous governance, explainability, fit-for-purpose applications, and human oversight. ECETOC supports the development of flexible and iterative frameworks advancing development, validation, transparency, accountability, and trust in AI applications in chemicals regulation.
This paper presents a novel control framework that integrates Physics-Informed Neural Networks (PINNs) with Model Predictive Control (MPC) for nonlinear dynamical systems. Unlike traditional MPC, which requires solving optimization problems in real time, the proposed method trains a single feedforward neural network to serve as an explicit controller that directly maps the current state, set-point, and disturbance signals to optimal control actions. The network is trained using a composite loss function that enforces the governing differential equations while incorporating control-oriented objectives such as set-point tracking, control smoothness, and soft constraints on states, inputs, and outputs. The proposed controller is validated on both single-input single-output (SISO) and multi-input multi-output (MIMO) water-tank benchmark systems, demonstrating accurate set-point tracking, effective measured disturbance rejection, and strong generalization across thousands of randomized test scenarios. A runtime comparison with a nonlinear MPC performing online optimization confirms that the explicit PINN-MPC approach achieves comparable control performance while requiring several orders of magnitude less computation time. These results highlight the scalability and computational efficiency of the proposed framework, positioning it as a novel paradigm for real-time control of nonlinear systems.
Rapid innovation in chemicals and materials calls for innovative integrated approaches that can assess their impacts across different areas. The Safe and Sustainable-by-Design (SSbD) framework, developed by the European Commission's Joint Research Centre (JRC), offers a comprehensive approach with which to evaluate the safety and sustainability of chemicals and materials across their lifecycle. While SSbD uses various modeling approaches to assess impacts on human health, the environment, and socioeconomic factors, these are often applied independently, hindering a holistic understanding of the complex interactions between these factors and thus the simultaneous optimization of function, cost, safety and sustainability. This review describes existing predictive models and available strategies for their integration to facilitate more comprehensive and holistic chemical and material impact assessments. Specifically, we examine three model integration strategies: consensus integration that combines model predictions for the same impact categories, weighted aggregation that combines different scores in a unified one, and pipeline integration that links models sequentially to create a more unified assessment. Furthermore, we address key concepts related to the uncertainty of model predictions and the applicability domain of models, highlighting how these evolve in integrated frameworks. Insights into the applications of these integration strategies and challenges will allow a more accurate, coherent, and sustainable approach to chemical and material safety and sustainability assessments.
The production of nanomaterials (NMs) has gained significant attention due to their unique properties and versatile applications in fields such as medicine, energy, and electronics. However, ensuring the large-scale synthesis of safe and sustainable NMs while maintaining their functionality remains a critical challenge. This study introduces the Safety by Process Control (SbPC) framework, a novel methodology integrating dynamic first-principles modeling, Model Predictive Control (MPC), and real-time safety monitoring. The framework employs a physics-based population balance model with a Method Of Moments (MOM) approximation to predict the evolution of key NM properties. A toxicity inferential sensor, built on experimental data, is integrated to facilitate real-time hazard assessment. The efficiency of the proposed framework is demonstrated using a continuous silver nanoparticle (Ag NP) production system as a case study. The proposed approach ensures the production of high-quality, safe, and sustainable NMs, aligning with Safe and Sustainable by Design (SSbD) principles and addressing gaps in current NM manufacturing processes. The framework’s adaptability to other NM types highlights its potential as a transformative tool for sustainable nanotechnology.
Advances in drug discovery and material design rely heavily on in silico analysis of extensive compound datasets and accurate assessment of their properties and activities through computational methods. Efficient and reliable prediction of molecular properties is crucial for rational compound design in the chemical industry. To address this need, we have developed predictive models for nine key properties, including the octanol/water partition coefficient, water solubility, experimental hydration free energy in water, vapor pressure, boiling point, cytotoxicity, mutagenicity, blood–brain barrier permeability, and bioconcentration factor. These models have demonstrated high predictive accuracy and have undergone thorough validation in accordance with OECD test guidelines. The models are seamlessly integrated into the Enalos Cloud Platform through Titania ( https://enaloscloud.novamechanics.com/EnalosWebApps/titania/ ), a comprehensive web-based application designed to democratize access to advanced computational tools. Titania features an intuitive, user-friendly interface, allowing researchers, regardless of computational expertise, to easily employ models for property prediction of novel compounds. The platform enables informed decision-making and supports innovation in drug discovery and material design. We aspire for this tool to become a valuable resource for the scientific community, enhancing both the efficiency and accuracy of property and toxicity predictions.
Traditional in vivo methodologies have long formed the foundation of chemical and material safety assessment, yet they are increasingly inadequate to meet modern regulatory, ethical, and sustainability demands. These conventional approaches are resource-intensive, ethically questionable, and often fail to accurately predict human or environmental toxicity, particularly for emerging pollutants such as PFAS, (nano-) pesticides, and 2D materials. In response, the EU has launched initiatives like the Chemical Strategy for Sustainability and the Zero Pollution Action Plan under the European Green Deal to promote innovation in safer, and more sustainable chemicals. Central to this transformation is the Safe and Sustainable by Design (SSbD) framework, developed by the European Commission’s Joint Research Center, which provides structured methodologies and metrics to integrate safety and sustainability into material innovation from the earliest stages of design.Building on this vision, the CHIASMA project aims to advance Next generation Safety Assessment (NGSA) by developing innovative New Approach Methodologies (NAMs) that combine experimental, computational, and Life Cycle Assessment (LCA) tools. Focusing on key biological systems and exposure routes, CHIASMA integrates Artificial Intelligence (AI), Machine Learning (ML), and Knowledge Graph (KG) technologies to enhance data interoperability and predictive accuracy. By embedding FAIR data principles and aligning with Good Laboratory Practice (GLP) standards, CHIASMA promotes transparency and regulatory acceptance. Fully aligned with SSbD principles, CHIASMA establishes a digital, interoperable infrastructure for predictive safety evaluation, that leverage on state-of-art experimental New Approach Methodologies (NAMs) bridging critical data gaps and supporting the transition towards sustainable, science-driven, and ethically responsible chemical and material innovation in Europe and beyond.
The deployment of nanoparticles (NPs) for targeted drug delivery in vivo holds immense potential for enhancing therapeutic efficacy while minimizing systemic side effects. However, the complexity of biological environments, including the biological barriers that need to be crossed for effective systemic delivery, presents significant challenges in optimizing NP delivery. This study demonstrates how a simple compartmental model facilitates the simulation and analysis of NP-mediated drug delivery, supporting targeted delivery optimization. The model involves reversible transport between five compartments related to drug delivery (administration site, off-target sites, target cell vicinity, target cell interior and excreta) that determine NP dynamics, including biodistribution, degradation, and excretion processes. This approach enables the estimation of delivery efficiency and the identification of critical factors affecting NP delivery through sensitivity analysis. A case study involving PEG-coated gold NPs delivered intravenously to the lungs demonstrates the model's capacity to describe observed biodistribution patterns and highlights key parameters influencing delivery outcomes. The model is exposed as a web application that provides a user-friendly graphical interface, enabling researchers to conduct in silico experiments with the goal of optimizing delivery strategies, thereby accelerating the development of precision nanomedicine. The model is made available both as a web application, via the Enalos Cloud Platform, and as a RESTful aaplication programming interface (API), providing a user-friendly graphical interface and programmatic access, respectively, enabling researchers to integrate the model into their own computational workflows. This study illustrates how simple compartmental modelling can be employed to guide the development of targeted drug delivery systems, contributing to more effective and personalized healthcare interventions.
In this work, we introduce a nonlinear economic-oriented model predictive control framework that can optimize the economic operation of wastewater treatment plants (WWTPs), while accounting for inlet flow disturbances. The proposed method utilizes an attention-based recurrent neural network (RNN) model to predict influent flow rate variations, and a WWTP reduced-order model specifically tailored for MPC integration. At each sampling instant, the proposed scheme recursively solves an optimal control problem, where the objective is to minimize the plant energy consumption. The inlet flow rate RNN predictions are integrated within the scheme and critical controller parameters, such as the prediction horizon, are optimized by considering the best RNN multi-step ahead prediction horizon. The proposed framework is applied to a modified benchmark simulation model no. 1 (BSM1) representation that corresponds to an actual WWTP and its performance is compared against different control schemes, outperforming the alternative methods in terms of optimizing WWTP performance.
The assessment of chemicals and materials has traditionally been fragmented, with health, environmental, social, and economic impacts evaluated independently. This disjointed approach limits the ability to capture trade-offs and synergies necessary for comprehensive decision-making under the Safe and Sustainable by Design (SSbD) framework. The EU INSIGHT project addresses this challenge by developing a novel computational framework for integrated impact assessment, based on the Impact Outcome Pathway (IOP) approach. Extending the Adverse Outcome Pathway (AOP) concept, IOPs establish mechanistic links between chemical and material properties and their environmental, health, and socio-economic consequences. The project integrates multi-source datasets (including omics, life cycle inventories, and exposure models) into a structured knowledge graph (KG), ensuring FAIR (Findable, Accessible, Interoperable, Reusable) data principles are met. INSIGHT is being developed and validated through four case studies targeting per- and polyfluoroalkyl substances (PFAS), graphene oxide (GO), bio-based synthetic amorphous silica (SAS), and antimicrobial coatings. These studies demonstrate how multi-model simulations, decision-support tools, and artificial intelligence-driven knowledge extraction can enhance the predictability and interpretability of chemical and material impacts. Additionally, INSIGHT incorporates interactive, web-based decision maps to provide stakeholders with accessible, regulatory-compliant risk and sustainability assessments. By bridging mechanistic toxicology, exposure modeling, life cycle assessment, and socio-economic analysis, INSIGHT advances a scalable, transparent, and data-driven approach to SSbD. This project aligns with the European Green Deal and global sustainability goals, promoting safer, more sustainable innovation in chemicals and materials through an integrated, mechanistic, and computationally advanced framework.
Current EU Strategies aim to rapidly advance the research, development and deployment of innovative advanced materials and chemicals to make Europe the first digitally enabled circular, climate-neutral and sustainable economy. To achieve this, an underlying adaptation of the research and innovation (R&I) process to the Safety-and-Sustainability-by-Design (SSbD) framework has been proposed. This perspective article provides an overview of already existing approaches providing guidance for implementing SSbD-like procedures in R&I in several different industrial sectors to ultimately replace substances of concern (SoC). Starting from the ECHA's Assessment of Alternatives (AoA) approach we put emphasis on the scoping phase during which the requirements for replacement will be identified. The limitations for the changes possible and trade-offs acceptable for the company need to be defined, in agreement with relevant stakeholders to be further involved in AoA scoping (e.g. for setting the trade-off levels). This includes listing the SSbD-relevant aspects in the different categories (i.e. functional performance, health, environment, social, and economic sustainability) in a customized manner, followed by weighting them in relation to their expected impact on the intended SSbD-guided multi-objective optimization procedure. An additional dimension is provided as to how to deal with uncertainties (e.g. data gaps or compromises in data quality, or which assessment methods and tools to employ); notably, it represents the company's own decision to herewith set the requirements and goals for replacement, and this can be done at different levels, such as the material or chemical itself, changes in the production processes, or within the entire system of a product's life cycle spanning across its entire value chain(s), which can be documented employing the use maps concept. Further, this article builds on the product life cycle and provides a general understanding of life cycle assessment (LCA) methodology, especially a deeper insight into prospective and anticipatory LCA, that will need to prove functional on real-life case studies from industry. Besides a clarification of these concepts, the article provides an interdisciplinary view, as required for implementing SSbD in small and medium-sized enterprises, with hints on the use of machine learning techniques for anticipatory LCA of new chemicals, materials, and products. Such methodologies will, in future, help extend classical LCA cases towards the data-scarce requirements of earlier material and product development stages.
The extensive conformational dynamics of partially disordered proteins hinders the efficiency of traditional in-silico structure-based drug discovery approaches due to the challenge of screening large chemical spaces of compounds, albeit with an excessive number of transient binding sites, quickly making this problem intractable. In this study, using the monomer of the AR-V7 transcription factor splicing variant related to prostate cancer as a test case, we present a deep ensemble docking pipeline that accelerates the screening of small molecule binders targeting partially disordered proteins at functional regions. By swiftly identifying the conformational ensemble of AR-V7 and reducing the dimension of binding sites by a factor of 90, we identify functionally relevant binding sites along the AR-V7 structural ensemble at phase separation-prone regions that have been experimentally shown to contribute to enhanced transcription activity and the onset of tumor growth. Following this, we combine physics-based molecular docking and multiobjective classification machine learning models to speed up the screening for binders in a larger chemical space able to target these functional multiple binding sites of AR-V7. This step increases the multibinding site hit rate of small molecules by a factor of 17 compared to naive molecular docking. Finally, assessing in atomistic molecular dynamics the effect of a selected binder on AR-V7 dynamics, we find that in the presence of the ChEMBL22003 compound, AR-V7 exhibits less conformational entropy, smaller solvent exposure of phase separation-prone regions, and higher solvent exposure of other protein regions, promoting this compound as a potential AR-V7 phase separation modulator.
In this study, we introduce a novel approach for predicting two key drug properties, blood–brain barrier (BBB) permeability and human intestinal absorption via Caco-2 permeability. Our methodology centers around a specialized neural network, the atom transformer-based Message Passing Neural Network (MPNN), which we have combined with contrastive learning techniques to enhance the process of representing and embedding molecular structures for more accurate property prediction. These innovative models focus on predicting BBB and Caco-2 permeability -two critical factors in drug absorption and distribution- which fall under the broader scope of ADMET (absorption, distribution, metabolism, excretion, and toxicity) properties. The models are readily accessible online through the Enalos Cloud Platform which offers a user-friendly, AI-powered, ready-to-use web service that significantly streamlines the drug design process, enabling users to easily predict and understand the behavior of potential drug compounds within the human body. Scientific Contribution Our study combines an atom-attention Message Passing Neural Network (AA-MPNN) with contrastive learning (CL), which significantly improves predictive accuracy. Our model leverages self-supervised learning to expand the chemical space used in training and self-attention mechanisms to focus on critical molecular features, enhancing both model accuracy and interpretability. Additionally, the ready-to-use web service based on our model democratizes access to predictive tools for the scientific and regulatory communities.
In this innovation report, we present the vision of the PINK project to foster Safe-and-Sustainable-by-Design (SSbD) advanced materials and chemicals (AdMas&Chems) development by integrating state-of-the-art computational modelling, simulation tools and data resources. PINK proposes a novel approach for the use of the SSbD Framework, whose innovative approach is based on the application of a multi-objective optimisation procedure for the criteria of functionality, safety, sustainability and cost efficiency. At the core is the PINK open innovation platform, a distributed system that integrates all relevant modelling resources enriched with advanced data visualisation and an AI-driven decision support system. Data and modelling tools from the, in large parts, independently developed areas of functional design, safety assessment, life cycle assessment & costing are brought together based on a newly created Interoperability Framework. The PINK In Silico Hub, as the user Interface to the platform, finally guides the user through the complete AdMas&Chems development process from idea creation to market introduction. Guided by two Developmental Case Studies, the process of building of the PINK Platform is iterative, ensuring industry readiness to implement and apply it. Additionally, the Industrial Demonstrator programme will be introduced as part of the final project phase, which allows industry partners and especially small and medium enterprises (SMEs) to become part of the PINK consortium. Feedback from the Demonstrators as well as other stakeholder-engagement activities and collaborations will shape the platform's final look and feel and, even more important, activities to assure long-term technical sustainability.
Partially disordered proteins can contain both stable and unstable secondary structure segments and are involved in various (mis)functions in the cell. The extensive conformational dynamics of partially disordered proteins scaling with extent of disorder and length of the protein hampers the efficiency of traditional experimental and in-silico structure-based drug discovery approaches. Therefore new efficient paradigms in drug discovery taking into account conformational ensembles of proteins need to emerge. In this study, using as a test case the AR-V7 transcription factor splicing variant related to prostate cancer, we present an automated methodology that can accelerate the screening of small molecule binders targeting partially disordered proteins. By swiftly identifying the conformational ensemble of AR-V7, and reducing the dimension of binding-sites by a factor of 90 by applying appropriate physicochemical filters, we combine physics based molecular docking and multi-objective classification machine learning models that speed up the screening of thousands of compounds targeting AR-V7 multiple binding sites. Our method not only identifies previously known binding sites of AR-V7, but also discovers new ones, as well as increases the multi-binding site hit-rate of small molecules by a factor of 10 compared to naive physics-based molecular docking.### Competing Interest StatementThe authors have declared no competing interest.
Modelling Data (MODA) reporting guidelines have been proposed for common terminology and for recording metadata for physics-based materials modelling and simulations in a CEN Workshop Agreement (CWA 17284:2018). Their purpose is similar to that of the Quantitative Structure-Activity Relationship (QSAR) model report form (QMRF) that aims to increase industry and regulatory confidence in QSAR models, but for a wider range of model types. Recently, the WorldFAIR project’s nanomaterials case study suggested that both QMRF and MODA templates are an important means to enhance compliance of nanoinformatics models, and their underpinning datasets, with the FAIR principles (Findable, Accessible, Interoperable, Reusable). Despite the advances in computational modelling of materials properties and phenomena, regulatory uptake of predictive models has been slow. This is, in part, due to concerns about lack of validation of complex models and lack of documentation of scientific simulations. The models are often complex, output can be hardware- and software-dependent, and there is a lack of shared standards. Despite advocating for standardised and transparent documentation of simulation protocols through its templates, the MODA guidelines are rarely used in practice by modellers because of a lack of tools for automating their creation, sharing, and storage. They also suffer from a paucity of user guidance on their use to document different types of models and systems. Such tools exist for the more well-established QMRF and have aided widespread implementation of QMRFs. To address this gap, a simplified procedure and online tool, Easy-MODA, has been developed to guide users through MODA creation for physics-based and data-based models, and their various combinations. Easy-MODA is available as a web-tool on the Enalos Cloud Platform (https://www.enaloscloud.novamechanics.com/insight/moda/). The tool streamlines the creation of detailed MODA documentation, even for complex multi-model workflows, and facilitates the registration of MODA workflows and documentation in a database, thereby increasing their Findability and thus Re-usability. This enhances communication, interoperability, and reproducibility in multiscale materials modelling and improves trust in the models through improved documentation. The use of the Easy-MODA tool is exemplified by a case study for nanotoxicity evaluation, involving interlinked models and data transformation, to demonstrate the effectiveness of the tool in integrating complex computational methodologies and its significant role in improving the FAIRness of scientific simulations.
Wastewater treatment plants (WWTPs) employ a series of complex chemical and biological processes, to transform an influent stream of contaminated water to an effluent suitable for return to the water cycle. To optimize the performance of WWTP control schemes, appropriate mathematical models capable of accurately simulating the plant dynamic behavior are essential. However, the development of reliable dynamic representations for these large-scale plants is challenging, mainly because of the complex biological reactions taking place and the significant fluctuations in the disturbances that affect the operation of WWTPs. First-principles models, such as the well-known benchmark simulation model no. 1 (BSM1), may be capable of capturing the highly nonlinear nature of WWTPs, but this comes at the cost of employing complex, high-order representations of the reactive units and settling processes. This complexity leads to highly complicated configurations that cannot be efficiently integrated in advanced process control schemes, like model predictive controllers (MPCs). Furthermore, the large number of unknown parameters in these models, along with the non-convex nature of the underlying functions, renders the use of conventional system identification techniques insufficient. To remedy these issues, in this work we introduce a reduced-order first-principles model for WWTPs, incorporating low order mathematical models for the chemical phenomena of the reactive units and the settling procedure. Furthermore, we present a novel system identification scheme, which is based on a customized cooperative particle swarm optimization approach; the scheme effectively handles the high-dimensionality and multimodality of the underlying nonlinear optimization problem, enabling accurate estimation of the model parameters. Comparison results between the dynamic behavior of the original BSM1 and the identified reduced-order model, indicate that the proposed approach is capable of accurately and robustly capturing the highly nonlinear nature of WWTPs, while being simple enough for incorporation in the design of MPC and other advanced control schemes. This represents a significant advancement over traditional models, offering a more practical and efficient approach for WWTP management and control.