In this work, we introduce a practical hybrid methodology that fuses first-principles knowledge with measured process data to improve model prediction and extrapolation. Using sequential-orthogonalized partial least squares (SO-PLS), we (i) remove variation already explained by physics so the data-driven component learns only from unexplained residual structures, and (ii) decompose each prediction into mechanistic and empirical contributions. This decomposition yields a simple diagnostic - the mechanistic-to-measured (M2M) ratio - that quantifies the relative influence of physics versus data and flags when the fit is dominated by empirical corrections, signaling potential degradation under extrapolation. We demonstrate the methodology on a simulated two-stage batch reactor with both well-specified and misspecified kinetics. Even with model misspecification, the mechanistic component adds useful information; however, the contribution analysis shows substantial empirical adjustment remains necessary. The M2M ratio enables practitioners to compare candidate models and select those more likely to remain reliable beyond the training domain.
This work presents a methodology for evaluating the effectiveness of hybrid modelling under varying conditions of mechanistic model quality and information available for model training, that is typically expressed as the amount of data available. While hybrid models – which integrate mechanistic and data-driven components – have gained significant attention in process systems engineering, their advantages over purely mechanistic or data-driven alternatives remain inadequately quantified. We address this gap by investigating two critical factors: (i) the impact of mechanistic model fidelity on hybrid model performance, and (ii) the influence of calibration dataset size on prediction accuracy. Our methodology is validated through an in-silico case study of baker's yeast cultivation and a real-world industrial application of ion-exchange chromatography in biopharmaceutical manufacturing. Results demonstrate that hybrid models consistently outperform purely mechanistic and data-driven approaches when the mechanistic component captures fundamental process behaviours, even with structural simplifications. Notably, hybrid models maintain superior predictive capability in extrapolative scenarios; however, when mechanistic knowledge is severely limited and insufficient information is available for compensation, hybridisation benefits diminish substantially. The work provides quantitative guidance for practitioners to determine when hybrid modelling represents a justified investment of modelling resources in process engineering applications.
Cement is regarded as the most widely used construction material worldwide; however, its production is also recognized as a major contributor to global CO2 emissions. Strict control of cement quality is therefore required to prevent excessive consumption of raw materials and energy, which would otherwise increase the process environmental footprint. Cement quality is largely governed by clinker quality, which is primarily characterized by two quality control parameters: free-lime content and alite fraction. At present, these are characterized by costly and time-consuming laboratory analyses that are not optimal for real time process control and optimization. Hence, in this work, a soft sensor for the real-time estimation of the clinker alite fraction is proposed. The developed soft sensor is designed to adapt to process drifts and operating condition changes, capture nonlinear and dynamic behavior, and retain interpretability through a Partial Least Squares (PLS) modelling framework. To this end, a novel recursively adaptive local dynamic soft-sensing strategy (ALD-PLS) is introduced and implemented within a multi-model ensemble structure referred to as Quasi-Ensemble PLS (QE-PLS). Unlike conventional ensemble approaches, where model diversity is generated through data resampling or training–testing partitioning, the proposed framework constructs multiple sub-models using combinations of model hyperparameters, evaluated on the same evolving dataset. As a result, improved predictive accuracy and robustness are achieved, while estimation uncertainty is quantified. The proposed QE-PLS soft sensor is shown to outperform, in terms of R2 and RMSE, PLS-based and single-instance implementations ALD-PLS for a similar task. The novel methodology is validated against data collected from an industrial cement production plant.
Lutein, a carotenoid involved in light harvesting and photoprotection, has demonstrated significant therapeutic benefits, particularly for ocular health. This study aims to optimize biomass and lutein production from an extremophilic microalga Coccomyxa onubensis in a fed-batch system. C. onubensis was cultivated in a membrane high-cell-density photobioreactor. Design of Dynamic Experiments was employed to optimize dynamic cultivation factors, including time-varying light intensity profiles (80-800 mu mol center dot m- 2 center dot s-1), nitrogen feeding strategies (60-600 mgN center dot L-1), and three constant COQ levels (0.04 %, 1 %, and 10 %). Biomass productivity peaked at 1.28 +/- 0.23 g center dot L- 1 center dot d-1 under 6.5 % COQ, with gradual light increase (peak 800 mu mol photons center dot m- 2 center dot s-1), and moderately high nitrogen (499 mg center dot L- 1 over 5 days). Conversely, lutein productivity reached 3.32 +/- 0.58 mg center dot L- 1 center dot d-1 under moderate light intensity (peak 414 mu mol photons center dot m- 2 center dot s-1) and limited nitrogen (217 mgN center dot L- 1 over 1 day) under 10 % COQ. This work provides optimized and validated cultivation strategies for impactful time varying factors, implementable at industrial level. DoDE proved to be valuable tool for multifactorial optimization and interpretation of complex interactions, significantly enhancing biomass and lutein production under acidophilic, high cell density conditions where bacterial contamination can be minimized.
Genome-scale metabolic models (GEMs) have been widely utilized to understand cellular metabolism. The application of GEMs has been advanced by computational methods that enable the prediction and analysis of intracellular metabolic states. However, the accuracy and biological relevance of these predictions often suffer from the many degrees of freedom and scarcity of available data to constrain the models adequately. Here, we introduce Neural-net EXtracellular Trained Flux Balance Analysis, (NEXT-FBA), a novel computational methodology that addresses these limitations by utilizing exometabolomic data to derive biologically relevant constraints for intracellular fluxes in GEMs. We achieve this by training artificial neural networks (ANNs) with exometabolomic data from Chinese hamster ovary (CHO) cells and correlating it with 13C-labeled intracellular fluxomic data. By capturing the underlying relationships between exometabolomics and cell metabolism, NEXT-FBA predicts upper and lower bounds for intracellular reaction fluxes to constrain GEMs. We demonstrate the efficacy of NEXT-FBA across several validation experiments, where it outperforms existing methods in predicting intracellular flux distributions that align closely with experimental observations. Furthermore, a case study demonstrates how NEXT-FBA can guide bioprocess optimization by identifying key metabolic shifts and refining flux predictions to yield actionable process and metabolic engineering targets. Overall, NEXT-FBA aims to improve the accuracy and biological relevance of intracellular flux predictions in metabolic modelling, with minimal input data requirements for pre-trained models.
The concepts of null space and orthogonal space have been developed in independent contexts and with different purposes: the former arises in the inversion of partial least-squares (PLS) regression models, and the latter in orthogonal PLS (O-PLS) modeling. In this study, we bridge PLS model inversion and O-PLS modeling by mathematically proving that the null space and the orthogonal space are the same space. We also provide a graphical interpretation of the equivalence between the two spaces, using both a simulated and a real case study.
The identification of highly productive cell lines is crucial in the development of bioprocesses for the production of therapeutic monoclonal antibodies (mAbs). Metabolomics data provide valuable information for cell line selection and allow the study of the relationship with mAb productivity and product quality attributes. We propose a novel robust machine learning procedure which, exploiting dynamic metabolomic data from the Ambr (R) 15 scale, supports the selection of highly productive cell lines during biopharmaceutical bioprocess development and scale-up. The metabolomic profiles dynamics allows to identify the cell lines with high productivity, already in the early stages of experimentation, and the biomarkers that are the most related to mAb productivity, finding at the same time the key metabolic pathways for discriminating mAb productivity. Specifically, tricarboxylic acid cycle pathways are predominant in the early stages of the cultivation, while amino and nucleotide sugar pathways influence in the late stages of the culture.
Batch process monitoring using principal component analysis requires sufficient historical manufacturing data to model the normal operating conditions of the process. However, when a new product is to be manufactured for the first time in a given facility, very limited historical data are available, thus entailing a small-data scenario. We thoroughly investigate and improve a data-driven methodology, previously reported in the literature (Tulsyan, Garvin & undey (2019). J. Process Control, 77, 114-133), that enables batch process monitoring under such type of scenarios. The methodology exploits machine learning algorithms (based on Gaussian process state-space models) to generate in-silico batch trajectory data from the few available historical ones, and then uses the overall pool of real and in-silico data to build a process monitoring model. We develop automatic procedures to tune the values of several parameters of this machine-learning framework, in such a way that the generation of consistent in-silico batch trajectory data can be streamlined, thus facilitating the deployment of the framework at an industrial level. Furthermore, we develop indicators and a metric to assist the in-silico data generation activity from a process monitoring-relevant perspective. Finally, using datasets from a benchmark simulated semi-batch process for the manufacturing of penicillin, we thoroughly investigate the appropriateness of the in-silico generated data for the purpose of process monitoring.
Membrane separation processes are precious assets for biorefineries to separate biomass from the solution containing the product after bioconversion in an effective and energy-efficient way. However, fouling can significantly reduce the benefits of membrane separations. Effects of fouling can be reversible, manifesting as short-term process disruption, or irreversible, causing long-term membrane degradation; the two actions typically affect one another. Understanding potential causes of membrane fouling is of paramount importance to mitigate this undesired phenomenon and improve process operation. In this study, we perform a comprehensive investigation of membrane fouling in the ultrafiltration operation of the world's first industrial-scale biorefinery manufacturing 1,4-biobutanediol via bioconversion of renewable raw materials. We use principal component analysis to extract information from sensor data spanning six months of plant operation. Furthermore, we resort to feature-oriented data-driven modeling to address the variability of batch duration, and we exploit process knowledge to enhance information on the effects of fouling. We show how this approach can provide valuable information on the effectiveness of the cleaning and control policies adopted by plant operators, and offer guidelines on how to improve the membrane maintenance schedule. We also resort to engineering judgment for model interpretation in order to identify potential causes of fouling, uncover a strong interaction between reversible and irreversible fouling, and plan experimental investigations to clarify some of the detected effects and assess new ones.
The development of cell cultures to produce monoclonal antibodies is a multi-step, time-consuming, and labor-intensive procedure which usually lasts several years and requires heavy investment by biopharmaceutical companies. One key aspect of process optimization is improving the feeding strategy. This step is typically performed though design of experiments (DoE) during process development, in such a way as to identify the optimal combinations of factors which maximize the productivity of the cell cultures. However, DoE is not suitable for time-varying factor profiles because it requires a large number of experimental runs which can last several weeks and cost tens of thousands of dollars. We here suggest a methodology to optimize the feeding schedule of mammalian cell cultures by virtualizing part of the experimental campaign on a hybrid digital model of the process to accelerate experimentation and reduce experimental burden. The proposed methodology couples design of dynamic experiments (DoDE) with a hybrid semi-parametric digital model. In particular, DoDE is used to design optimal experiments with time-varying factor profiles, whose experimental data are then utilized to train the hybrid model. This will identify the optimal time profiles of glucose and glutamine for maximizing the antibody titer in the culture despite the limited number of experiments performed on the process. As a proof-of-concept, the proposed methodology is applied on a simulated process to produce monoclonal antibodies at a 1-L shake flask scale, and the results are compared with an experimental campaign based on DoDE and response surface modeling. The hybrid digital model requires an extremely limited number of experiments (nine) to be accurately trained, resulting in a promising solution for performing in silico experimental campaigns. The proposed optimization strategy provides a 34.9% increase in the antibody titer with respect to the training data and a 2.8% higher antibody titer than the optimal results of two DoDE-based experimental campaigns comprising different numbers of experiments (i.e., 9 and 31), achieving a high antibody titer (3,222.8 mg/L) —very close to the real process optimum (3,228.8 mg/L).
In the development of data-driven soft sensors for product quality assessment in multi-unit manufacturing processes, the only information that is typically used as an input to the model is real-time measurements from field sensors. However, even if detailed knowledge of the mechanistic behavior of the process may not be available, information about the sequence of processing units, and their connectivity, is available, typically in graphical form through process flow diagrams. In this study, we investigate the use of sequential-orthogonalized partial least-squares (SO-PLS) regression as a way to capture connectivity information from a process flow diagram, and transfer it into a data-driven model to be used as a soft sensor in a multi-unit process. Connectivity between units is captured and translated into a block order that establishes a sequence for block regressions. Orthogonalization between two blocks is then carried out with the aim of eliminating overlapping data and retaining information that is unique to each block. Product quality is finally predicted by summing the contributions from each block, and the accuracy of prediction is enhanced due to the embedded dual feature-extraction procedure, which combines orthogonalization and latent-variable extraction. The effectiveness of the proposed approach is illustrated by comparing the quality prediction performance of two soft sensors for a simulated multi-unit continuous process: one using standard PLS and one using SO-PLS. Superior performance of the SO-PLS soft sensor is achieved, even more markedly so when fewer field measurements are available to build the soft sensor.
Data-driven modeling has significantly transformed problem-solving in the process industry, especially in designing new products digitally by finding the process conditions that are required to manufacture a product with assigned quality. This can be achieved by utilizing historical process data via latent-variable model inversion, notably through the extensively used partial least-squares (PLS) regression model. Despite the development of numerous PLS model-inversion techniques, from straightforward algebraic manipulation of the model equations to the formulation and solution of complex nonlinear optimization problems, there lacks a comprehensive discussion on their comparative benefits and limitations. This paper offers a systematic analysis of PLS model inversion strategies, especially those based on optimization problems. We delve into aspects such as optimization in the latent or input variable spaces, the nature of constraints (soft vs. hard), and the feasibility of analytical solutions. We outline a clear hierarchical structure of the available methods based on the successive inclusion of constraints, and propose a general formulation of the PLS model inversion by optimization problem that encompasses all available methods. We support our theoretical analysis with a numerical case study, and provide the code to reproduce it and to solve general latent-variable model inversion problems according to the proposed formulation.
Membrane filtration is commonly used in biorefineries to separate cells from fermentation broths containing the desired products. However, membrane fouling can cause short-term process disruption and long-term membrane degradation. The evolution of membrane resistance over time can be monitored to track fouling, but this calls for adequate sensors in the plant. This requirement might not be fulfilled even in modern biorefineries, especially when multiple, tightly interconnected membrane modules are used. Therefore, characterization of fouling in industrial facilities remains a challenge. In this study, we propose a hybrid modeling strategy to characterize both reversible and irreversible fouling in multi-module biorefinery membrane separation systems. We couple a linear data-driven model, to provide high-frequency estimates of trans-membrane pressures from the available measurements, with a simple nonlinear knowledge-driven model, to compute the resistances of the individual membrane modules. We test the proposed strategy using real data from the world's first industrial biorefinery manufacturing 1,4-bio-butanediol via fermentation of renewable raw materials. We show how monitoring of individual resistances, even when done by simple visual inspection, offers valuable insight on the reversible and irreversible fouling state of the membranes. We also discuss the advantage of the proposed approach, over monitoring trans-membrane pressures and permeate fluxes, from the standpoints of data variability, effect of process changes, interaction between module in multi-module systems, and fouling dynamics.
Roller compaction is a key unit operation in a dry granulation line for pharmaceutical tablet manufacturing. During product development, one would like to find the roller compactor (RC) settings that are required to achieve a desired ribbon solid fraction. These settings can be determined from the compression profile of the powder mixture being compacted and a mathematical model that interprets it. However, establishing compression profiles in an RC requires relatively large amounts of powder, which are expensive and may not be available during drug development. As a cost-effective alternative to an RC, a compactor simulator (CS) can be used, which is a small-scale equipment that uses minimal amounts of powder to build the compression profile. However, since the working principles of a CS and an RC are different, the compression profiles obtained from the two devices for a given powder are also different. In this study, we propose a transfer learning approach that allows the RC compression profile of a given powder to be easily predicted from the compression profile obtained in a CS for the same powder. Based on the well-known Johanson model and on the mass correction factor theory, we examine the compaction behavior of six formulations, two of which including active ingredients, and we find that the mass correction factor does not depend significantly on the powder being compacted. We develop a simple, generalized correlation (transfer model) that allows the mass correction factor to be predicted solely as a function of the pressure at which the compaction is carried out. By using the proposed transfer model, the prediction of the RC compression profiles for the validation powders is significantly improved over the case where a constant value of the mass correction factor is used.
A novel approach based on supervised machine -learning is proposed to predict the solubility of drugs and druglike molecules in mixtures of organic solvents. Similar to quantitative structure -property relationship (QSPR) models, different solvent types are identified by molecular descriptors, which, in this study, are considered as UNIFAC subgroups. To overcome the potential lack of UNIFAC subgroups for the complex Active Pharmaceutical Ingredients (APIs) currently developed in the pharmaceutical industry, the API molecule is considered as a unique entity in the proposed modelling approach. Therefore, API solubility is predicted as a function of temperature, functional subgroups of the solvents and composition of the solvent mixture; in turn, regressors ' correlation is handled through Partial Least -Squares (PLS) regression. The method is developed and tested with experimental data of a real API and 14 organic solvents that are industrially employed for crystallisation. Solubility predictions are accurate and precise for single solvents, binary mixtures and ternary mixtures of organic solvents at different compositions and temperatures, with a determination coefficient R 2 >= 0.90. To further test the applicability of the model, the proposed approach is applied to 9 literature organic solubility datasets of drugs and drug -like compounds and compared to benchmark solubility models in the literature. Results show that the proposed approach provides satisfactory predictions: the majority of validation and calibration data have R 2 = 0.95 -0.99; the ratio between RMSE (root mean squared error) of the proposed method and the range of measured solubility values is from 1 to 3 orders of magnitude smaller than the RMSE ratio obtained by the benchmark models.
The use of computationally demanding knowledge-driven models to optimize a process might encounter substantial numerical challenges. Because a model is an abstraction and approximation of the process, calculating the exact model optimum might not be necessary because its industrial implementation is bound to be an approximate one. Here we are exploring an alternative optimization route through a surrogate model. Because one of the decision variables affecting the optimization is time-varying, the Design of Dynamic Experiments is used to estimate the surrogate model. The process considered here is a freeze-drying process widely used in the pharmaceutical industry. The model used is a stochastic model describing the process in great detail. It is shown that the proposed data-driven route calculates the optimum in about 8 h, as opposed to 22 h for the knowledge-driven model, while sacrificing only < 15% in the computed value of the process performance.
The management of trade-off between experimental design space exploration and information maximization is still an open question in the field of optimal experimental design. In classical optimal experimental design methods, the uncertainty of model prediction throughout the design space is not always assessed after parameter identification and parameters precision maximization do not guarantee that the model prediction variance is minimized in the whole domain of model utilization. To tackle these issues, we propose a novel model-based design of experiments (MBDoE) method that enhances space exploration and reduces model prediction uncertainty by using a mapping of model prediction variance (G-optimality mapping). This explorative MBDoE (eMBDoE) named G-map eMBDoE is tested on two models of increasing complexity and compared against conventional factorial design of experiments, Latin Hypercube (LH) sampling and MBDoE methods. The results show that G-map eMBDoE is more efficient in exploring the experimental design space when compared to a standard MBDoE and outperforms classical design of experiments methods in terms of model prediction uncertainty reduction and parameters precision maximization.