We present an extended Bayesian Transfer Learning framework that integrates heterogeneous data sources for developing robust kinetic models in industrial catalyst development. The methodology incorporates hard and soft constraints into Monte Carlo Markov Chain (MCMC) sampling, enabling multi-objective optimization with arbitrary model structures. Applied to continuous lumping models of zeolite hydrocracking catalysts, our approach strategically combines absolute measurements from pilot plant data with relative performance rankings from high-throughput screening; information that cannot be directly included due to scale-up limitations. By transferring information through prior distributions and ranking constraints, we obtain a set of coherent models that respect both absolute performance levels and relative activity and selectivity trends. Validation on unseen stacking tests confirms predictions. The impact of ranking constraints depends on design space overlap: they prove essential when catalyst generations explore different operating regimes but less critical when conditions overlap substantially. This enables efficient experimental design where costly pilot campaigns explore new conditions while inexpensive screening maintains portfolio consistency. The Bayesian framework provides rigorous uncertainty quantification. This methodology addresses a fundamental industrial challenge: optimally leveraging diverse information from experimental programs spanning different scales, costs, and information content.
This work presents a methodology for kinetic model parameter fitting with Transfer Learning to improve model robustness and mitigate over-fitting of poorly sensitized parameters. Datasets from lab, pilot plant, or industrial scale experiments for the same process often cover different regions of the design space. Therefore, some of the model’s features lack sufficient variations to obtain a robust estimate of the model parameters from one of the datasets alone. Combining such, heterogeneous, datasets is a challenging task and requires extensive expert knowledge. Transfer Learning with Monte Carlo Markov Chains (MCMC) is used to retain information from different datasets. This method uses the Bayes theorem to impose a prior distribution on the model parameters when sampling the likelihood distribution via repeated model inference for random parameter variation. The aim of the methodology is to leverage the respective strengths and mitigate the weaknesses of different datasets, which will lead to a more robust estimation of the kinetic model parameters.In this work we use the MCMC algorithm for parameter identification for a hydrodenitrogenation (HDN) model used to simulate the pretreatment reactor of hydrocracking units. Industrial follow-up and pilot plant data was used in this study. A deactivation model was developed to account for loss of catalyst activity during long cycles of industrial units. The kinetic parameters are well sensitized in the pilot plant dataset, but feed variability is low. The industrial dataset has more variation of feedstock descriptors but temperature, pressure, and LHSV effects are poorly sensitized. The comparison to a naive approach of identifying selected subsets of the parameters on each dataset shows that the Transfer Learning method provides less over-fitting by retaining information from both datasets.
Mechanistic models of catalytic reactors are robust but often fail to generalize across feedstocks and operating conditions due to incomplete kinetic representations and parameter uncertainty, while purely data-driven models lack physical consistency and require large datasets that are costly to obtain in process engineering. To address these challenges, this work proposes a physics-informed hybrid framework that integrates a kinetic ordinary differential equation (ODE) into a neural network through automatic differentiation, differential residual loss minimization, and physical boundary constraints, taking both operating conditions and feedstock descriptors as inputs. The approach is applied to industrial hydrotreating (HDT), focusing on the prediction of nitrogen slip concentration in vacuum gas oil (VGO) hydrodenitrogenation (HDN).Two methodological developments are proposed: a Delaunay-triangulation-based strategy to select the input locations at which the physical residual and boundary terms are enforced during hybrid training, and a feedstock oriented evaluation across 30 randomized splits in which entire feedstocks are withheld from training to assess extrapolation. The hybrid model consistently outperforms both the mechanistic and purely data-driven baselines, achieving lower RMSE, MAE, and temperature deviations, while trend analyses confirm stronger preservation of the underlying physical dependencies, even under extrapolation to unseen feedstocks. A systematic comparison using a simplified kinetic formulation further shows that the hybrid architecture remains effective when constrained by incomplete mechanistic knowledge, outperforming data-driven models and approaching the accuracy of the complete kinetic expression.These results highlight the robustness, flexibility, and practical relevance of physics-informed hybrid modeling for reactor systems, offering a promising pathway toward reliable and quickly developed digital twins in catalytic process engineering.
Accurately predicting the nitrogen content of co-processed petroleum feedstocks after hydrotreatment is a common challenge in refining, exacerbated by the limited experimental data available for each new feedstock. Transfer learning addresses this scarcity by leveraging knowledge from a related, data-rich source to improve predictions on a data-scarce target. For the source, the hydrodenitrogenation (HDN) of fossil vacuum gas oil, a rich dataset and an accurate prior governing ordinary differential equation (ODE) are available; for the target, the HDN of co-processed vacuum gas oil and tire-pyrolysis oil, the ODE structure is unknown and an additional descriptor (the tire fraction) is present, making the problem heterogeneous. To recover this unknown ODE, we propose HTL-HODE, a heterogeneous transfer learning hybrid ODE in which small neural networks encode its unknown components while the rest is kept explicit and transferred from the source. For comparison, we also develop physics-informed variants (HTL-PINNs) of data-driven heterogeneous transfer networks. On simulated data, HTL-HODE reduces the test mean absolute error (MAE) by up to 52% relative to a no-transfer baseline; on experimental data, the reduction reaches up to 60% in the best case. This suggests that the hybrid form captures the underlying HDN kinetics, making it a promising tool for uncovering the governing ODE of unknown systems, in a hybrid sense. The HTL-PINN variants did not match HTL-HODE's accuracy. Overall, when a kinetic structure is available, embedding it explicitly in a hybrid model is more effective than a black-box network penalised by a physics residual.
Classical Fault Detection and Diagnosis (FDD) methods, including many data-driven approaches, assume a static normal operating space and interpret deviations from a fixed reference as fault indicators. When Operating Conditions (OCs) vary over time, this assumption breaks down: legitimate transitions trigger false alarms while moderate faults go undetected. We propose a spatiotemporal Graph Neural Network (GNN) framework that decomposes the normal operating space into OC-specific subspaces linked by transition functions, with a dual learning objective combining reconstruction loss and a Deep Support Vector Data Description (DeepSVDD) one-class term. The framework learns adjacency matrices through Graph Attention Networks (GATs) and integrates spatial modelling with temporal encoding to represent process dynamics under evolving OCs. This paper evaluates the foundational components of the framework — fault detection, spatial graph learning via GATv2, and temporal encoding — on the Tennessee Eastman Process (TEP) benchmark, with training performed exclusively on fault-free data. The spatiotemporal architecture achieves competitive detection performance from reconstruction error alone, with similarity-based feature selection improving both accuracy and graph structure diversity. We then evaluate the physical interpretability of the learned attention matrices against 22 ground-truth sensor pairs derived from the TEP control structure and process topology. The GATv2 attention does not recover all the necessary known physical pairs across multiple hyperparameter configurations, suggesting a structural limitation of reconstruction-driven attention rather than a tuning issue. This result challenges a common assumption in GNN-based FDD: that learned attention weights provide a basis for fault diagnosis and root-cause analysis. The architecture detects faults effectively, but the learned graph does not encode the physical topology needed for interpretable diagnosis, motivating physics-informed graph construction.
Hydrotreating is a process that uses existing refining facilities to upgrade waste tire pyrolysis oils, containing high amounts of bio-carbon, by reducing heteroatoms (e,g,. nitrogen, sulfur and oxygen) and aromatics content. A NiMo/ gamma-Al2O3 catalyst was used to assess hydrotreating reactivity differences between three tire pyrolysis oils and a petroleum straight run vacuum gas oil in a semibatch reactor at 140 bar for four hours at three different temperatures ranging from 352 degrees C to 397 degrees C. Hydrodenitrogenation performance for tire pyrolysis oils is higher than their vacuum gas oil counterpart despite their larger content of nitrogen in the feed. Hydrodesulfurization conversion is similar for all the studied tire pyrolysis oils with temperature having a stronger effect on vacuum gas oil. For hydrodeoxygenation, residual oxygen was low (0.1 wt%) for all tire pyrolysis oils. The conversion of aromatic carbon content for tire pyrolysis oils (10%-25%) is significantly reduced compared to vacuum gas oil (55 %). On a feedstock basis, vacuum gas oil, containing less light hydrocarbons (370 degrees C -), yields higher amounts of bottom residues (370 degrees C +). Meanwhile, tire pyrolysis oils produce more naphtha and kerosene products as temperatures increase. Hydrogen consumption calculation shows that TPOs require increased quantities of hydrogen compared to VGO.
Accurate modeling of complex industrial processes often relies on costly mechanistic simulations grounded in physical principles. In this paper, we investigate the subject of $\text{CO}_{2}$ capture, a major environmental challenge, through the absorption column of an amine-based post-combustion process. The modeling of such unit at industrial scale faces two difficulties: (i) theoretical models, efficient at laboratory scale, might fail to fully reflect the complexity of the numerous intertwined phenomena occurring in the absorber, (ii) the cost and uncertainty of industrial observation data make purely data-driven approaches unfeasible. To tackle both this low data regime and inaccurate physical models, we envision this $\text{CO}_{2}$ capture problem through the lens of Physics-informed Machine Learning (PiML). We present a hybrid (data+knowledge) model where the scarce observation data complement the physical model, while the latter ensures that the predictions remain physically consistent. Beyond the standard use of simulation data for learning and the embedding of physical laws as regularization, the originality of our PiML algorithm compared to other methods in the literature lies in a physical prior assumption about the network architecture and its countercurrent flow learning process inspired by the column's operation. Our experimental results showcase a significant improvement in accuracy and highlight the potential of our augmented model for generalizing across domains, especially when data is scarce.
A review of uncertainty quantification techniques is provided for a variety of situations involving uncertainties in model inputs (independent variables). The situations of interest are divided into three categories: (i) when model prediction uncertainties are quantified based on uncertainties in uncertain inputs, (ii) when parameter estimate uncertainties are calculated by propagation of uncertainties from measured inputs and outputs, and (iii) when model prediction uncertainties are quantified based on corresponding uncertainties in measured inputs and uncertain parameter estimates. For all three situations, linearization-based and Monte Carlo-based techniques are reviewed and details for their corresponding algorithms are presented. Recommendations are provided on which uncertainty quantification techniques are best for different types of chemical engineering models based on the amount of input uncertainty and nonlinearity over the range of plausible input and parameter values.
Error-in-variables model (EVM) methods are used for parameter estimation when independent variables are uncertain. During EVM parameter estimation, output measurement variances are required as weighting factors in the objective function. These variances can be estimated based on data from replicate experiments. However, conducting replicates is complicated when independent variables are uncertain. Instead, pseudo-replicate runs may be performed where the target values of inputs for repeated runs are the same, but the true input values may be different. Here, we propose a method to estimate output-measurement variances for use in multivariate EVM estimation problems, based on pseudo-replicate data. We also propose a bootstrap technique for quantifying uncertainties in resulting parameter estimates and model predictions. The methods are illustrated using a case study involving n-hexane hydroisomerization in a well-mixed reactor. Case-study results reveal that assumptions about input uncertainties can have important influences on parameter estimates, model predictions and their confidence intervals.
Hydrocracking is a crucial refinery process that transformsheavymolecules (i.e., vacuum gas oil (VGO)) into lighter and highly valuedproducts such as naphtha, kerosene, and diesel. It is a two-step process.The hydrotreatment (HDT) reactor uses a more robust catalyst, whichessentially serves to remove heteroatoms from the VGO feed in orderto satisfy product quality constraints and avoid poisoning the moredelicate zeolite-based HCK catalysts. The second hydrocracking (HCK)reactor uses a commercial zeolite catalyst with a carefully selectedbalance of acid and metallic sites. For hydrotreatment simulation,the kinetic model is decomposed in several ODEs (ordinary differentialequations). Catalyst vendors develop more and more catalysts. Foreach new catalyst (new generation), the kinetic parameters must berefitted. This task is costly and time-consuming. In this article,in order to reduce the required number of experimental points, a Bayesiantransfer approach is proposed to fit the parameters of catalyst (n + 1), using the past knowledge of catalyst (n) to add more information. A method for the choice of the prior isproposed and can be used for any type of parametric model. This approachis applied and shows an improvement in the prediction performanceand robustness compared to a classical fitting method. In our case,only 10 pilot plant points on catalyst (n + 1) arerequested to refit an HDN kinetic model.
This paper describes the development of product property models for a new hydrocracking catalyst, developed by IFPEN. The naphtha density is used as an illustrative example. A Bayesian Transfer Learning approach is used to transfer the ordinary least squares (OLS) parameters from a source dataset consisting of 3980 points from industrial units operating previous (n-1) catalyst generations to a target dataset with 96 points from pilot plant tests with the new catalyst (n). Robustness of the new model is greatly improved because information from the larger dataset is included in the new model. The Transfer Learning approach is shown to give good results and lead to a more robust model with respect to feedstock descriptors.