Danuglipron (PF-06882961), a potent, orally bioavailable small-molecule glucagon-like peptide-1 receptor (GLP-1R) agonist, is currently being developed for glycemic control among patients with Type-2 diabetes (J. Med. Chem. 2022, 65, 8208-8226; JAMA Netw. Open. 2023; 6(5): e2314493). The earlier synthesis of danuglipron suffered from chemoselective issues due to the competing nitrile hydrolysis in the final saponification step, which resulted in highly convoluted operations and extensive chromatographic purifications. We found that the methyl ester could be converted to trifluoroethyl ester, and the latter underwent hydrolysis to carboxylic acid in a much cleaner reaction profile. A thorough design of experiments (DOE) was conducted to expand the operating time window of the process to aid the process robustness during manufacturing. The improved process increased the yield by similar to 20% and reduced the process mass intensity (PMI) by 86%.
In vitro dissolution testing is a regulatory required critical quality measure for solid dose pharmaceutical drug products. Setting the acceptance criteria to meet compendial criteria is required for a product to be filed and approved for marketing. Statistical approaches for analyzing dissolution data, setting specifications and visualizing results could vary according to product requirements, company's practices, and scientific judgements. This paper provides a general description of the steps taken in the evaluation and setting of in vitro dissolution specifications at release and on stability.
Statistical design of experiments (DoE) is used to aid in the development and execution of preparative liquid chromatography (LC) for large-scale purification of active pharmaceutical ingredients (API) and pharmaceutical intermediates. Four purification case studies were undertaken. In case study 1, a normal phase preparative silica method is developed and modeled. After initial method screening, DoE results were used to set mobile phase composition, flowrate, and sample diluent. Of the three particle sizes studied (10 µm, 20 µm, 50 µm) only 10 µm silica resin was able to produce purified API at the yield (>96%) and productivity (> 1 kg/kg-resin/day) necessitated by the project. The second case study uses DoE studies to identify critical process parameters of column load, mobile phase solvent ratio and basic modifier level for a low-resolution, preparative, chiral separation. Trade-offs between purity, yield and productivity are quantified in a tight separation which made compromising on process outcomes a necessity. The third case study troubleshoots a loss of yield experienced during operation of a process-scale reverse-phase LC purification. DoE is used to identify a critical interaction between levels of acetonitrile and phosphoric acid in the mobile phase. An operating region which increased yield from around 85% to 97% was defined and implemented. The fourth case study was initially designed as a preparative chromatography purification of API. DoE was used to screen mobile phase solubility. These experiments uncovered conditions where API is soluble, and impurities are not. The solubility model in acetonitrile/water mixtures is further defined via a response surface DoE. The resulting targeted solvent mixture allows bulk purification via dissolution of API while three less-polar impurities remain in the solid phase and are removed by filtration. These four case studies demonstrate the efficiency of DoE and response surface modeling as tools for process development and optimization.
The Dynamic Response Surface Methodology (DRSM; Klebanov, N.; Georgakis, C. Ind. Eng. Chem. Res 2016, 55, 4022) is a data-driven modeling methodology for time-resolved measurements from Design of Experiments data sets. To extend its use to semi-infinite domains, a second version was proposed (Wang, Z. Y.; Georgakis, C. Ind. Eng. Chem. Res 2017, 56, 10770), which uses an exponential transformation of time as the independent variable. Here, we further improve upon this later version by imposing regression constraints, motivated by our qualitative understanding of the physical nature of the reaction system. We apply it to the modeling of several pharmaceutical case studies provided by our industrial collaborators. This constrained DRSM approach is able to model these compositional data more parsimoniously and with greater accuracy than the initial DRSM, eliminating the existence of oscillations near the end of some experiments from the model predictions.
The dynamic response surface methodology (DRSM) estimates an input–output for time-resolved measurements from design of experiments data sets. The improved DRSM algorithm, denoted by DRSM-2c, has been recently presented by Dong et al. In this paper, we examine the identification of the underlying reaction stoichiometry. Given a number of stoichiometries proposed as candidates, we show how the use of DRSM enhances the ability of target factor analysis (TFA) to identify the reaction stoichiometries active in the mixture, using both the parallel approach and the sequential approach. Besides a projection score, we introduce statistical tests in the use of TFA to better identify the correct stoichiometries. Based on different data sets provided by industrial collaborators, we show that the improved DRSM algorithm is helpful for the stoichiometric identification.
Statisticians at Pfizer who support Chemistry, Manufacturing, and Controls (CMC), and Regulatory Affairs (Reg CMC) have developed many statistical R-based computational tools to enable high efficiency, consistency, and fast turnaround in their routine statistical support to drug product and manufacturing process development. Most tools have evolved into web-based applications for convenient access by statisticians and colleagues across the company. These tools cover a wide range of areas, such as product stability and shelf life or clinical use period estimation, process parameter criticality assessment, and design space exploration through experimental design and parametric bootstrapping. In this article, the general components of these R-programmed web-based computational tools are introduced, and their successful applications are demonstrated through an application of estimating a drug product shelf life based on stability data.
Mathematical modeling of chemical reaction kinetics has been proven to aid the development of new reactions and processes. Chemical kinetic modeling is a well-established principle in chemical engineering that uses fundamental knowledge of the reaction mechanism to predict conversion data. In pharmaceutical drug development, elementary-type kinetics are hardly common, because of the nature of the complex organic reaction mixtures and low-level impurities. Thus, data-driven modeling plays an important role in understanding the relationship between reaction parameters and reaction profiles. Advances in reaction automation technologies, such as high-throughput platforms and autosamplers, enable greater data collection to enrich our understanding of chemical reactions. As a result, statistical analysis has shifted from conventional end-point analysis to modeling the entire reaction profile using more advanced statistical models. Data-driven approaches are especially useful in early stage of development where not enough time or material is available for a proper kinetic model development. For the same modeling task, regardless of the underlying approach, we strongly feel that a systematic process of model development needs to be applied. We developed a rigorous and general modeling workflow describing how to apply kinetic models and statistical models to a set of dynamic reaction data. In particular, a semiparametric model was applied. An industrial case study is presented with a methyl ester chemoselective hydrolysis reaction, to showcase the performance and robustness of the two modeling approaches and their impacts, side by side, on parameter effect estimation, reaction robustness range finding, and reaction optimization and operation window prediction. New and innovative visualization techniques are shared in this article for efficient data and model result interpretation.