Shake flasks are a foundational tool in early process development by allowing high throughput exploration of the design space. However, lack of online data at this scale can hamper rapid decision making. Oxygen transfer rate (OTR) monitoring has been readily applied as an online process characterization tool at the benchtop bioreactor scale. Recent advances in modern sensing technology have allowed OTR monitoring to be available at the shake flask level. It is now possible to multiplex time-of-action (e.g., Induction, temperature shift, pH shift, feeding initiation, point of harvest) characterization studies by relying on careful analysis of OTR profile kinetics. As a result, there is potential to save time and capital expenditures while exploring process intensification studies though accurate and physiologically relevant online data. In this article, we detail the application of OTR monitoring to characterize the impact that recombinant protein production has on an inducible CHO cell line expressing Palivizumab. We then test out time-of-action studies to intensify protein production outcomes. We observe that recombinant protein expression causes a metabolic load that diminishes potential biomass growth. As a result, when compared to a control standard process, delaying induction and temperature shift has the potential to improve viable cell densities (VCD) by 2-fold thus increasing recombinant protein yield by over 30 %. The study also demonstrates that OTR can serve as a useful tool to detect cessation of exponential growth. Consequently, time-of-action points that are characteristic of inducible systems can be formulated accurately and reliably to maximize production performance.
Key hydrodynamic-related parameters such as volumetric power input (P/V), impeller configuration, aeration strategy, and maximum gas sparge rate, as well as an appropriate feeding strategy, must be carefully selected to improve production yields in bioreactor. In this study, the feeding regimen was found to have an important impact on cell growth and productivity of a cumate-inducible CHO fed-batch cell culture. A low-volume feeding regimen avoided a rapid increase in osmolality, allowing for prolonged cell viability and a 33% increase in volumetric titer compared to the high-volume feeding regimen. Both sparged air and oxygen were used for dissolved oxygen (DO) control, utilizing three levels of airflow rates. An optimum airflow rate of 0.0031 vvm was found to improve cell growth, longevity, and thus final titer. A larger air cap required increased gas flow rates, which led to an earlier cell mortality. Scale-up from 1-L to 10-L bioreactor using constant P/V and air cap volumetric gas flow rate (vvm) allowed for comparable cell growth and productivity. Further investigation of the effect of mixing and aeration was done by maintaining P/V and vvm constant throughout the cell culture, which further improved product titers at 11 days after induction. Our study also demonstrates that keeping a constant volume by removing a culture amount equal to the feed volume added at each sampling event can significantly improve the final volumetric titer. This finding shows the benefit of developing a concentrated feed to reduce the volume increase, which in turn could greatly ease the scale-up task.
One strategy to enhance the production of biological therapeutics is using transient perfusion in the preculture (N-1 stage) to seed the production culture (N stage) at ultra-high cell densities (>10 x 106 viable cells/mL). This very high seeding density improves cell culture performance by shortening the timeline and/or achieving higher final product concentrations. Typically, an N-1 seed train employs bioreactors with alternating tangential flow filtration (ATF) or tangential flow filtration (TFF) perfusion systems or Wave cell bag bioreactor with integrated filtration membrane, which have costs and technical complexity. Here, we propose an alternative method using semi-continuous transient perfusion through media exchange in shake flasks, which is suitable for benchtop-scale intensification process development. Daily media exchange was necessary to prevent nutrient limitations. The observed limitation of maximum viable cell densities (VCD) in various flask sizes was demonstrated to be due to oxygen limitations through the measurements of maximum oxygen transfer rates (OTR) using the sulfite system. By increasing agitation frequency from 200 to 300 RPM, maximum OTR in 500-mL shake flasks was increased by 62.3%, allowing an increase in maximum VCD of 29.6%. However, in 1000-mL shake flasks, an increase in agitation rate resulted in early cell death. After demonstrating that media exchange in shake flasks by centrifugation had no significant impact on cell growth rates, metabolism, and productivity, a benchtop bioreactor was seeded from semi-continuous transient perfusion cell expansion. The ultra-high cell density seeding resulted in a 49.3% increase in space-time-yield (STY) when compared to a standard low seeding density culture.
Fed-batch recombinant therapeutic protein (RTP) production processes utilizing Chinese Hamster Ovary (CHO) cells can take a long period of time (>10 days). Within this period, not all critical features may be measured routinely, and in fact, some are only measured once the process is terminated, complicating decision making. As a consequence, utilizing routine current day bioreactor online data to aid in next day predictions is a promising strategy for model predictive control-based feeding strategies. The article details the development of a proposed soft sensor that merges current day bioreactor online data and offline historical sampling data to generate predictions about the next day of the production process. This approach demonstrated the ability to track product titer, cell growth, key metabolites, and cumulative glucose consumption across the 17-day process with low normalized root mean squared error (nRMSE = 0.24) and low normalized mean absolute error (nMAE = 0.18) as well as high linearity with respect to ground data (average R2 = 0.97). It was also demonstrated that the same model architecture could effectively soft sense product titer and metabolic profiles (glucose, lactate, ammonia) without having sampling day's offline data as inputs to the model. This suggests that the proposed model could act as a true soft sensor of hard-to-determine variables such as the trimeric SARS-CoV-2 spike protein that relies on end-of-process measurements to acquire the data (labor-intensive semi-quantitative SDS-PAGE gels or ELISA assay). Instantaneous specific glucose consumption rates were also predicted and showed good agreement with experimental measurements, further offering opportunities for online glucose control.
The recent COVID-19 pandemic revealed an urgent need to develop robust cell culture platforms which can react rapidly to respond to this kind of global health issue. Chinese hamster ovary (CHO) stable pools can be a vital alternative to quickly provide gram amounts of recombinant proteins required for early-phase clinical assays. In this study, we analyze early process development data of recombinant trimeric spike protein Cumate-inducible manufacturing platform utilizing CHO stable pool as a preferred production host across three different stirred-tank bioreactor scales (0.75, 1, and 10 L). The impact of cell passage number as an indicator of cell age, methionine sulfoximine (MSX) concentration as a selection pressure, and cell seeding density was investigated using stable pools expressing three variants of concern. Multivariate data analysis with principal component analysis and batch-wise unfolding technique was applied to evaluate the effect of critical process parameters on production variability and a random forest (RF) model was developed to forecast protein production. In order to further improve process understanding, the RF model was analyzed with Shapley value dependency plots so as to determine what ranges of variables were most associated with increased protein production. Increasing longevity, controlling lactate build-up, and altering pH deadband are considered promising approaches to improve overall culture outcomes. The results also demonstrated that these pools are in general stable expressing similar level of spike proteins up to cell passage 11 (~31 cell generations). This enables to expand enough cells required to seed large volume of 200-2000 L bioreactor.
Protein expression from stably transfected Chinese hamster ovary (CHO) clones is an established but time-consuming method for manufacturing therapeutic recombinant proteins. The use of faster, alternative approaches, such as non-clonal stable pools, has been restricted due to lower productivity and longstanding regulatory guidelines. Recently, the performance of stable pools has improved dramatically, making them a viable option for quickly producing drug substance for GLP-toxicology and early-phase clinical trials in scenarios such as pandemics that demand rapid production timelines. Compared to stable CHO clones which can take several months to generate and characterize, stable pool development can be completed in only a few weeks. Here, we compared the productivity and product quality of trimeric SARS-CoV-2 spike protein ectodomains produced from stable CHO pools or clones. Using a set of biophysical and biochemical assays we show that product quality is very similar and that CHO pools demonstrate sufficient productivity to generate vaccine candidates for early clinical trials. Based on these data, we propose that regulatory guidelines should be updated to permit production of early clinical trial material from CHO pools to enable more rapid and cost-effective clinical evaluation of potentially life-saving vaccines.
The growing interest in the use of lentiviral vectors (LVs) for various applications has created a strong demand for large quantities of vectors. To meet the increased demand, we developed a high cell density culture process for production of LV using stable producer clones generated from HEK293 cells, and improved volumetric LV productivity by up to fivefold, reaching a high titer of 8.2 × 107 TU/mL. However, culture media selection and feeding strategy development were not straightforward. The stable producer clone either did not grow or grow to lower cell density in majority of six commercial HEK293 media selected from four manufacturers, although its parental cell line, HEK293 cell, grows robustly in these media. In addition, the LV productivity was only improved up to 53% by increasing cell density from 1 × 106 and 3.8 × 106 cells/mL at induction in batch cultures using two identified top performance media, even these two media supported the clone growth to 5.7 × 106 and 8.1 × 106 cells/mL, respectively. A combination of media and feed from different companies was required to provide diverse nutrients and generate synergetic effect, which supported the clone growing to a higher cell density of 11 × 106 cells/mL and also increasing LV productivity by up to fivefold. This study illustrates that culture media selection and feeding strategy development for a new clone or cell line can be a complex process, due to variable nutritional requirements of a new clone. A combination of diversified culture media and feed provides a broader nutrients and could be used as one fast approach to dramatically improve process performance.
Training Deep Learning (DL) models with missing labels is a challenge in diverse engineering applications. Missing value imputation methods have been proposed to try to address this problem, but their performance is affected with Massive Proportion of Missing Labels (MPML). This paper presents a approach for handling MPML in Multivariate Long-Term Time Series Forecasting. It is an two-step process where interpolation (using Gaussian Processes Regression (GPR) and domain knowledge from experts) and prediction model are separated to enable the integration of prior domain knowledge. First, a set of samples of the possible interpolation of the missing outputs are generated by the GPR based on the domain knowledge. Second, the observed input sensor data and interpolated labels from GPR are used to train the prediction model. We evaluated our approach with the development of a soft-sensor with one real datasets to forecast the biomass during recombinant adeno-associated virus (rAAV) production in bioreactors. Our experimental results demonstrate the potential of the approach through quantitative evaluation of the generated forecasts in a case that would be extremely difficult to train a DL model due to MPML.
Background: Chinese hamster ovary ( CHO) cells are extremely important host cells for recombinant DNA technology with their utility requiring optimization of growth. The ability to test conditions using in silico models of growing CHO cells can help advance the optimization of bioreactor conditions towards higher viable cell concentration and increased mAb production. Methods: A new kinetic model of CHO cell metabolism is presented, tested and provided in this publication. RNASeq data from CHO cells was used to guide the selection of major metabolic pathways that were included in the kinetic model. The kinetic model includes 37 completely described reactions processing 45 species (metabolites). This model is based on previously published kinetic characteristics for this system and was evaluated against metabolomics data for mAb producing CHO cells grown under two different feeding regimes. Results and Conclusion: This work provides a kinetic model for energy metabolism of CHO cells. Application of the model offers insights into possible causes of the different performance of two feeding strategies, thus suggesting possible problems and future optimization routes.
REOLYSIN® (pelareorep) is a proprietary isolate of the reovirus T3D (Type 3 Dearing) strain which is currently being tested in clinical trials as an anticancer therapeutic agent. Reovirus genomes are composed of ten segments of double-stranded ribonucleic acid (RNA) characterized by genome size: large (L1, L2, and L3), medium (M1, M2, and M3), and small (S1, S2, S3, and S4). The objective of this work was to evaluate the homogeneity and genetic stability of REOLYSIN®. Sanger sequencing (SS) performed on test articles derived from the Master Virus Bank (MVB) and Working Virus Bank (WVB) identified many modifications when compared to GenBank reference sequences. Massively parallel sequencing (MPS) using Roche-454 sequencing was performed on REOLYSIN® (100 L scale) and resulted in 69,821,115 bases and an average of 335 bases per read. Twenty-nine high confidence differences relative to the GenBank reference sequence were identified in REOLYSIN® by MPS. Of those, 27 were previously identified by SS in the virus bank-derived test articles. Of the remaining two nucleotide differences, one was predicted to be silent at the amino acid level (L3 genome-T3163C, codon 1054, 86 % of the population was “T” and 13 % of the population were reported as “C”). The other modification was in the noncoding region (M1 genome-A2284A to A2284G), and A2284G was present in 97 % of the population. The results obtained from MPS were comparable to those from SS; both demonstrate a high level of homogeneity at the amino acid level and genetic stability of REOLYSIN®. Finally, phylogenetic analysis of the REOLYSIN® L1 genome segment showed close evolutionary relationship with its human homologs, serotypes Lang and Dearing.
Adenovirus production is currently operated at low cell density because infection at high cell densities still results in reduced cell-specific productivity. To better understand nutrient limitation and inhibitory metabolites causing the reduction of specific yields at high cell densities, adenovirus production in HEK 293 cultures using NSFM 13 and CD 293 media were evaluated. For cultures using NSFM 13 medium, the cell-specific productivity decreased from 3,400 to 150 vp/cell (or 96% reduction) when the cell density at infection was increased from 1 to 3 x 10(6) cells/mL. In comparison, only 50% of reduction in the cell-specific productivity was observed under the same conditions for cultures using CD 293 medium. The effect of medium osmolality was found critical on viral production. Media were adjusted to an optimal osmolality of 290 mOsm/kg to facilitate comparison. Amino acids were not critical limiting factors. Potential limiting nutrients including vitamins, energy metabolites, bases and nucleotides, or inhibitory metabolites (lactate and ammonia) were supplemented to infected cultures to further investigate their effect on the adenovirus production. Accumulation of lactate and ammonia in a culture infected at 3 x 10(6) cells/mL contributed to about 20% reduction of the adenovirus production yield, whereas nutrient limitation appeared primarily responsible for the decline in the viral production when NSFM 13 medium was used. Overall, the results indicate that multiple factors contribute to limiting the specific production yield at cell densities beyond 1 x 10(6) cells/mL and underline the need to further investigate and develop media for better adenoviral vector productions.
Capacitance measurements at one single frequency are already an established tool for the on-line monitoring of the viable and total cell density in animal cell culture processes. Recently available systems allow for the automatic measurement over a wide range of frequencies. As changes in the characteristic fall of capacitance with increasing frequency (β-dispersion) change with the cell size distribution, we hypothesized that we could get information about the cell size distribution from on-line capacitance measurements.
In this paper, we present a two-phase, hybrid model for generating training data for Named Entity Recognition systems. In the first phase, a trained annotator labels all named entities in a text irrespective of type. In the second phase, naive crowdsourcing workers complete binary judgment tasks to indicate the type(s) of each entity. Decomposing the data generation task in this way results in a flexible, reusable corpus that accommodates changes to entity type taxonomies. In addition, it makes efficient use of precious trained annotator resources by leveraging highly available and cost effective crowdsourcing worker pools in a way that does not sacrifice quality.
Machine learning and text mining offer new models for text analysis in the humanities by searching for meaningful patterns across many hundreds or thousands of documents. In this study, we apply comparative text mining to a large database of 20th century Black Drama in an effort to examine linguistic distinctiveness of gender, race, and nationality. We first run tests on the plays of American versus non-American playwrights using a variety of learning techniques to classify these works, identifying those which are incorrectly classified and the features which distinguish the plays. We achieve a significant degree of performance in this cross-classification task and find features that may provide interpretative insights. Turning our attention to the question of gendered writing, we classify plays by male and female authors as well as the male and female characters depicted in these works. We again achieve significant results which provide a variety of feature lists clearly distinguishing the lexical choices made by male and female playwrights. While classification tasks such as these are successful and may be illuminating, they also raise several critical issues. The most successful classifications for author and character genders were accomplished by normalizing the data in various ways. Doing so creates a kind of distance from the text as originally composed, which may limit the interpretive utility of classification tools. By framing the classification tasks as binary oppositions (male/female, etc), the possibility arises of stereotypical or “lowest common denominator” results which may gloss over important critical elements, and may also reflect the experimental design. Text mining opens new avenues of textual and literary research by looking for patterns in large collections of documents, but should be employed with close attention to its methodological and critical limitations.
The Encyclopedie of Denis Diderot and Jean le Rond d'Alembert was one of the most important and revolutionary intellectual products of the French Enlightenment. Mobilizing many of the great – and the notsogreat – philosophes of the 18th century, the Encyclopedie was a massive reference work for the arts and sciences, which sought to organize and transmit the totality of human knowledge while at the same time serving as a vehicle for critical thinking. In its digital form, it is a highly structured corpus; some 55,000 of its 77,000 articles were labeled with classes of knowledge by the editors making it a perfect sandbox for experiments with supervised learning algorithms. In this study, we train a Naive Bayesian classifier on the labeled articles and use this model to determine class membership for the remaining articles. This model is then used to make binary comparisons between labeled texts from different classes in an effort to extract the most important features in terms of class distinction. Reapplying the model onto the original classified articles leads us to question our previous assumptions about the consistency and coherency of the ontology developed by the Encyclopedists. Finally, by applying this model to another corpus from 18th century France, the Journal de Trevoux, or Memoires pour l'Histoire des Sciences & des BeauxArts, new light is shed on the domain of Literature as it was understood and defined by 18th century writers.
Reolysin®, a human reovirus type 3, is being evaluated in the clinic as an oncolytic therapy for various types of cancer. To facilitate the optimization and scale-up of the current process, a high performance liquid chromatography (HPLC) method has been developed that is rapid, specific and reliable for the quantification of reovirus type 3 particles. Using an anion-exchange column, the intact virus eluted from the contaminants in 9.78min at 350mM NaCl in 50mM HEPES, pH 7.10 in a total analysis time of 25min. The virus demonstrated a homogenous peak with no co-elution of other compounds as analyzed by photodiode array analysis. The HPLC method facilitated the optimization of the purification process which resulted in the improvement of both total and infectious particle recovery and contributed to the successful scale-up of the process at the 20L, 40L and 100L production scale. The method is suitable for the analysis of crude virus supernatants, crude lysates, semi-purified and purified preparations and therefore is an ideal monitoring tool during process development and scale-up.