The biopharmaceutical industry is transitioning towards more efficient and cost-effective production methods, driven by the need for more affordable treatments. Process intensifications techniques, such as perfusion are the means by which the biopharmaceutical industry is trying to lower costs, enhance productivity and product quality, and reduce facility footprints. The dynamic nature of perfusion processes presents considerable challenges for real-time monitoring and control, requiring the advancement of process analytical technologies (PAT) with at-line, in-line, or on-line capabilities. Raman spectroscopy has emerged as a pivotal technology, providing real-time, noninvasive measurements of multiple analytes simultaneously, contingent upon the availability of sufficient data for calibration modeling. This study outlines the implementation of an automated data generation workflow for Raman calibration modeling within a high-throughput perfusion miniature bioreactor system, specifically the Ambr 250 HT Perfusion. Additionally, we demonstrate the effectiveness of Raman calibration models in monitoring various cell culture parameters within perfusion cultivations, spanning multiple cell lines and monoclonal antibody products. Finally, we present the feasibility of a Raman-based bleed-rate control system and how it compares to the conventional cell counter-based approach.
Abstract Cell culture bioprocess data are typically collected across many timepoints and batches, where numerous analytes covary with each other and, critically, with elapsed process time. This time dependence can inflate performance metrics and compromise the validity of multivariate models. We introduce time-adjusted performance evaluation (TAPE), a regression-agnostic validation technique that quantifies and separates time-driven from time-independent predictivity. TAPE pairs leave-one-group-out cross-validation with per-timepoint centering to decompose performance into between-timepoint (time-dependent) and within-timepoint (time-decoupled) parts by comparing predicted and observed deviations from each timepoint mean. Applying TAPE to orthogonal partial least squares models across five Chinese hamster ovary cell culture datasets (three Raman spectroscopy, one metabolomics, and one transcriptomics), several ostensibly strong models’ predictivity was largely explained by timepoint means alone. After removing between-timepoint variation, only models with sample–response relationships independent of time retained good predictivity. For Raman, only models for Raman-active analytes (glucose, lactate) remained predictive, whereas Raman-inactive ones (K+, NH4 +) did not. In the omics studies, the models for titer, viable cell density, growth rate, and death rates were predominantly time-driven. By quantifying time’s contribution to model performance, TAPE helps prevent misleadingly good performance metrics and supports more reliable multivariate modeling of time-series bioprocess data.
Integrating cell segmentation with tracking is critical for achieving a detailed and dynamic understanding of cellular behavior. This integration facilitates the study and quantification of cell morphology, movement, and interactions, offering valuable insights into developmental processes, drug response, and disease mechanisms. Traditional segmentation and tracking methods rely on labor-intensive annotations—such as full segmentation masks or bounding boxes for every cell in each frame—which severely limit throughput and scalability. PointTrack addresses these challenges by introducing a weakly-supervised pipeline that leverages sparse point annotations in the initial frame to automatically generate high-fidelity cell masks and maintain consistent object identities across subsequent frames. PointTrack begins with multiple sparse point annotations per cell in the first image of a time-lapse sequence. An interactive segmentation engine uses these annotations to delineate accurate cell boundaries, eliminating the need for exhaustive mask drawing. A dedicated point-tracking module then propagates the annotated points through the entire video, adapting to changes in cell appearance and motion with minimal user intervention. By reducing annotation effort to a few clicks per cell while preserving mask quality and identity continuity, PointTrack significantly accelerates the preparation of large-scale microscopy datasets. Extensive evaluation on two diverse time-lapse datasets—CTMC and DeepCell—demonstrates robust segmentation and reliable tracking performance under varying imaging modalities, cell densities, and noise conditions. Qualitative examples highlight precise delineation of overlapping cells and seamless trajectory recovery through transient occlusions. With streamlined annotation requirements and automated processing, PointTrack offers an effective and scalable solution for large-scale microscopy-based cell segmentation and tracking.
Accurate estimation of growth and metabolic rates is essential for understanding and optimizing bioprocesses, yet traditional methods often fail when faced with sparse or noisy concentration data. We present MetRaC, a probabilistic framework based on Bayesian inference and Nested Sampling that addresses these challenges by integrating biological knowledge directly into the model structure. The approach transforms raw concentration measurements into pseudo-concentrations that account for distortions caused by bioreactor volume changes (e.g., feed additions, sample withdrawals), and models metabolic rates as linear combinations of basis functions to yield continuous rate profiles from discrete data. Using in-silico simulations, we evaluated the framework under a range of experimental conditions and compared its performance with a conventional rate calculation method. We further analyzed the influence of key experimental design parameters - sampling frequency, sample volume, and measurement noise - on both rate estimation accuracy and concentration reconstruction quality. Results demonstrate that the proposed framework delivers accurate, robust metabolic rate estimates even under severe data sparsity and noise, offering a powerful tool for improving bioprocess characterization and optimization.
Morphological profiling is a common approach to investigate the modes of action (MOAs) of compounds. Most methods rely on fixed-cell assays, which provide only a single snapshot at a predefined time point and overlook the dynamic nature of cellular responses. In contrast, live-cell imaging tracks responses over time, offering deeper insight into compound-specific effects and mechanisms; however, time-series analysis of image data remains challenging due to limited analytical tools.We present Live Cell Temporal Profiling (LCTP), a workflow for morphological profiling of label-free live-cell time series data that yields interpretable, biologically relevant results. We showcase LCTP in an MOA classification study using label-free data. The workflow integrates established deep-learning components, cell segmentation, live/dead classification, and single-cell feature extraction, with data-driven models to capture MOA-specific temporal phenotypes and produce time-resolved profiles that can be compared across compounds and cell lines.We assess MOA classification performance using double-blinded cross-validation simulating a real-world screening scenario. LCTP significantly improves MOA classification over single–time point analysis, consistently across both cell lines used in the study. Time-resolved phenotypic modeling reveals transient, sustained, and delayed responses, clarifying compound-specific temporal effects and mechanisms across MOAs.The presented workflow is modular: each step removes irrelevant information, enriching signal, and enabling straightforward updates as technologies evolve and as new technologies become available, while supporting reuse across studies broadly. We believe LCTP adds substantial value to high-throughput compound screening, showing that live-cell imaging combined with this workflow yields informative visualizations of temporal effects and improved MOA classification.
Accurate cell tracking in microscopy is essential for studying biological dynamics like proliferation and migration. Traditional fully supervised methods demand dense pixel-wise masks for every frame, making them impractical for large-scale use. Recent methods like SAT reduce annotation effort by using sparse point-based supervision, but still require multiple positive and negative points per cell, which remains labor-intensive. BoxTrack offers a lightweight and annotation-efficient alternative, requiring only a single bounding box per cell in the first frame. Without relying on any point-level annotations, it performs end-to-end instance segmentation and tracking over entire sequences. This simplification leads to a substantial reduction in annotation cost while improving performance over SAT. On the CTMC dataset, BoxTrack improves Multiple Object Tracking Accuracy (MOTA) by +15.96 https://github.com/nabeelkhalid92/Box-it-Track-it .
Digital twins of mammalian cell cultures hold great potential for predictive bioprocess modeling, yet their development is challenged by the nonlinear dynamics and metabolic complexity of these systems. We present a hybrid computational framework that integrates mechanistic and data-driven modeling to construct predictive digital twins for Chinese hamster ovary (CHO) cell cultures producing monoclonal antibodies. The framework couples ordinary differential equation (ODE) models with constraint-based metabolic modeling and machine learning components trained on Bayesian-estimated metabolic rates. Applied to 23 CHO fed-batch cultures, viable cell density, product titer, and key metabolite concentrations are accurately predicted under varying feeding and media conditions within a unified simulation engine, where empirical variability is incorporated through multivariate statistical constraints derived from experimental data. Cross-validation analyses demonstrated strong generalization across process variations, highlighting the framework’s capacity to capture both biochemical constraints and adaptive cellular behavior. This hybrid modeling approach provides a mechanistically interpretable yet data-adaptive foundation for constructing bioprocess digital twins. By bridging statistical, mechanistic, and machine learning methodologies, it advances the computational representation of CHO cell culture systems and offers a generalizable strategy for predictive modeling in complex biological production processes.
Biopharmaceuticals are medical compounds derived from biological sources and are often manufactured by living cells, primarily Chinese hamster ovary (CHO) cells. CHO cells display variation among cell clones, leading to growth and productivity differences that influence the product's quantity and quality. The biological and environmental factors behind these differences are not fully understood. To identify metabolites with a consistent relationship to productivity or cell death over time, we analyzed the extracellular metabolome of 11 CHO clones with different growth and productivity characteristics over 14 days. However, in bioreactor processes, metabolic profiles and process variables are both strongly time-dependent, confounding the metabolite-process variable relationship. To address this, we customized an existing hierarchical approach for handling time dependency to highlight metabolites with a consistent correlation to a process variable over a selected timeframe. We benchmarked this new method against conventional orthogonal partial least squares (OPLS) models. Our hierarchical method highlighted several metabolites consistently related to productivity or cell death that the conventional method missed. These metabolites were biologically relevant; most were known already, but some that had not been reported in CHO literature before, such as 3-methoxytyrosine and succinyladenosine, had ties to cell death in studies with other cell types. The metabolites showed an inverse relationship with the response variables: those positively correlated with productivity were typically negatively correlated with the death rate, or vice versa. For both productivity and cell death, the citrate cycle and adjacent pathways (pyruvate, glyoxylate, pantothenate) were among the most important. In summary, we have proposed a new method to analyze time-dependent omics data in bioprocess production. This approach allowed us to identify metabolites tied to cell death and productivity that were not detected with traditional models.
Multiclass data sets and large-scale studies are increasingly common in omics sciences, drug discovery, and clinical research due to advancements in analytical platforms. Efficiently handling these data sets and discerning subtle differences across multiple classes remains a significant challenge. In metabolomics, two-class orthogonal projection to latent structures discriminant analysis (OPLS-DA) models are widely used due to their strong discrimination capabilities and ability to provide interpretable information on class differences. However, these models face challenges in multiclass settings. A common solution is to transform the multiclass comparison into multiple two-class comparisons, which, while more effective than a global multiclass OPLS-DA model, unfortunately results in a manual, time-consuming model-building process with complicated interpretation. Here, we introduce an extension of OPLS-DA for data-driven multiclass classification: orthogonal partial least squares-hierarchical discriminant analysis (OPLS-HDA). OPLS-HDA integrates hierarchical cluster analysis (HCA) with the OPLS-DA framework to create a decision tree, addressing multiclass classification challenges and providing intuitive visualization of interclass relationships. To avoid overfitting and ensure reliable predictions, we use cross-validation during model building. Benchmark results show that OPLS-HDA performs competitively across diverse data sets compared to eight established methods. This method represents a significant advancement, offering a powerful tool to dissect complex multiclass data sets. With its versatility, interpretability, and ease of use, OPLS-HDA is an efficient approach to multiclass data analysis applicable across various fields.
Morphological profiling is a powerful method for identifying the modes of action (MOAs) of chemical compounds. However, most approaches rely on fixed-cell assays that capture only a single time point, missing the dynamic nature of cellular responses. While live-cell imaging captures these temporal effects, its potential for MOA classification using high-dimensional features remains unexplored. We strategically convert large-scale time-series image data ( > 82,000 images) into interpretable phenotypic trajectories and show that label-free live-cell imaging captures meaningful temporal phenotypic signatures that improve MOA classification compared to single-time-point analyses. The results are consistent across two cell lines and six MOA groups. This workflow enables mechanistic insight from live-cell images and supports the growing interest in label-free methods, which offer advantages in preparation time and cost, and preserve the native state of cells. Our findings demonstrate that live-cell imaging, combined with data-driven analysis, offers a powerful tool for high-throughput drug screening. ### Competing Interest Statement P.J., J.T., and G.L. are employed by Sartorius. J.C.P. and O.S. declare ownership in Phenaros Pharmaceuticals. Swedish Research Council, https://ror.org/03zttf063, 2024-04576, 2024-03566 FORMAS, , 2022-00940 Swedish Cancer Foundation, , 22 2412 Pj 03 H Horizon Europe Grant Agreement, , 101057442 (REMEDI4ALL)
Single-cell RNA-seq methods can be used to delineate cell types and states at unprecedented resolution but do little to explain why certain genes are expressed. Single-cell ATAC-seq and multiome (ATAC + RNA) have emerged to give a complementary view of the cell state. It is however unclear what additional information can be extracted from ATAC-seq data besides transcription factor binding sites. Here, we show that ATAC-seq telomere-like reads counter-inituively cannot be used to infer telomere length, as they mostly originate from the subtelomere, but can be used as a biomarker for chromatin condensation. Using long-read sequencing, we further show that modern hyperactive Tn5 does not duplicate 9 bp of its target sequence, contrary to common belief. We provide a new tool, Telomemore, which can quantify nonaligning subtelomeric reads. By analyzing several public datasets and generating new multiome fibroblast and B-cell atlases, we show how this new readout can aid single-cell data interpretation. We show how drivers of condensation processes can be inferred, and how it complements common RNA-seq-based cell cycle inference, which fails for monocytes. Telomemore-based analysis of the condensation state is thus a valuable complement to the single-cell analysis toolbox.
Microscopic imaging plays a pivotal role in various fields of science and medicine, offering invaluable insights into the intricate world of cellular biology. At the heart of this endeavor lies the need for accurate identification and characterization of individual cells within these images. Deep learning-based cell segmentation, which involves delineating cells from complex microscopic images, is pivotal for cell analysis. It serves as the foundation for extracting meaningful information about cell morphology, spatial organization, and interactions. However, traditional deep-learning models for cell segmentation require extensive and expensive annotation masks for each cell in the image, posing a significant challenge. To address this issue, this study introduces CellBoxify, a novel pipeline that streamlines cell instance segmentation. Unlike traditional methods, CellBoxify operates solely on bounding box annotations, making it approximately seven times faster than manual segmentation mask annotation for each cell. The proposed approach’s effectiveness is evident in its performance on the LIVECell dataset, a well-known resource for cell segmentation research. Achieving 83.40
Adoptive cell therapy (ACT) requires the in vitro expansion of T cells, a process where currently several variables are poorly controlled. As the state and quality of the cells affects the treatment outcome, the lack of insight is problematic. To get a better understanding of the production process and its degrees of freedom, we have generated a multiome CD4 T cell single-cell atlas. We find in particular a JUNB+ epigenetic state, orthogonal to traditional CD4 T cell subtype categorization. This new state is present but overlooked in previous transcriptomic CD4 T cell atlases. We characterize it to be highly proliferative, having condensed and actively remodeled chromatin, and correlating with exhaustion. JUNB+ subsets are also linked to memory formation, as well as circadian rhythm, connecting several important processes into one state. To dissect JUNB regulation, we also derived a gene regulatory network (GRN) and developed a new explainable machine learning package, Nando. We propose potential upstream drivers of JUNB, verified by other atlases and orthogonal data. We expect our results to be relevant for optimizing in vitro ACT conditions as well as modulation of gene expression through novel gene editing.### Competing Interest StatementJ.T. is employed at Sartorius. I.S.M is employed at Umea university but partially funded by Sartorius. Other authors declare no conflict of interest.
Cells play a fundamental role in sustaining life by performing numerous functions crucial for the survival of living organisms. The detection of cells holds paramount importance in the validation and analysis of biological hypotheses, as it offers valuable insights into the behavior, function, diagnosis, and treatment of diseases. By accurately detecting and studying cells, researchers can unravel the complexities of cellular processes, leading to advancements in understanding diseases and the development of effective therapeutic interventions. In the domain of microscopic image analysis, substantial efforts have been devoted to the quantification of cells through segmentation masks and bounding boxes. However, these methods are time-consuming and resource-intensive. To tackle this challenge, we've introduced a novel approach focused on cell detection using solely their centerpoints. The proposed pipeline drastically cuts down on annotation efforts while still delivering commendable performance. By leveraging the proposed method, we aim to enhance efficiency in cell detection, paving the way for more expedient and resource-effective analysis in biological research and medical diagnostics.
Cell Painting is an established community-based microscopy-assay platform that provides high-throughput, high-content data for biological readouts. In November 2022, the JUMP-Cell Painting Consortium released the largest publicly available Cell Painting dataset with CellProfiler features, comprising more than 2 billion cell images. This dataset is designed for predicting the activity and toxicity of 115k drug compounds, with the aim to make cell images as computable as genomes and transcriptomes. In this context, our paper introduces a scalable and computationally efficient data analytics workflow created to meet the needs of researchers. This data-driven workflow facilitates the comparison of drug treatment effects through significant and biologically relevant insights. The workflow consists of two parts: first, the Equivalence score (Eq. score), a straightforward yet sophisticated metric highlighting relevant deviations from negative controls based on cell image morphology; second, the scalability of the workflow, by utilizing the Eq. scores on a large scale to predict and classify the subtle morphological changes in cell image profiles. By doing so, we show classification improvements compared to using the raw CellProfiler features on the CPJUMP1-pilot dataset on three types of perturbations. We hope that our workflow's contributions will enhance drug screening efficiency and streamline the drug development process. As this process is resource-intensive, every incremental improvement is valuable. Through our collective efforts in advancing the understanding of high-throughput image-based data, we aim to reduce both the time and cost of developing new, life-saving treatments.
Examining boiler failure causes is crucial for thermal power plant safety and profitability. However, traditional approaches are complex and expensive, lacking precise operational insights. Although data-driven approaches hold substantial potential in addressing these challenges, there is a gap in systematic approaches for investigating failure root causes with unlabeled data. Therefore, we proffered a novel framework rooted in data mining methodologies to probe the accountable operational variables for boiler failures. The primary objective was to furnish precise guidance for future operations to proactively prevent similar failures. The framework was centered on two data mining approaches, Principal Component Analysis (PCA) + K-means and Deep Embedded Clustering (DEC), with PCA + K-means serving as the baseline against which the performance of DEC was evaluated. To demonstrate the framework’s specifics, a case study was performed using datasets obtained from a waste-to-energy plant in Sweden. The results showed the following: (1) The clustering outcomes of DEC consistently surpass those of PCA + K-means across nearly every dimension. (2) The operational temperature variables T-BSH3rm, T-BSH2l, T-BSH3r, T-BSH1l, T-SbSH3, and T-BSH1r emerged as the most significant contributors to the failures. It is advisable to maintain the operational levels of T-BSH3rm, T-BSH2l, T-BSH3r, T-BSH1l, T-SbSH3, and T-BSH1r around 527 °C, 432 °C, 482 °C, 338 °C, 313 °C, and 343 °C respectively. Moreover, it is crucial to prevent these values from reaching or exceeding 594 °C, 471 °C, 537 °C, 355 °C, 340 °C, and 359 °C for prolonged durations. The findings offer the opportunity to improve future operational conditions, thereby extending the overall service life of the boiler. Consequently, operators can address faulty tubes during scheduled annual maintenance without encountering failures and disrupting production.
Cellular imaging plays a pivotal role in understanding various biological processes and diseases, making accurate cell segmentation indispensable for many biomedical applications. However, traditional methods for cell segmentation often rely on manual annotation, which is labor-intensive and time-consuming. Deep learning-based approaches for cell segmentation have shown promising results, but they require a vast amount of annotated data for training. In this context, this study presents CellGenie, an end-to-end pipeline designed to address the challenge of data scarcity in deep learning-based cell segmentation. This research proposes an innovative approach for automatic synthetic data generation tailored for microscopic image analysis. Leveraging the rich information provided by the LIVECell dataset, CellGenie generates synthetic microscopic images along with their corresponding segmentation masks for individual cells. By seamlessly integrating this synthetic data into the training process, this study enhances the performance of cell segmentation models beyond the limitations of existing annotated dataset. Furthermore, extensive experimentations are conducted to evaluate the efficacy of the generated data across various experimental scenarios. The results demonstrate the substantial impact of synthetic data generation in improving the robustness and generalization of cell segmentation models.
Cells are essential to life because they provide the functional, genetic, and communication mechanisms essential for the proper functioning of living organisms. Cell segmentation is pivotal for any biological hypothesis validation/analysis i.e., to get valuable insights into cell behavior, function, diagnosis, and treatment. Deep learning-based segmentation methods have high segmentation precision, however, need fully annotated segmentation masks for each cell annotated manually by the experts, which is very laborious and costly. Many approaches have been developed in the past to reduce the effort required to annotate the data manually and even though these approaches produce good results, there is still a noticeable difference in performance when compared to fully supervised methods. To fill that gap, a weakly supervised approach, PACE, is presented, which uses only the point annotations and the bounding box for each cell to perform cell instance segmentation. The proposed approach not only achieves 99.8% of the fully supervised performance, but it also surpasses the previous state-of-the-art by a margin of more than 4%.
Raman spectroscopy is widely used in monitoring and controlling cell cultivations for biopharmaceutical drug manufacturing. However, its implementation for culture monitoring in the cell line development stage has received little attention. Therefore, the impact of clonal differences, such as productivity and growth, on the prediction accuracy and transferability of Raman calibration models is not yet well described. Raman OPLS models were developed for predicting titer, glucose and lactate using eleven CHO clones from a single cell line. These clones exhibited diverse productivity and growth rates. The calibration models were evaluated for clone-related biases using clone-wise linear regression analysis on cross validated predictions. The results revealed that clonal differences did not affect the prediction of glucose and lactate, but titer models showed a significant clone-related bias, which remained even after applying variable selection methods. The bias was associated with clonal productivity and lead to increased prediction errors when titer models were transferred to cultivations with productivity levels outside the range of their training data. The findings demonstrate the feasibility of Raman-based monitoring of glucose and lactate in cell line development with high accuracy. However, accurate titer prediction requires careful consideration of clonal characteristics during model development.