Chinese hamster ovary (CHO) cells are used to produce almost 90% of therapeutic monoclonal antibodies (mAbs) and antibody fusion proteins (Fc-fusion). The annotation of non-canonical translation events in these cellular factories remains incomplete, limiting our ability to study CHO cell biology and detect host cell protein (HCP) impurities in the final antibody drug product. We utilised ribosome footprint profiling (Ribo-seq) to identify novel open reading frames (ORFs) including N-terminal extensions and thousands of short ORFs (sORFs) predicted to encode microproteins. Mass spectrometry-based HCP analysis of eight commercial antibody drug products (7 mAbs and 1 Fc-fusion protein) using the extended protein sequence database revealed the presence of microprotein impurities. We present evidence that microprotein abundance varies with growth phase and can be affected by the cell culture environment. In addition, our work provides a vital resource to facilitate future studies of non-canonical translation and the regulation of protein synthesis in CHO cell lines.
Single cell RNA-seq (scRNA-seq) has recently been shown to provide a powerful method for the analysis of transcriptional heterogeneity in Chinese hamster ovary (CHO) cells. A potential drawback of current scRNA-seq platforms is that the cost can limit the complexity of experimental design and therefore the utility of the approach. In this manuscript, we report the use of oligonucleotide barcoding to perform multiplexed CHO cell scRNA-seq to study the impact of tunicamycin (TM), an inducer of the unfolded protein response (UPR). For this experiment, we treated a CHO-K1 GS cell line with 10μg/ml tunicamycin and acquired samples at 1, 2, 4 and 8 hr post-treatment as well as a non-treated TM-control. We transfected cells with sample-specific polyadenylated ssDNA oligonucleotide barcodes enabling us to pool all cells for scRNA-seq. The sample from which each cell originated was subsequently determined by the oligonucleotide barcode sequence. Visualisation of the transcriptome data in a reduced dimensional space confirmed that cells were not only separable by sample but were also distributed according to time post-treatment. These data were subsequently utilised to perform weighted gene co-expression analysis (WGCNA) and uncovered groups of genes associated with TM treatment. For example, the expression of one group of coexpressed genes was found to increase over the time course and were enriched for biological processes associated with ER stress. The use of multiplexed single cell RNA-seq has the potential to reduce the cost associated with higher sample numbers and avoid batch effects for future studies of CHO cell biology. Highlights ### Competing Interest Statement The authors have declared no competing interest. * CLR : center-log-ratio CHO : Chinese hamster ovary ER : Endoplasmic reticulum GO : Gene Ontology MAD : Median Absolute Deviations ME : Module Eigengene scRNA-seq : single cell RNA-seq SBO : short barcode oligonucleotide SFT : Scale Free Topology TOM : Topological Overlap Measure TM : Tunicamycin t-SNE : t-stochastic neighbour embedding UPR : unfolded protein response UMAP : uniform manifold approximation and projection WGCNA : weighted gene coexpression network analysis ssDNA : single stranded DNA
Affiliations 7 1National Institute for Bioprocessing Research and Training, Fosters Avenue, Blackrock, Co. Dublin, 8 Ireland. 9 2School of Chemical and Bioprocess Engineering, University College Dublin, Belfield, Dublin, Ireland. 10 3Bioprocess R&D, Pfizer Inc. Andover, Massachusetts, USA. 11 4National Institute for Cellular Biotechnology, Dublin City University, Dublin 9, Ireland. 12 5Barnett Institute, Northeastern University, 360 Huntington Ave, Boston, Massachusetts 02115, USA. 13
Chinese hamster ovary (CHO) cells are used to produce almost 90% of therapeutic monoclonal antibodies (mAbs). The annotation of non-canonical translation events in these cellular factories remains incomplete, limiting not only our ability to study CHO cell biology but also detect host cell protein (HCP) contaminants in the final mAb drug product. We utilised ribosome footprint profiling (Ribo-seq) to identify novel open reading frames (ORFs) including N-terminal extensions and thousands of short ORFs (sORFs) predicted to encode microproteins. Mass spectrometry-based HCP analysis of four commercial mAb drug products using the extended protein sequence database revealed the presence of microprotein impurities for the first time. We also show that microprotein abundance varies with growth phase and can be affected by the cell culture environment. In addition, our work provides a vital resource to facilitate future studies of non-canonical translation as well as the regulation of protein synthesis in CHO cell lines.
Mass spectrometry (MS) has emerged as a powerful approach for the detection of Chinese hamster ovary (CHO) cell protein impurities in antibody drug products. The incomplete annotation of the Chinese hamster genome, however, limits the coverage of MS-based host cell protein (HCP) analysis. In this study, we performed ribosome footprint profiling (Ribo-seq) of translation initiation and elongation to refine the Chinese hamster genome annotation. Analysis of these data resulted in the identification of thousands of previously uncharacterised non-canonical proteoforms in CHO cells, such as N-terminally extended proteins and short open reading frames (sORFs) predicted to encode for microproteins. MS- based HCP analysis of adalimumab with the extended protein sequence database, resulted in the detection of CHO cell microprotein impurities in a mAb drug product for the first time. Further analysis revealed that the CHO cell microprotein population is altered over the course of cell culture and in response to a change in cell culture temperature. The annotation of non-canonical Chinese hamster proteoforms permits a more comprehensive characterisation of HCPs in antibody drug products using MS. Highlights Analysis of translation initiation and elongation using ribosome footprint profiling provides a refined annotation of the Chinese hamster genome. 7,769 novel Chinese proteoforms were identified including those initiating at near cognate start codons. 941 N-terminal extensions of annotated genes were identified. 5,553 short open reading frames (sORFs) predicted to encode microproteins (i.e., proteins < 100 aa) were also characterised. The annotation of non-canonical proteins increases the coverage of MS-based host-cell protein analysis in monoclonal antibody drug products. 8 microproteins were found in adalimumab drug product. Transcripts annotated as non-coding can contain short open reading frames (sORFs) predicted to encode peptides (or microproteins) which are found to undergo changes in expression and translational regulation at reduced cell culture temperature. 95 of the novel proteoforms of which 79 were microproteins were subsequently identified in a second CHO K1 cell line using LC-MS/MS based proteomics. A comparison of protein abundance revealed that 13 microproteins were found to be differentially expressed between the exponential growth and stationary phases of cell culture.
A variety of mechanisms including transcriptional silencing, gene copy loss, and increased susceptibility to cellular stress have been associated with a sudden or gradual loss of monoclonal antibody (mAb) production in Chinese hamster ovary (CHO) cell lines. In this study, we utilized single-cell RNA-seq (scRNA-seq) to study a clonally derived CHO cell line that underwent production instability leading to a dramatic reduction of the levels of mAb produced. From the scRNA-seq data, we identified subclusters associated with variations in the mAb transgenes and observed that heavy chain gene expression was significantly lower than that of the light chain across the population. Using trajectory inference, the evolution of the cell line was reconstructed and was found to correlate with a reduction in heavy and light chain gene expression. Genes encoding for proteins involved in the response to oxidative stress and apoptosis were found to increase in expression as cells progressed along the trajectory. Future studies of CHO cell lines using this technology have the potential to dramatically enhance our understanding of the characteristics underpinning efficient manufacturing performance as well as product quality.
Our ability to study Chinese hamster ovary (CHO) cell biology has been revolutionised over the last decade following the development of next generation sequencing technology and publication of reference DNA sequences for CHO cells and the Chinese hamster. RNA sequencing has not only enabled the association of transcript expression with bioreactor conditions and desirable bioprocess phenotypes but played a key role in the characterisation of protein coding and small noncoding RNAs. The annotation of long noncoding RNAs, and therefore our understanding of their role in CHO cell biology, has been limited to date. In this manuscript, we use high-resolution RNASeq data to more than double the number of annotated lncRNA transcripts for the CHO K1 genome. In addition, the utilisation of strand-specific sequencing enabled the identification of more than 1,000 new antisense and divergent lncRNAs. The utility of monitoring lncRNA expression is demonstrated through an analysis of the transcriptomic response to a reduction of cell culture temperature and identification of simultaneous sense/antisense differential expression for the first time in CHO cells. To enable further studies of lncRNAs, the transcripts annotated in this study have been made available for the CHO cell biology community.
RNA sequencing (RNASeq) has been widely used to associate alterations in Chinese hamster ovary (CHO) cell gene expression with bioprocess phenotypes; however, alternative messenger RNA (mRNA) splicing, has thus far, received little attention. In this study, we utilized RNASeq for transcriptomic analysis of a monoclonal antibody (mAb) producing CHO K1 cell line subjected to a temperature shift. More than 2,465 instances of differential splicing were observed 24 hr after the reduction of cell culture temperature. A total of 1,197 of these alternative splicing events were identified in genes where no changes in abundance were detected by standard differential expression analysis. Ten examples of alternative splicing were selected for independent validation using quantitative polymerase chain reaction in the mAb-producing CHO K1 cell line used for RNASeq and a further two CHO K1 cell lines. This analysis provided evidence that exon skipping and mutually exclusive splicing events occur in genes linked to the cellular response to changes in temperature and mitochondrial function. While further work is required to determine the impact of these changes in mRNA sequence on cellular phenotype, this study demonstrates that alternative splicing analysis can be utilized to gain a deeper understanding of post-transcriptional regulation in CHO cells during biopharmaceutical production.
A regulatory mechanism that limits the number of complete protein molecules that can be synthesized from a single mRNA molecule of the human AMD1 gene encoding adenosylmethionine decarboxylase 1. The translation of messenger RNA (mRNA) sequences into proteins is regulated at many levels. Pavel Baranov and colleagues describe a new regulatory mechanism that occurs on the AMD1 mRNA and is conserved in vertebrates. They find that ribosomes occasionally read through the stop codon into the 3′ untranslated region. When these ribosomes encounter the next in-frame stop codon they halt but remain attached, leading to a ribosome queue. When this queue backs up over the true stop codon, the mRNA can no longer be translated. The authors propose that this mechanism serves to give each mRNA a restricted lifetime, to prevent aberrant translation if damage were to accumulate in the coding region over time. In addition to acting as template for protein synthesis, messenger RNA (mRNA) often contains sensory sequence elements that regulate this process1,2. Here we report a new mechanism that limits the number of complete protein molecules that can be synthesized from a single mRNA molecule of the human AMD1 gene encoding adenosylmethionine decarboxylase 1 (AdoMetDC). A small proportion of ribosomes translating AMD1 mRNA stochastically read through the stop codon of the main coding region. These readthrough ribosomes then stall close to the next in-frame stop codon, eventually forming a ribosome queue, the length of which is proportional to the number of AdoMetDC molecules that were synthesized from the same AMD1 mRNA. Once the entire spacer region between the two stop codons is filled with queueing ribosomes, the queue impinges upon the main AMD1 coding region halting its translation. Phylogenetic analysis suggests that this mechanism is highly conserved in vertebrates and existed in their common ancestor. We propose that this mechanism is used to count and limit the number of protein molecules that can be synthesized from a single mRNA template. It could serve to safeguard from dysregulated translation that may occur owing to errors in transcription or mRNA damage.
Although stop codon readthrough is used extensively by viruses to expand their gene expression, verified instances of mammalian readthrough have only recently been uncovered by systems biology and comparative genomics approaches. Previously our analysis of conserved protein coding signatures that extend beyond annotated stop codons predicted stop codon readthrough of several mammalian genes, all of which have been validated experimentally. Four mRNAs display highly efficient stop codon readthrough, and these mRNAs have a UGA stop codon immediately followed by CUAG (UGA_CUAG) that is conserved throughout vertebrates. Extending on the identification of this readthrough motif, we here investigated stop codon readthrough, using tissue culture reporter assays, for all previously untested human genes containing UGA_CUAG. The readthrough efficiency of the annotated stop codon for the sequence encoding vitamin D receptor (VDR) was 6.7%. It was the highest of those tested but all showed notable levels of readthrough. The VDR is a member of the nuclear receptor superfamily of ligand-inducible transcription factors and binds its major ligand, calcitriol, via its C-terminal ligand-binding domain. Readthrough of the annotated VDR mRNA results in a 67 amino-acid-long C-terminal extension that generates a VDR proteoform named VDRx. VDRx may form homodimers and heterodimers with VDR but, compared to VDR, VDRx displayed a reduced transcriptional response to calcitriol even in the presence of its partner retinoid X receptor. _______________________________________ INTRODUCTION Context dependent codon meaning enriches gene expression. Depending on the nature of relevant context features, the efficiency of specification of the alternative meaning can be set at widely different levels or be subject to regulatory influences. The majority of known occurrences of such dynamic redefinition of codon meaning involve UGA and UAG. Since, in the nearly universal genetic code, these codons usually specify translation termination, specification of an alternative meaning generally involves tRNA competition with release factor for their reading in the ribosomal A-site. In what is commonly http://www.jbc.org/cgi/doi/10.1074/jbc.M117.818526 The latest version is at JBC Papers in Press. Published on January 31, 2018 as Manuscript M117.818526 Copyright 2018 by The American Society for Biochemistry and Molecular Biology, Inc. by gest on July 5, 2018 hp://w w w .jb.org/ D ow nladed from by gest on July 5, 2018 hp://w w w .jb.org/ D ow nladed from by gest on July 5, 2018 hp://w w w .jb.org/ D ow nladed from by gest on July 5, 2018 hp://w w w .jb.org/ D ow nladed from Novel variant of the human vitamin D receptor 2 termed stop codon readthrough, a near-cognate tRNA performs the decoding with utility deriving from a proportion of the product having a C-terminal extension with an additional function. In these instances, the identity of the amino acid specified by the UGA or UAG is often, but not always, unimportant. However, when the non-universal amino acids, selenocysteine or pyrrolysine are specified, the selected features are these particular amino acids because of their distinctive properties. [Paradoxically in a species where the meaning of UGA, UAA and UAG has, throughout the body of coding sequences, been reassigned to specify amino acids, their meaning is dynamically redefined, in a context-dependent manner, to specify termination (1, 2).] Stop codon readthrough is well-known in viral decoding, especially of RNA viruses (3). Just as there are select organisms where RNA editing and ribosomal frameshifting are common, cephalopods (4) and Euplotes ciliates (5) respectively, so too stop codon readthrough is unusually common in Drosophila (6–8) and related insects (9). However, few instances of stop codon readthrough are known in vertebrate gene decoding. Until relatively recently hardly any instances of experimentally verified conserved mammalian readthrough were known (10, 11), although one of the reported occurrences is at least subject to substantial doubt (12, 13). Recent advances in sequencing technologies paved the way for the advent of ribosome profiling which has identified several potential human readthrough candidates (8, 14, 15). Sequencing advances have also propelled comparative genomics which led to the identification of seven mammalian mRNAs whose expression likely involves stop codon readthrough (7, 16, 17). Subsequent experimental analysis confirmed extended inframe decoding beyond the annotated stop codon (13, 17–21). Two of these mRNAs, ACP2 and SACM1L, have predicted RNA secondary structures immediately 3’ of their stop codons and 3’ structural elements are well known stimulators of functionally utilized stop codon readthrough (22–24). The four mRNAs with the highest readthrough efficiencies, in the tissue culture cells tested so far, are OPRL1, OPRK1, AQP4 and MAPK10. Their readthrough efficiencies range from 6-17% and all have UGA stop codons immediately followed by CUAG. For these four genes, this motif is conserved not only in mammals but throughout vertebrates and the importance of the UGA_CUAG motif was confirmed using a systematic mutagenesis approach (17). UGA_CUAG was also subsequently shown to promote readthrough in mRNAs encoding human malate and lactate dehydrogenases (17, 18, 20). Several earlier studies indicated that a cytidine 3′-adjacent to the stop codon influences readthrough in both prokaryotes and eukaryotes (25, 26) but subsequent studies showed that the termination context effect is not limited to a single 3’ nucleotide. The 3’ motif, CARYYA, can stimulate efficient readthrough, especially in plant viruses (27–29). In yeast, a similar sequence (CARNBA) can stimulate readthrough (30). Very recently, using reporters expressed in mammalian cell-lines, a comprehensive systematic mutagenesis study identified UGA_CUA among the most highly efficient autonomous readthrough signals (31). Indeed, several alphaviruses employ stop codon readthrough on UGA_CUAG including Middelburg, Ross river, Getah and also Chikungunya (24). Readthrough has also been identified on UGA_CUA in Mimivirus and Megavirus which are the best characterized representatives of an expanding new family of giant viruses infecting Acanthamoeba (32). A search of all human genes for CUAG immediately following a UGA stop codon indicated that there are 23 instances. Four have positive evolutionary coding potential, as measured by PhyloCSF (33) and these are the four candidates we previously confirmed (17). However, functional readthrough cannot be ruled out for those genes with UGA_CUAG and negative PhyloCSF scores. In fact, readthrough of both malate and lactate dehydrogenases (both harboring UGA_CUAG and both having negative PhyloCSF scores) allows translation of a short peroxisome-targeting motif which has been verified experimentally (18, 20). Here, we investigated stop codon readthrough in all previously untested human mRNAs with UGA_CUAG. Consistent with our previous study showing that UGA-CUA alone can support ~1.5% readthrough (17), all candidates by gest on July 5, 2018 hp://w w w .jb.org/ D ow nladed from Novel variant of the human vitamin D receptor 3 tested here displayed levels of readthrough ranging from ~1.3 – 6.7%. The mRNA encoding the vitamin D receptor (VDR) displayed the highest level of readthrough in this study and was selected for further investigation, however, several other mRNAs, including ATP10D, CDH23, DDX58, SIRPB1 and TMEM86B also display highly efficient readthrough (~5.0%). The VDR is a member of the nuclear receptor superfamily of ligand-inducible transcription factors. While it is expressed in most tissues, it is most abundant in bone, intestine, kidney and the parathyroid gland. Consistent with its role as a transcription factor, it’s expression in tissues and tissue culture cells is low (34). Calcitriol (or 1α,25dihydroxyvitamin D3) is the ligand for the VDR which mediates the actions of the hormone by ligand-inducible heterodimerization with its partner, retinoid X receptor (RXR). Insufficient concentrations of either calcitriol or the VDR impair calcium and phosphate absorption and hypocalcemia develops which can develop into either rickets in children or else osteomalacia in adults. Dietary vitamin D deficiency is the most common cause of rickets and osteomalacia worldwide. Here we identify a C-terminally extended proteoform of the VDR generated by stop codon readthrough and investigate the effect of this extension on VDR function. RESULTS Following identification of the UGA_CUAG readthrough motif (17), searches of all human mRNAs for CUAG immediately following a UGA stop codon identified 23 instances (Supp. Table 1). This is a significant depletion of this combination of four nucleotides compared to expectations based on the frequencies of the individual nucleotides in those positions immediately following a UGA stop codon (39 expected, one-sided binomial p-value 0.004). Six of these were previously described and shown by us and others to promote efficient readthrough (13, 17, 18, 20). To experimentally test the remaining 17 potential readthrough candidates, surrounding sequences were cloned in-frame between Renilla and firefly luciferase genes. Recently, we described a modification to the classical dual luciferase reporter system (35) that avoids potential distortions, sometimes observed using fused dual reporters, by incorporating ‘StopGo’ sequences on either side of the polylinker (13). The advantage is that reporter activities and/or stabilities are not influenced by the product/s of the test sequences. HEK293T cells were transfected and lysates assayed by dual luciferase assay. Readthrough efficiencies were determined by comparing relative luciferase activities (firefly/Renilla) of test constructs against controls for each construct in which the TGA stop codon is changed to TGG (Trp). All 17 stop codon cont
Biopharmaceuticals such as monoclonal antibodies have revolutionised the treatment of a variety of diseases. The production of recombinant therapeutic proteins, however, remains expensive due to the manufacturing complexity of mammalian expression systems and the regulatory burden associated with administrating these medicines to patients in a safe and efficacious manner. In recent years, academic and industrial groups have begun to develop a greater understanding of the biology of host cell lines, such as Chinese hamster ovary (CHO) cells and utilise that information for process development and cell line engineering. In this review, we focus on ribosome footprint profiling (RiboSeq), an exciting next generation sequencing (NGS) method that provides genome-wide information on translation, and discuss how its application can transform our understanding of therapeutic protein production.
Translation initiation is typically restricted to AUG codons, and scanning eukaryotic ribosomes inefficiently recognize near-cognate codons. We show that queuing of scanning ribosomes behind a paused elongating ribosome promotes initiation at upstream weak start sites. Ribosomal profiling reveals polyamine-dependent pausing of elongating ribosomes on a conserved Pro-Pro-Trp (PPW) motif in an inhibitory non-AUG-initiated upstream conserved coding region (uCC) of the antizyme inhibitor 1 (AZIN1) mRNA, encoding a regulator of cellular polyamine synthesis. Mutation of the PPW motif impairs initiation at the uCC's upstream near-cognate AUU start site and derepresses AZIN1 synthesis, whereas substitution of alternate elongation pause sequences restores uCC translation. Impairing ribosome loading reduces uCC translation and paradoxically derepresses AZIN1 synthesis. Finally, we identify the translation factor eIF5A as a sensor and effector for polyamine control of uCC translation. We propose that stalling of elongating ribosomes triggers queuing of scanning ribosomes and promotes initiation by positioning a ribosome near the start codon.
•Ribosome footprint profiling (RiboSeq) has illuminated translation in exquisite detail.•A range of RiboSeq derived methods is available to study ribosomal elongation and initiation.•RiboSeq is poised to enable enhanced translation control in mammalian cell factories for biopharmaceutical production.
Although stop codon readthrough is used extensively by viruses to expand their gene expression, verified instances of mammalian readthrough have only recently been uncovered by systems biology and comparative genomics approaches. Previously, our analysis of conserved protein coding signatures that extend beyond annotated stop codons predicted stop codon readthrough of several mammalian genes, all of which have been validated experimentally. Four mRNAs display highly efficient stop codon readthrough, and these mRNAs have a UGA stop codon immediately followed by CUAG (UGA_CUAG) that is conserved throughout vertebrates. Extending on the identification of this readthrough motif, we here investigated stop codon readthrough, using tissue culture reporter assays, for all previously untested human genes containing UGA_CUAG. The readthrough efficiency of the annotated stop codon for the sequence encoding vitamin D receptor (VDR) was 6.7%. It was the highest of those tested but all showed notable levels of readthrough. The VDR is a member of the nuclear receptor superfamily of ligand-inducible transcription factors, and it binds its major ligand, calcitriol, via its C-terminal ligand-binding domain. Readthrough of the annotated VDR mRNA results in a 67 amino acid-long C-terminal extension that generates a VDR proteoform named VDRx. VDRx may form homodimers and heterodimers with VDR but, compared with VDR, VDRx displayed a reduced transcriptional response to calcitriol even in the presence of its partner retinoid X receptor.
Abundant evidence for translation within the 5' leaders of many human genes is rapidly emerging, especially, because of the advent of ribosome profiling. In most cases, it is believed that the act of translation rather than the encoded peptide is important. However, the wealth of available sequencing data in recent years allows phylogenetic detection of sequences within 5' leaders that have emerged under coding constraint and therefore allow for the prediction of functional 5' leader translation. Using this approach, we previously predicted a CUG-initiated, 173 amino acid N-terminal extension to the human tumour suppressor PTEN. Here, a systematic experimental analysis of translation events in the PTEN 5' leader identifies at least two additional non-AUG-initiated PTEN proteoforms that are expressed in most human cell lines tested. The most abundant extended PTEN proteoform initiates at a conserved AUU codon and extends the canonical AUG-initiated PTEN by 146 amino acids. All N-terminally extended PTEN proteoforms tested retain the ability to downregulate the PI3K pathway. We also provide evidence for the translation of two conserved AUG-initiated upstream open reading frames within the PTEN 5' leader that control the ratio of PTEN proteoforms.