Deep learning-based codon optimization with large-scale synonymous variant datasets enables generalized tunable protein expression

David A. Constant,Jahir M. Gutierrez,Anand V. Sastry,Rebecca Viazzo, Nicholas R. Smith, Jubair Hossain,David A. Spencer,Hayley Carter,Abigail B. Ventura, Michael T. M. Louie,Christa Kohnert,Rebecca Consbruck, Joshua Bennett, Kenneth A. Crawford,John M. Sutton,Anneliese Morrison,Andrea K. Steiger, Kerianne A. Jackson, Jennifer T. Stanton,Shaheed Abdulhaqq,Gregory Hannum,Joshua Meier, Matthew Weinstock,Miles Gander

biorxiv(2023)

引用 3|浏览25
暂无评分
摘要
Increasing recombinant protein expression is of broad interest in industrial biotechnology, synthetic biology, and basic research. Codon optimization is an important step in heterologous gene expression that can have dramatic effects on protein expression level. Several codon optimization strategies have been developed to enhance expression, but these are largely based on bulk usage of highly frequent codons in the host genome, and can produce unreliable results. Here, we develop deep contextual language models that learn the codon usage rules from natural protein coding sequences across members of the Enterobacterales order. We then fine-tune these models with over 150,000 functional expression measurements of synonymous coding sequences from three proteins to predict expression in E. coli . We find that our models recapitulate natural context- specific patterns of codon usage and can accurately predict expression levels across synonymous sequences. Finally, we show that expression predictions can generalize across proteins unseen during training, allowing for in silico design of gene sequences for optimal expression. Our approach provides a novel and reliable method for tuning gene expression with many potential applications in biotechnology and biomanufacturing. ### Competing Interest Statement The authors are current or former employees, contractors, or executives of Absci Corpo- ration and may hold shares in Absci Corporation. Methods and compositions described in this manuscript are the subject of one or more pending patent applications.
更多
查看译文
关键词
generalized tunable protein expression,codon optimization,synonymous variant datasets,learning-based,large-scale
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要