Rapid directed evolution guided by protein language models and epistatic interactions.
Saved in:
| Title: | Rapid directed evolution guided by protein language models and epistatic interactions. |
|---|---|
| Authors: | Tran, Vincent Q. (AUTHOR), Nemeth, Matthew (AUTHOR), Bartie, Liam J. (AUTHOR), Chandrasekaran, Sita S. (AUTHOR), Fanton, Alison (AUTHOR), Moon, Hyungseok C. (AUTHOR), Hie, Brian L. (AUTHOR), Konermann, Silvana (AUTHOR), Hsu, Patrick D. (AUTHOR) |
| Source: | Science. 5/7/2026, Vol. 392 Issue 6798, p1-13. 13p. |
| Subjects: | Epistasis (Genetics), Machine learning, Amino acid sequence, Artificial neural networks, Protein engineering, Biotechnology, Mutagenesis |
| Abstract: | Protein engineering is limited by the inefficient search through a high-dimensional sequence space to find combinations of synergistic mutations. Traditional approaches use stepwise mutation stacking, whereas machine learning methods require extensive datasets or multiple experimental rounds and are bottlenecked by costly, length-limited gene synthesis. We present MULTI-evolve (where MULTI stands for model-guided, universal, targeted installation of multimutants), a rapid evolution framework that systematically engineers multimutants. Our approach combines protein language models or existing functional data with epistatic modeling to predict synergistic combinations. Proposed multimutants are built through MULTI-assembly, a mutagenesis method enabling high-efficiency assembly across multikilobase sequences. Applying MULTI-evolve to three proteins achieved up to 10-fold improvements with a single round of machine learning–guided directed evolution. MULTI-evolve provides a streamlined approach for end-to-end, multimutant engineering for a broad range of protein types and functions. Editor's summary: Protein sequences are highly degenerate, with many possible functional sequences and complex interactions between pairs of otherwise neutral or deleterious mutations. Finding desired states in this space can be experimentally demanding, and there has been much interest in using protein language models to find improved variants. Tran et al. developed a machine learning framework for engineering multiple mutations at a time to access synergistic epistatic effects and accelerate traditional mutant screening. Combined with a tailored mutagenesis tool, this framework enables rapid mutation selection and assembly over large amounts of genetic material. The authors demonstrate the utility of their tool on a number of protein systems with applications in biotechnology. —Michael A. Funk INTRODUCTION: Proteins perform a variety of functions that can be enhanced or adapted into tools for research and medicine. The function of each protein is determined by its amino acid sequence of length N, where a single mutation can enhance its function, and multiple mutations can combine in ways that are difficult to predict, with both positive and negative effects. Protein engineering requires tediously searching through a vast landscape of 20N sequences. RATIONALE: Traditional directed evolution methods have improved the search for enhanced multimutants, but it requires performing slow, stepwise experimental mutation stacking. Conventional machine learning (ML)–guided methods are able to search larger sequence spaces, but the predictive power of the models require extensive experimental variant screening and many engineering rounds to accurately model the complex epistatic interactions for a given protein. Lastly, many methods rely on costly, length-limited gene synthesis to synthesize complex multimutants. We sought to overcome the challenges of existing methodologies and develop a universal approach for rapid evolution of any protein of interest. We reasoned that optimizing several key elements—protein language model (PLM)–based discovery, neural network–mediated epistatic modeling for multimutant extrapolation, and multisite mutagenesis—and integrating them into a single end-to-end workflow would enable the rapid design and synthesis of hyperactive multimutants. RESULTS: We developed MULTI-evolve (where MULTI stands for model-guided, universal, targeted installation of multimutants), an ML-guided end-to-end framework for engineering any protein of interest. MULTI-evolve efficiently identifies enhancing mutations with an ensemble of PLMs and uses epistatic modeling with double mutants and neural networks to extrapolate to enhanced multimutants. MULTI-evolve uses a single round of ML-guided directed evolution, as opposed to multiple rounds, and constructs the proposed multimutants using MULTI-assembly, which synthesizes multimutants containing up to nine mutations with up to 70% efficiency. Through computational benchmarking, we demonstrate that the PLM ensemble approach increases the discovery of function-enhancing mutations across 73 protein datasets and the neural network–based approach for multimutant extrapolation surpasses standard methods across 12 protein datasets. We applied MULTI-evolve to three distinct proteins, including engineered soybean ascorbate peroxidase (APEX), CRISPR-Cas13d, and an anti-CD122 antibody. We improved the catalytic activity of APEX greater than 100-fold and the trans-splicing activity of CRISPR-Cas13d by 10-fold. We perform multiobjective optimization of the anti-CD122 antibody to simultaneously improve the expression by sixfold and the binding affinity by threefold. Across all proteins, we demonstrate the ability of MULTI-evolve to learn from distinct sets of epistatic interactions and engineer multimutants. CONCLUSION: MULTI-evolve enables a new paradigm for protein engineering in which complex multimutants can be rapidly designed and synthesized for any protein of interest. MULTI-evolve enables engineering of enhanced multimutants with a single round of ML-guided directed evolution and can perform multiobjective optimization of therapeutically relevant properties. With advances in PLMs and computational oracles to predict protein-ligand interactions, MULTI-evolve will be able to incorporate additional functional information for nominating variants, enhancing its ability to sift through the vast sequence landscape and discover hyperfunctional proteins for biological research, biotechnology, and therapeutic development. MULTI-evolve enables rapid directed evolution of hyperactive multimutants.: A lab-in-the-loop framework combines PLMs to identify beneficial mutations, experimental screening of double mutants to capture epistatic interactions, and neural network modeling to propose optimized combinations. An integrated method for multisite mutagenesis, regardless of protein length, enables efficient gene synthesis for streamlined experimental validation. [ABSTRACT FROM AUTHOR] |
| Copyright of Science is the property of American Association for the Advancement of Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Psychology and Behavioral Sciences Collection |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | Protein engineering is limited by the inefficient search through a high-dimensional sequence space to find combinations of synergistic mutations. Traditional approaches use stepwise mutation stacking, whereas machine learning methods require extensive datasets or multiple experimental rounds and are bottlenecked by costly, length-limited gene synthesis. We present MULTI-evolve (where MULTI stands for model-guided, universal, targeted installation of multimutants), a rapid evolution framework that systematically engineers multimutants. Our approach combines protein language models or existing functional data with epistatic modeling to predict synergistic combinations. Proposed multimutants are built through MULTI-assembly, a mutagenesis method enabling high-efficiency assembly across multikilobase sequences. Applying MULTI-evolve to three proteins achieved up to 10-fold improvements with a single round of machine learning–guided directed evolution. MULTI-evolve provides a streamlined approach for end-to-end, multimutant engineering for a broad range of protein types and functions. Editor's summary: Protein sequences are highly degenerate, with many possible functional sequences and complex interactions between pairs of otherwise neutral or deleterious mutations. Finding desired states in this space can be experimentally demanding, and there has been much interest in using protein language models to find improved variants. Tran et al. developed a machine learning framework for engineering multiple mutations at a time to access synergistic epistatic effects and accelerate traditional mutant screening. Combined with a tailored mutagenesis tool, this framework enables rapid mutation selection and assembly over large amounts of genetic material. The authors demonstrate the utility of their tool on a number of protein systems with applications in biotechnology. —Michael A. Funk INTRODUCTION: Proteins perform a variety of functions that can be enhanced or adapted into tools for research and medicine. The function of each protein is determined by its amino acid sequence of length N, where a single mutation can enhance its function, and multiple mutations can combine in ways that are difficult to predict, with both positive and negative effects. Protein engineering requires tediously searching through a vast landscape of 20N sequences. RATIONALE: Traditional directed evolution methods have improved the search for enhanced multimutants, but it requires performing slow, stepwise experimental mutation stacking. Conventional machine learning (ML)–guided methods are able to search larger sequence spaces, but the predictive power of the models require extensive experimental variant screening and many engineering rounds to accurately model the complex epistatic interactions for a given protein. Lastly, many methods rely on costly, length-limited gene synthesis to synthesize complex multimutants. We sought to overcome the challenges of existing methodologies and develop a universal approach for rapid evolution of any protein of interest. We reasoned that optimizing several key elements—protein language model (PLM)–based discovery, neural network–mediated epistatic modeling for multimutant extrapolation, and multisite mutagenesis—and integrating them into a single end-to-end workflow would enable the rapid design and synthesis of hyperactive multimutants. RESULTS: We developed MULTI-evolve (where MULTI stands for model-guided, universal, targeted installation of multimutants), an ML-guided end-to-end framework for engineering any protein of interest. MULTI-evolve efficiently identifies enhancing mutations with an ensemble of PLMs and uses epistatic modeling with double mutants and neural networks to extrapolate to enhanced multimutants. MULTI-evolve uses a single round of ML-guided directed evolution, as opposed to multiple rounds, and constructs the proposed multimutants using MULTI-assembly, which synthesizes multimutants containing up to nine mutations with up to 70% efficiency. Through computational benchmarking, we demonstrate that the PLM ensemble approach increases the discovery of function-enhancing mutations across 73 protein datasets and the neural network–based approach for multimutant extrapolation surpasses standard methods across 12 protein datasets. We applied MULTI-evolve to three distinct proteins, including engineered soybean ascorbate peroxidase (APEX), CRISPR-Cas13d, and an anti-CD122 antibody. We improved the catalytic activity of APEX greater than 100-fold and the trans-splicing activity of CRISPR-Cas13d by 10-fold. We perform multiobjective optimization of the anti-CD122 antibody to simultaneously improve the expression by sixfold and the binding affinity by threefold. Across all proteins, we demonstrate the ability of MULTI-evolve to learn from distinct sets of epistatic interactions and engineer multimutants. CONCLUSION: MULTI-evolve enables a new paradigm for protein engineering in which complex multimutants can be rapidly designed and synthesized for any protein of interest. MULTI-evolve enables engineering of enhanced multimutants with a single round of ML-guided directed evolution and can perform multiobjective optimization of therapeutically relevant properties. With advances in PLMs and computational oracles to predict protein-ligand interactions, MULTI-evolve will be able to incorporate additional functional information for nominating variants, enhancing its ability to sift through the vast sequence landscape and discover hyperfunctional proteins for biological research, biotechnology, and therapeutic development. MULTI-evolve enables rapid directed evolution of hyperactive multimutants.: A lab-in-the-loop framework combines PLMs to identify beneficial mutations, experimental screening of double mutants to capture epistatic interactions, and neural network modeling to propose optimized combinations. An integrated method for multisite mutagenesis, regardless of protein length, enables efficient gene synthesis for streamlined experimental validation. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 00368075 |
| DOI: | 10.1126/science.aea1820 |